Jul 2026
TinyMem: Condensing Multimodal Memory for Long-Form Video Action Detection.
This paper introduces TinyMem, a model built upon compact multimodal memory for long-form video action detection that outperforms a range of state-of-the-art models on AVA v2.2 while using 5 times fewer memory tokens than the baseline with dense visual memory embeddings.
Rui Tian, Qi Dai, Hang-Rui Hu et al.
· IEEE Transactions on Pattern... · 0 citations