Real-time full-duplex speech models can listen while speaking, enabling natural interaction without rigid turn boundaries. Existing benchmarks evaluate turn-taking, interruption handling and multi-round dialogue, but largely centre on a designated user rather than an assistant participating in a shared conversation amo...
Chengqian Ma, Wen-Hao Feng, Wei-Xuan Jin et al.· 0 citations
Unified multimodal models combine understanding, generation, and editing within a single network, offering a promising foundation for versatile multimodal applications. However, growing multimodal contexts make KV cache storage and access increasingly costly. Existing KV cache compression methods are typically tailored...
Wan-Qi Yang, Yue-Xiao Ma, Mei Xie et al.· 0 citations
Frame selection is essential for applying Large Multimodal Models (LMMs) to long videos due to severe frame redundancy and limited context windows. Since the appropriate frame budget varies with the downstream LMM, reasoning demands, and latency constraints, a practical selector should serve multiple budgets. However,...
Wang Chen, Yu Chen, Xiang Wang et al.· 0 citations
WaveZip is proposed, a joint signal-frequency-domain framework for efficient video inference that requires no task-specific training and can be seamlessly integrated into off-the-shelf LVLMs to boost inference efficiency.
Yuhui Zeng, Wang Chen, Jin-Fa Huang et al.· arXiv.org· 1 citation
TimePLE is proposed, which reformulates VTG from endpoint prediction to interval-native grounding by predicting a single joint distribution over valid temporal intervals, and curate 90K-scale grounded samples and human-verify 3K-scale benchmark annotations.
Yuhui Zeng, Xin-Yu Mao, Xiaokun Liu et al.· arXiv.org· 0 citations
YOLO-PEFT is proposed, a structure-aware framework that formulates adapter placement as an auditable constraint-planning problem that replaces manual target-module trial and error with explicit, inspectable planning while preserving verified train-save-merge-export paths.
Xuchen Lin, Wenjie Nie, Jinlong Peng et al.· 0 citations
This work introduces A 2 -Judger, a novel MLLM-based A gentic instantiation of A uto Judger equipped with semantic-aware retrieval and dynamic memory that significantly improves sample efficiency while maintaining reliable evaluation results.
Xuanwen Ding, Chengjun Pan, Zejun Li et al.· 0 citations
SDAR employs a symbolic reasoning engine to guide agentic decision making, connecting low-level visual cues with structured symbolic representations of events, and enables interpretable reasoning chains that capture causal relationships, contextual dependencies, and event categories.
Guangyao Chen, Liqin Luo, Jun Peng et al.· Visual Intelligence· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.