Semantic parsing (SP)-based knowledge base question answering aims to answer natural language questions by generating executable logical forms (LFs) over knowledge bases (KBs). When applying Large Language Models (LLMs) to this task, a key challenge over large, heterogeneous KBs is selecting question-related schema ele...
Guang-Ze Gao, Zi-Xuan Li, Si-Kui Zhang et al.· 0 citations
The integration of novel view synthesis (NVS) and open-vocabulary segmentation (OVS) has recently yielded powerful feed-forward 3D foundation models. However, their inherent reliance on static-scene assumptions leads to severe misalignment of spatial features in unconstrained dynamic environments. To bridge this critic...
MMAgent-R$^2$, an agentic mRAG framework that integrates visual reranking and active rejection as its internal verification mechanism, is proposed and achieves joint optimization of external retrieval, internal verification, and answer generation via GRPO training.
Tao Zhang, Ziqi Zhang, Zongyang Ma et al.· arXiv.org· 0 citations
Gaussian Mixture Modeling for Event-Aware Visual Allocation is proposed, which leverages Gaussian Mixture Models to model event-level structure from discrete frame-wise observations and achieves comparable performance to baseline selection methods while utilizing only approximately half of the visual token budget.
Yifan Lu, Ziqi Zhang, C. Yuan et al.· arXiv.org· 0 citations
A Semantic-Retrieval-Augmented Detector (SRA-Det) is proposed that uses an attention-based module to retrieve multiple semantic facets from token-level text features, and a soft-min matching rule that behaves like a differentiable logical AND over these facets, ensuring that all key attributes are satisfied.
Li Yang, Boyu Cai, Wei Liu et al.· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.