The Document Question Answering (DocQA) task necessitates the synergistic interpretation of visual and textual information embedded within documents. Although Retrieval-Augmented Generation (RAG) has enhanced the capabilities of Large Vision-Language Models (LVLMs), existing approaches still encounter significant bottl...
Jia-Yuan Wang, Jie Lian, Fu Zhao et al.· Proceedings of the 32nd ACM...· 0 citations
Center-guided spectral diffusion is proposed, which replaces traditional alignment with generative modeling and outperforms state-of-the-art methods on several datasets and alleviates the noise amplification problem commonly found in traditional alignment methods.
Jia-Yuan Wang, Jie Lian, Yong-Quan Shi et al.· Proceedings of the Thirty-Fi...· 0 citations
A novel framework for multi-channel data fusion that integrates a multi-channel encoding module with continuous feature dynamics to initialize and evolve the latent node representations over static graph topologies, and effectively mitigates the performance degradation typically associated with deep graph architectures...
Na Song, Zi-Han Fang, Wei-Dong Zhang et al.· Proceedings of the Thirty-Fi...· 0 citations
Unsupervised person re-identification (USL-ReID) typically relies on clustering to generate pseudo-labels, but significant cross-view appearance variations often cause images of the same identity to be split into different clusters. Training on such noisy pseudo-labels severely degrades the learned representations. The...
Xuan Tan, Qi-Xian Zhang, Ding Qi et al.· IEEE Transactions on Image P...· 0 citations
This work introduces MMHRAG, a novel Multi-Modal Hierarchical Retrieval-Augmented Generation framework that achieves cross-modal interaction on DocQA for the first time, and designs a Summarizing Agent to resolve logical conflicts, information redundancy, and granularity discrepancies among retrieved cross-modal eviden...
Jiayuan Wang, Jie Lian, Fu Zhao et al.· Proceedings of the 32nd ACM...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.