This work presents Omni-Embed-Mini, a 0.9B-parameter model that maps text, speech, audio, images, video, and visually-rich documents into a single shared cosine space without updating any text-side parameter, and is competitive with the closed gemini-embedding-2, edging ahead of it on the overall-modality average.
Mohammed Irfan Kurpath, Jaseel Muhammad Kaithakkodan, Sahal Shaji Mullappilly et al.· 0 citations
An AVQA-specific two-stage framework that adapts established cross-modal reconstruction and dependency-modeling principles to the supervision constraints of TM-AVQA is developed, a setting in which modality-complete, audio-missing, and visual-missing samples may occur during both training and testing.
Jin-Xing Zhou, Zhangbin Li, Di Hu et al.· IEEE Transactions on Pattern...· 0 citations
An estimated 1 billion people worldwide live with vision impairment, yet current vision-language models (VLMs) produce descriptions too vague for safe navigation by blind and low-vision (BLV) users. Large VLMs can generate high-quality audio-description-compliant narrations but cannot run on mobile devices; small VLMs...
Rishabh C. Choudhary, S. Raj, Umesh Goyal et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.