CERES is proposed, a closed-loop multimodal indexing framework that builds a three-level semantic pyramid, mines implicit concepts via a co-occurrence-aware router, performs scale-routed cross-attention into a lightweight U-Net generator, and verifies coverage by re-indexing the generated image with the same frozen VLM...
Guangyuan Dong, Chuang Liu, Hao-Yu Wang et al.· 1 citation
Multimodal information systems increasingly route generated visual content back through the same vision-language index that informed its production, so the output must remain retrievable by the queries it was meant to serve. When the scene contains entities at vastly different scales, existing language-guided generator...
Guangyuan Dong, Chuang Liu, Yangchen Zeng et al.· 0 citations
DSF–MarianMT, a semantic fusion enhanced neural machine translation framework built upon MarianMT, is proposed, demonstrating that the proposed framework effectively enhances semantic representation learning and translation fidelity for Chinese–English neural machine translation.