Skip to content

Boosting scene captioning with pairwise semantic spatial reasoning and contextualized large language model integration

Aug 2026 · Signal, Image and Video Processing · Vol 20 · 0 citations · 40 references
Computer Science

TL;DR

This work extends its previous work in scene graph generation to integrate image captioning, incorporating large language models to interpret scene graph structures and produce high-quality captions encompassing complex actions and relationships.

View source

Similar papers

Conference Jul 2026

3D vision-language question answering with explicit scene graphs and local topology priors

TA-LMM, a 3D visual question answering method built on explicit scene graphs and local topological priors, is proposed, suggesting that explicit local topological priors can improve scene consistency in 3D visual question answering.

Kaixin Wu, Kunlin Zhou, Boxin Li et al. · 0 citations
Preprint Aug 2026

Modeling Scientific Experiment Scenes: Dataset and Model

The Cross-Modal Dual-Path Generator (CM-DPG), a model for robust open-vocabulary SGG that enhances object-level semantic representations through joint visual-textual encoding and improves relational reasoning using complementary visual and geometric cues, is proposed.

Ming-Hao Zou, Qingtian Zeng, Shangkun Liu et al. · 0 citations

SPOT: Structured Prompting with Object-centric Tokens for open-world scene graphs

SPOT is introduced, a structured prompting framework that augments open-source VLMs with spatial reasoning abilities for scene graph generation with minimal training, and achieves competitive or superior relation prediction compared to large proprietary models.

Unknown authors · 0 citations
Conference Jul 2026

Multimodal Context-Enriched Visual Representation Learning for Enhanced Vision–Language Image Captioning

Image captioning models can produce rapid Sentences, without visual relationships, or insert non-existing plausible objects. A common cause is to compress image evidence into visual symbols that carry a weak neighbourhood context. The multimodal context-enhanced visual representation learning framework (MCVRL) addresse...

E. Divya, Johnson Kolluri, Kiran Siripuri · 0 citations
Open access Aug 2026

Image captioning using a transformer with topic–word semantic modeling and multimodal feature fusion

Despite the recent advances in Transformer-based image captioning models, reliance on implicit semantic representations and the lack of integration of topic-level and word-level semantic information with appearance and geometric features remain challenging. To address these limitations, we propose a semantic modeling f...

Ali Abdullah Yahya, Majjed Al-Qatf, Ammar Hawbani et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.