Skip to content
Conference

Improving Visual Question Answering Via Rule- Guided Neuro-Symbolic Learning

Aug 2026 · International Conference on Information Security and Cryptology · pp. 1160-1167 · 0 citations · 20 references

Abstract

Visual Question Answering (VQA) involves models combining reasoning through visual scenes and natural language questions and typically related to compositional and relational reasoning. Even though deep neural models have demonstrated high empirical results on VQA benchmarks, they are often based on implicit associations, and their reasoning is not always transparent. In order to overcome this shortcoming, this paper will discuss a lightweight neuro-symbolic method that incorporates symbolic consistency constraints into neural visual reasoning. The proposed framework suggests a rule-based learning model, where the scene annotations in the form of symbols are applied to penalize logically inconsistent predictions towards training. This method couples a regular neural processing of visual and linguistic information with a symbolic consistency loss based on facts of the scene. The study shows the test of the suggested approach on the CLEVR data, a diagnostic task that aims at testing compositional visual reasoning. Through experimental findings, it is shown that the suggested neuro-symbolic model shows consistent gains in accuracy of answer prediction with a near-zero rule violation rate with each training epoch. Qualitative analysis also suggests that symbolic guidance brings about less reasoning inconsistencies and increases interpretability without adding complex symbolic engines. These results have indicated that slight symbolic integration is capable of positively contributing to the faithfulness and reliability of neural visual reasoning systems.

View source

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.