Experimental results reveal that leveraging LLMs for graph refinement can improve classification accuracy, particularly for kNN graphs and some backbones, particularly for kNN graphs and some backbones.
Abstract
While the growing availability of image data has driven significant advances, labeling datasets remains costly and time-consuming. Therefore, semi-supervised approaches such as Graph Convolutional Networks (GCNs), which learn from both labeled and unlabeled data, have emerged as a promising solution. One of the primary challenges in applying GCNs to image classification is graph construction, since, unlike in citation networks or similar domains, images typically do not come with a predefined structural representation. For visual data, most studies construct graphs based on the similarity between feature vectors from pretrained deep learning backbones, typically by employing kNN or reciprocal kNN algorithms. Although Large Language Models (LLMs) have shown remarkable capability in capturing high-level semantics, their integration with GCNs for image classification remains underexplored. Aiming to fill this gap, our approach uses a Vision Language Model (VLM) to generate textual image descriptions, which are then processed by an LLM to estimate semantic similarity scores between connected images. These scores guide the pruning of edges in kNN and reciprocal kNN graphs, filtering out semantically irrelevant neighbors. Experimental results reveal that leveraging LLMs for graph refinement can improve classification accuracy, particularly for kNN graphs and some backbones. The source code is publicly available at http://gcnllm.lucasvalem.com.
Few-shot open-set recognition (FSOR) presents unique challenges due to the limited labeled samples and the presence of unknown classes during inference. While recent few-shot learning methods leverage class-level textual information to enhance feature discrimination, they often overlook image-grounded descriptions that...
Few-shot image classification remains difficult because a model must identify novel classes from only one or a few labeled examples while preserving discriminative local information. Metric-learning methods based on Earth Mover’s Distance (EMD) improve local correspondence by representing an image as a set of regional...
Huie Zhang, Mary Jane C. Samontet· International journal of com...· 0 citations
A novel broad graph convolutional network (BGCN) paradigm is proposed, which completely eliminates hidden layers and instead expands the receptive field through network width, effectively circumventing the inherent limitations of over-smoothing and overfitting in existing deep and high-order GCN models.
Alex Hay-Man Ng, Xun Liu, Fang-Yuan Lei et al.· Complex & Intelligent Sy...· 0 citations
Building upon recent Deep Neural Network architectures, current approaches lying in the intersection of Computer Vision and Natural Language Processing have achieved unprecedented breakthroughs in tasks like automatic captioning or image retrieval. Most of these learning methods, though, rely on large training sets of...
U.gayathri, M. Balaram· Journal of Science & Technol...· 0 citations
GTAP (Graph Topology-Aware Pre-training), a self-supervised initialization framework for within-dataset graph classification, improves over a matched GCN trained from scratch and achieves competitive accuracy against published baselines on most datasets with available results.
This narrative review examines deep learning for static visual scene classification across indoor, outdoor, and aerial imagery and concludes that useful progress should be assessed through reproducible gains under matched protocols and credible transfer to new environments, rather than isolated accuracy values.
Sandeep Kumar Aazad, Kritika Chaudhary, Taniya Saini et al.· Journal of Computing and Dat...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.