Sep 2026· Journal of Computing and Data Technology· 0 citations
TL;DR
This narrative review examines deep learning for static visual scene classification across indoor, outdoor, and aerial imagery and concludes that useful progress should be assessed through reproducible gains under matched protocols and credible transfer to new environments, rather than isolated accuracy values.
Abstract
Scene classification assigns semantic categories to images by integrating information about objects, spatial layout, texture, and environmental context. Deep learning has shifted this task from manually designed descriptors toward transferable representations, but reported improvements remain difficult to interpret when datasets, supervision, and evaluation protocols differ. This narrative review examines deep learning for static visual scene classification across indoor, outdoor, and aerial imagery. It organizes the literature along three complementary axes: learning signal, representation architecture, and prediction task. The discussion connects convolutional feature extraction, local and global feature aggregation, semantic and graph-based reasoning, and generative representation learning with selected developments in vision transformers, self-supervised pretraining, and vision-language models. Benchmark characteristics are examined alongside the distinction between single-label, multi-label, and zero-shot evaluation. A source-checked numerical example illustrates why domain-specific pretraining can improve some benchmarks while leaving others unchanged or worse. The synthesis identifies dataset bias, incomplete split reporting, spatial leakage, semantic ambiguity, and deployment cost as persistent obstacles to reliable comparison. It also establishes a reporting framework covering class-wise performance, uncertainty, computational resources, and distribution shift. The review concludes that useful progress should be assessed through reproducible gains under matched protocols and credible transfer to new environments, rather than isolated accuracy values.
: Fine-grained visual categorization (FGVC) presents a class of recognition problems in which the discriminative signal is spatially concentrated, visually subtle, and easily destroyed by the preprocessing and augmentation strategies that serve coarse recognition well. Where standard image classification requires a mod...
Richard Adusei, G. Abdul-Salaam· Journal of Artificial Intell...· 0 citations
Deep learning, especially CNN- and transformer-based approaches, has significantly impacted computer vision techniques such as object detection, segmentation, and scene classification. Despite the many successful scene understanding pipelines currently in use, most still rely primarily on global image features or objec...
M. Srividya, V. Rachapudi· Discover Computing· 0 citations
Few-shot image classification remains difficult because a model must identify novel classes from only one or a few labeled examples while preserving discriminative local information. Metric-learning methods based on Earth Mover’s Distance (EMD) improve local correspondence by representing an image as a set of regional...
Huie Zhang, Mary Jane C. Samontet· International journal of com...· 0 citations
Existing class-incremental learning (CIL) methods for remote sensing (RS) scene classification often tend to be training-intensive or rely on static visual features that may inadequately capture the complex interclass similarity and intraclass diversity inherent in RS imagery. Moreover, directly reusing features from m...
Wen-Liang Du, Ji-Cun He, Jia-Qi Zhao et al.· IEEE Transactions on Geoscie...· 0 citations
How much of an image does a detector need to see? Small RGB regions can retain useful evidence of image synthesis even when they reveal little of the full scene. Motivated by single-patch detection, we study patch-wise supervision: a shared backbone classifies explicit crops, each crop receives its own loss, and patch...
Zhi-Da Zhang, Tao Wu, Si-Yu Liu et al.· 0 citations
Learning rate (LR) initialization and decay remain important factors in the optimization of deep vision networks. Although these models exhibit a clear hierarchical structure, training typically starts from a single global learning rate, with little explicit consideration of stage depth. This paper investigates a simpl...
Qiang He, Qiu Zong, Yi-Qi Wang et al.· IEEE Access· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.