Skip to content

VICAL: Vicinal Consistency Alignment for Long-Tailed Visual Recognition

Sep 2026 · 0 citations · 52 references
Computer Science

TL;DR

VICAL is introduced, a consistency-driven framework that improves long-tailed recognition not by enforcing expert diversity, but by reducing prediction variance, and suggests that multi-expert models benefit more from variance reduction than diversity maximization.

Abstract

Multi-expert models have become the dominant paradigm for long-tailed learning, largely attributed to their presumed ability to benefit from expert diversity. However, we revisit this central assumption and reveal that diversity induced by logit adjustment or explicit regularizers does not guarantee better ensemble accuracy. Our work suggests that multi-expert models benefit more from variance reduction than diversity maximization. We introduce \textbf{VICAL}, a \textbf{VI}cinal \textbf{C}onsistency \textbf{AL}ignment framework that improves long-tailed recognition not by enforcing expert diversity, but by reducing prediction variance. Specifically, our approach comprises two key components: Self-Consistency Learning and Deep Ensemble Distillation. Self-Consistency Learning discourages reliance on unstable high-frequency information, smoothing the local loss landscape and mitigating overfitting, especially for tail classes. Deep Ensemble Distillation promotes cross-expert low-frequency semantic agreement using a low-resolution view, thereby sidestepping optimization conflicts with established knowledge. Extensive experiments on CIFAR-LT, ImageNet-LT, and iNaturalist 2018 show that VICAL consistently outperforms state-of-the-art methods, validating the effectiveness of our consistency-driven design. Our code is available at \href{https://github.com/FlamieZhu/Vicinal-Consistency-Alignment}{VICAL}.

View source

Similar papers

Preprint Sep 2026

0.5%>100%: Bidirectional Reciprocal Learning for Referring Image Segmentation

Recent advances in vision foundation models (VFMs) have shown remarkable capabilities across diverse unimodal visual tasks. However, adapting VFMs to referring image segmentation (RIS) typically necessitates precise vision-language alignment via full fine-tuning, incurring substantial computational overhead and risking...

Xiaoqiang Lu, Li-Cheng Jiao, Ling-Ling Li et al. · 0 citations
#machine learning Preprint Sep 2026

Unifying Distributional Training for One-Step Visual Generation

Distributional training provides collective supervision for one-step visual generation by matching real and generated features in frozen representation spaces. We introduce a unified theoretical framework that separates distribution modeling from matching discrepancy and connects global objectives to pointwise feature...

Chi Zhang, Hao-Yan Shi, Yue-Yi Liu et al. · 2 citations
Preprint Aug 2026

CLEAR: Class-wise Expert Aggregation with Structured Sampling for Long-Tailed Classification

Long-tailed classification poses a reliability challenge because models trained on imbalanced data are unevenly reliable across frequent and underrepresented classes. While existing methods address imbalance through re-balancing, adjustment, representation learning, or multi-expert modeling, they rarely estimate which...

Gawon Lim · 0 citations
#machine learning Preprint Sep 2026

StarWM: Self-Supervised Trained Attention Routing for Robust World Models

StarWM is proposed, which uses a cross-attention module trained on self-supervised dynamics to decide where reconstruction applies and preserves state attributes with near-perfect fidelity through long-horizon imagination while systematically discarding distractors.

Zeqiang Zhang, Fabian Wurzberger, Maximilian Otte et al. · 0 citations
Preprint Sep 2026

UnfoldCRF: Structured Mask Refinement with Image-Conditioned Latent Regions

Learned mask refiners improve segmentation accuracy, but it is hard to tell how much of the improvement comes from explicit structure rather than from extra capacity, and whether it holds up when the mask generator or its error distribution changes. UnfoldCRF treats refinement as inference in a conditional random field...

Chun-Ming He, Rihan Zhang, Lei Xu et al. · 0 citations
Preprint Aug 2026

DeCo-MIL: Debiased Counterfactual Reasoning for Long-Tailed Whole Slide Image Analysis

Multiple instance learning (MIL) is widely used for weakly supervised whole slide image (WSI) analysis. However, under long-tailed distributions, MIL-based WSI analysis faces a nested dual long-tail: an inter-slide class long tail and an intra-slide long tail of instance-level discriminative evidence. The two long tail...

Xiao-Xiao Li, Xi-Tong Ling, Jiawen Li et al. · 0 citations

Related blog posts

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.