It is demonstrated that structured distortion of training data can induce compositional generalization in a standard CNN classifier without architectural modification, and suggests that structured complication of training signals can influence both the internal organization of learned representations and their compositional interpretation - a principle that may underlie the role of natural language in cognitive development.
Abstract
We demonstrate that structured distortion of training data - which we term complexity induction - can induce compositional generalization in a standard CNN classifier without architectural modification. Using synthetic images of colored geometric shapes, we encode classes as flat string labels (e.g.,"red-circle") with no explicit attribute decomposition, and exclude certain color-shape combinations from training entirely. We apply two distortion methods derived from Jaccard string similarity between class names: mixed labels (soft target distributions encoding inter-class overlap) and expanded dataset (false training samples with structurally motivated incorrect labels). Both methods induce the ability to predict unseen class combinations, and act at different levels: mixed labels activate the classifier for unseen combinations by exploiting the CNN's natural embedding structure, while expanded training improves the embedding factorization itself. A control with random (unstructured) false labels confirms that the effect depends on the structure of the distortion, not on noise per se. These results suggest that structured complication of training signals can influence both the internal organization of learned representations and their compositional interpretation - a principle that may underlie the role of natural language in cognitive development.
Flow-Based Distribution Matching is introduced, a non-adversarial framework that learns this reference-directed geometry through spherical conditional velocity regression that achieves performance nearly on par with DM and remains competitive with existing SSL methods.
Yu-Ling Jiao, Wen-Sen Ma, Hou-Duo Qi et al.· 0 citations
This work studies raw pixels, SD-VAE latents and DINOv2 as well as MAE representation-autoencoder features within a unified masked autoregressive rectified-flow model and shows that compression, reconstruction fidelity, token dimensionality, and visible semantic clustering do not individually predict generative behavio...
Marcel Plocher, B. Schölkopf, Andreas Geiger et al.· 0 citations
Visual classifiers are expected to generalize under data shifts, target shifts, and their combinations, yet most existing methods focus on domain invariance while failing to address intra-image predictive sufficiency. We investigate the structural hypothesis that each image contains a sample-adaptive oracle intra-image...
Aether is introduced, a simple plug-in method that applies diffusion-style random perturbations in the embedding space via controlled alpha-mixing, specifically designed to provide isotropic regularization that remains semantically consistent.
Hyesong Choi, Daeun Kim, Song Park et al.· 0 citations
Class-wise Covariance Regularization is proposed, which aligns the predicted covariance structure of class confidences with the semantic correlations encoded in pretrained text embed-dings with the geometric consistency of the class space throughout fine-tuning, resulting in more stable and interpretable confidence dis...
Ao Zhou, Zhi-Wei Jiang, Zi-Feng Cheng et al.· 0 citations
A lightweight CNN that attributes more models at higher accuracy than prior work, reaching 98.9% on 25-class DRAGON and 95.0% on 27-class OpenFake; runs in a few milliseconds at a cost nearly independent of candidate-set size; and remains robust to transformations encountered in the wild.
Asaf Livne, Amir Jevnisek, S. Avidan· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.