Experimental results demonstrate that GERIS outperforms existing instance selection methods in terms of classification accuracy and robustness, and offers a transparent, theoretically grounded mechanism for data filtering.
Abstract
In this paper, we propose GERIS, a game-theoretic framework for instance selection in the data augmentation phase of license plate recognition systems. During augmentation, synthetic license plate images are generated and transformed using stochastic noise to simulate real-world conditions. However, certain noise configurations lead to highly distorted, unreadable images that degrade model performance by introducing instance-dependent label noise. GERIS formulates a non-cooperative game in which each noise vector competes for inclusion in the training set based on its similarity to labeled data and its contribution to model reliability. By identifying and pruning low-quality instances, GERIS improves the overall quality of the augmented dataset. Unlike traditional black-box learning methods, GERIS offers a transparent, theoretically grounded mechanism for data filtering. Experimental results demonstrate that GERIS outperforms existing instance selection methods in terms of classification accuracy and robustness.
A plug-in, SmoothOperator (SmoothOP), which sets the smoothing coefficient of each sample from its prominence, an embedding-space signal measuring how clearly the sample's own class stands out against its strongest competing class, integrates into four existing spherical representation learning methods at minimal train...
PAPT++ is introduced, a risk-aware adversarial generation-training framework for SDG that progressively exposes the classifier to challenging yet semantically consistent variations.
Zhi-Peng Xu, De Cheng, Xinyang Jiang et al.· 0 citations
Diffusion models have achieved remarkable progress in visual generation, yet their practical deployment for multi-class procedural scene synthesis still lacks standardized training schemes and comprehensive performance verification. This study adopts DDPM as the base architecture and introduces classifier-free guidance...
Jing-Yi Fan· International Conference on...· 0 citations
Aesthetic image cropping aims to identify the optimal crop of an image in terms of aesthetics and composition. While supervision based on annotated data is fundamental, the field has been hindered by a long-standing problem: existing datasets suffer from (1) human subjectivity and (2) rigid discreteness confined to fix...
Zi-Qing Zhang, Xiao Liu, Kai Liu et al.· 0 citations
Multimodal synthetic datasets combine structured attributes with free text, but are often evaluated separately. Such metrics can remain high after tabular--text pairings are disrupted. We present a projection-based evaluator for tabular--text synthetic data. A fixed sentence encoder maps text to embeddings, \(k\)-means...
Ye-Feng Yuan, Zhan Shi, Liang Cheng et al.· 0 citations
Large language models (LLMs) can generate synthetic training data for text classification, but the quality of generated samples is heterogeneous: some fall in correct class regions of the embedding space while others land in peripheral or cross-class zones. We propose a geometric filtering framework that evaluates each...
Exploring how generative AI could make machine vision more accessible to businesses. The post GenEye in a Box: Making Machine Vision Something You Can Just Ask For appeared first on GPT-Lab.
MIT News · Artificial Intelligence· news.mit.eduSep 29, 2026
Professor Sherry Turkle’s new book, “Artificial Intimacy,” offers a withering critique of chatbots and the antisocial dynamics she believes they encourage.