Aug 2026· Machine Learning and Knowledge Extraction· 0 citations· 39 references
TL;DR
It is shown that the proposed TAME method consistently surpasses baselines on neural classifiers, while remaining competitive with strong coreset baselines on tree-based classifiers (RF, XGBoost) with a lightweight validation method.
Abstract
Dataset distillation has achieved strong results in computer vision, but is largely underexplored in the tabular domain. We introduce a tabular dataset distillation method that projects the data through many random embedders (views) to achieve invariance to transformation and to focus on the consistency between real and synthetic sets. The synthetic set is determined through a formulation of distribution matching between the many-view projection of the original and distilled dataset. Building on this approach, our proposal achieves three goals: (1) we formulate the Tabular Alignment via Moment Embeddings (TAME) method and, by extensive empirical evaluation, we show its efficiency; (2) we evaluate TAME on a benchmark of 18 tabular datasets, with strong baselines and evaluation metrics; and (3) we present a structured set of studies analyzing the impact of losses, dataset geometry, embedder architecture, instances per class (IPC) budget and downstream classifiers. We show that the proposed TAME method consistently surpasses baselines on neural classifiers, while remaining competitive with strong coreset baselines on tree-based classifiers (RF, XGBoost). Performance is further increased, especially for tree-based classifiers with a lightweight validation method. Extensive evaluation, including on additional large-scale sets and ablation experiments, allow a better understanding of the method.
Results indicate that TTA improves OOD performance, with composite and photometric strategies providing the best trade-off between robustness and variance, in contrast to frequency-domain transformations that alter the encoder's feature-to-intensity mapping consistently degrade performance.
Malena Loza, Felipe Grijalva, Eva Milara et al.· 0 citations
Self-supervised representation-guided generative dataset distillation (SRG) is proposed, a framework that translates the SSL geometry into diffusion guidance and consistently outperforms the evaluated generative baselines across multiple datasets and IPC settings.
Mingzhuo Li, Guang Li, Linfeng Ye et al.· 0 citations
Multi-step matching models, including flow and diffusion models, produce high-quality outputs but incur substantial inference costs and may reproduce unwanted components of their training datasets. We introduce Inverse Distillation Unlearning (IDU), a unified framework that simultaneously distills a teacher multi-step...
Aleksei Leonov, N. Kornilov, Zhen-He Zhang et al.· 0 citations
DeCO uses attention rollout from a pretrained TransFG teacher to identify informative patches, applies spatial diversification to reduce redundant coverage, and organizes the resulting regions into class-wise evidence banks and consistently outperforms representative coreset and dataset-distillation baselines under dif...
Chuixuan Fan, Guang Li, Shi-Jie Wang et al.· 0 citations
The Segment Anything Model (SAM) has gained widespread recognition in general visual segmentation owing to its strong zero-shot generalisation and segmentation performance, yet its large network architecture severely limits deployment on mobile and edge devices. Most existing lightweight SAM methods adopt decoupled dis...
Multimodal synthetic datasets combine structured attributes with free text, but are often evaluated separately. Such metrics can remain high after tabular--text pairings are disrupted. We present a projection-based evaluator for tabular--text synthetic data. A fixed sentence encoder maps text to embeddings, \(k\)-means...
Ye-Feng Yuan, Zhan Shi, Liang Cheng et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.