TabSOM is proposed, a tabular-to-image encoding built on the Self-Organizing Map, which provides a spatial layout in which every input feature occupies a fixed canvas position derived from its component plane via collision-free Hungarian assignment and a graph that captures pairwise feature relationships derived from the SOM component planes.
Abstract
Tabular-to-image methods have emerged as novel approaches to leverage the high predictive performance of convolutional neural networks and vision transformers. They convert tabular data into image representations, mapping each feature at a fixed pixel location derived from a dimensionality-reduction method (e.g., t-SNE, UMAP, PCA). However, they encode only the marginal value of each feature and discard information about feature relationships. We propose TabSOM, a tabular-to-image encoding built on the Self-Organizing Map (SOM), which provides: (i) a spatial layout in which every input feature occupies a fixed canvas position derived from its component plane via collision-free Hungarian assignment; and (ii) a graph that captures pairwise feature relationships derived from the SOM component planes. The resulting image stacks two multi-scale node channels: one encodes feature values at fixed scales, while the other encodes pairwise feature interactions as spatial connections between related features. Two SOM-derived interpretability approaches are introduced: a prototype-inspired partial dependence plot and a class--separation importance score. Benchmarked against twelve existing tabular-to-image methods across public binary-classification datasets, TabSOM ranks first or second on every dataset and achieves the lowest variance of any method evaluated. Interpretability obtained with TabSOM was validated against Random Forest, XGBoost, and SHAP, the class-separation score shows reasonable agreement with established baselines on the top-ranked features while capturing complementary structural information from input data. These results demonstrate that TabSOM provides an effective and interpretable approach for applying deep learning architectures to tabular data, bridging the performance--interpretability gap in this domain.
Experiments show that VG-TIE is competitive with other tabular-to-image methods while providing interpretability on feature importance and ranking similar to intrinsic interpretable methods, and highlights the potential of the proposed image-based transformation to provide an effective framework that expands the use of...
David Chushig-Muzo, L. López-Ramos, Ángeles Rodríguez de Cara et al.· 0 citations
Results indicate that TTA improves OOD performance, with composite and photometric strategies providing the best trade-off between robustness and variance, in contrast to frequency-domain transformations that alter the encoder's feature-to-intensity mapping consistently degrade performance.
Malena Loza, Felipe Grijalva, Eva Milara et al.· 0 citations
Vision Transformers (ViTs) rely on positional encoding (PE) because self-attention has no native notion of token order or image-grid location, yet the information-theoretic properties of different PE strategies and their downstream consequences for model behaviour remain insufficiently characterised. We present a syste...
D. Bandur, M. Bandur, B. Jakšić· IEEE Transactions on Pattern...· 0 citations
GraLoD, a plug-and-play framework that treats restoration scale as a spatially varying and stage-dependent continuous variable, is proposed and minimal-sufficient footprint calibration (MSFC) together with structure-aware regularization (SAR) is introduced to encourage restoration-effective and spatially coherent LOD a...
Self-supervised pre-training with Vision Transformers, including Masked Autoencoders (MAE), is difficult to apply to gigapixel scientific images. Random masking is poorly matched to the structured, multi-scale morphology of scientific data, while uniform tokenization produces prohibitively long sequences that make $O(N...
En-Zhi Zhang, Du Wu, Rui Zhong et al.· 0 citations
Related blog posts
MIT News · Artificial Intelligence· news.mit.eduJul 15, 2026
Assistant Professor Pat Pataranutaporn describes a new interface that lets everyday users glimpse inside an AI's neural network before their chatbot ever says a word.
Microsoft Research Blog· microsoft.comJul 13, 2026
Cryptographic code supports vital protections in modern computing systems. Learn how a new method helps verify code as developers write it while preserving speed and adaptability as it gets implemented and evolves. The post Verifying Rust cryptography in SymCrypt, from standards to code appeared first on Microsoft Research.
MIT News · Artificial Intelligence· news.mit.eduJul 6, 2026
PhD student Rachel Sava, winner of the Envisioning the Future of Computing Prize, explores transformative improvements and dystopian risks of neural technology.