Experiments show that VG-TIE is competitive with other tabular-to-image methods while providing interpretability on feature importance and ranking similar to intrinsic interpretable methods, and highlights the potential of the proposed image-based transformation to provide an effective framework that expands the use of deep learning across tabular data domains.
Abstract
Tabular-to-image encoding methods enable the application of models based on both convolutional neural networks and vision transformers to tabular data, transforming feature vectors into images. Existing methods employ linear and nonlinear dimensionality reduction techniques (e.g., Principal Component Analysis (PCA), t-SNE, and UMAP) to determine pixel positions, resulting in images whose spatial layout do not inherently reflect feature relationships. This paper introduces Visibility Graphs for Tabular-to-Image Encoding (VG-TIE), a novel method that encodes the structure of feature values using Natural Visibility Graph (NVG) and Horizontal Visibility Graph (HVG) into a two-dimensional space obtained through PCA. The resulting images are model-agnostic and intrinsically interpretable. Each pixel corresponds to an input feature, its intensity reflects the magnitude and direction of deviation from the population mean, and edges represent formally defined visibility relationships between features. VG-TIE provides two interpretability methods: (i) feature ranking from node degree distributions; and (ii) local and global feature importance from pixel intensity combined with Grad-CAM. Experiments on six public tabular datasets show that VG-TIE is competitive with other tabular-to-image methods while providing interpretability on feature importance and ranking similar to intrinsic interpretable methods. The results highlight the potential of the proposed image-based transformation to provide an effective framework that expands the use of deep learning across tabular data domains.
TabSOM is proposed, a tabular-to-image encoding built on the Self-Organizing Map, which provides a spatial layout in which every input feature occupies a fixed canvas position derived from its component plane via collision-free Hungarian assignment and a graph that captures pairwise feature relationships derived from t...
David Chushig-Muzo, M. D. de Cara, Eva Milara et al.· 1 citation
A significant difference is revealed between the visual evidence reflected in CLIP representations and that readily accessible to human perception, raising broader questions about the relationship between artificial and human vision and, ultimately, between artificial and human aesthetic judgment.
Vision Transformers (ViTs) rely on positional encoding (PE) because self-attention has no native notion of token order or image-grid location, yet the information-theoretic properties of different PE strategies and their downstream consequences for model behaviour remain insufficiently characterised. We present a syste...
D. Bandur, M. Bandur, B. Jakšić· IEEE Transactions on Pattern...· 0 citations
The Visual-Spatial Latent Graph Network (VSLG-Net), a parameter-compact transformer-based framework with dual-branch for local and global context perception in attention mechanisms, is proposed, which achieves competitive performance on the outdoor YFCC100M benchmark and remains competitive on the indoor SUN3D benchmar...
Wei Lv, Han-Lin Guo, Zhi Shen et al.· Signal, Image and Video Proc...· 0 citations
This work proposes Structure-Anchored Rectified Flow (SA-RF), which maintains correspondence through separate chromaticity/intensity stems, a scale-matched condition pyramid, and HybridAda, and introduces BC-IHV, a learnable Box--Cox polar color space whose analytically invertible intensity mapping controls the inverse...
Yihao Ai, Zheng Chen, Yuanhao Cai et al.· 0 citations
Exploring how generative AI could make machine vision more accessible to businesses. The post GenEye in a Box: Making Machine Vision Something You Can Just Ask For appeared first on GPT-Lab.
Microsoft Research Blog· microsoft.comAug 11, 2026
Radiology AI is evolving beyond report generation. CARE-X explores a unified approach that combines flexible reasoning, calibrated predictions, and measurement-based tools for chest X-ray interpretation. The post Introducing CARE-X: Towards Clinically Useful Radiology VLMs with Auxiliary Supervision, Reward-Aligned Learning, and Tool-Augmented Measurement appeared first on Microsoft Research.
Students in MIT’s Concourse program delve deeply into the human condition, debate challenging questions, and learn to develop judgment about issues that can’t be quantified.
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.