Skip to content

Category

machine learning

1,973 papers

#artificial intelligence Preprint Jan 2025

Gradient Heterogeneity Complements Hessian Heterogeneity in Transformer Optimization

This study provides a theoretical analysis showing that gradient heterogeneity, together with Hessian heterogeneity, degrades the convergence of gradient-based methods such as SGD, while sign-based methods are substantially less sensitive to this effect.

Akiyoshi Tomihari, Issei Sato · 6 citations · ⚡1

Diffusion Models for Smarter UAVs: Decision-Making and Modeling

Simulation results confirm the effectiveness and benefits of DMs in generating neighbor velocity estimates in a four-UAV swarm coordination task using Deep Reinforcement Learning (DRL), and explore the integration of DMs with RL and DT.

Yousef Emami, Hao Zhou, Luís Almeida et al. · 9 citations

HyPE-GT: where Graph Transformers meet Hyperbolic Positional Encodings

Comprehensive experiments on four molecular benchmarks, including the four large-scale Open Graph Benchmark datasets, substantiate the effectiveness of hyperbolic positional encodings in enhancing the performance of Graph Transformers and provide extensive theoretical underpinnings to offer insights into the working mechanism of the HyPE framework.

Kushal Bose, Swagatam Das · 2 citations
#machine learning Preprint Aug 2026

TokEval: A Tokenizer Evaluation Suite

Language model tokenizers are typically selected with minimal evaluation, despite the fact that their design choices directly impact model capabilities. This can be partly attributed to a limited understanding of which tokenizer properties affect which aspects of downstream performance. We introduce TokEval, a framework of tokenizer evaluation metrics that goes beyond standard measures like fertility and compression rate to capture linguistically and structurally meaningful properties, e.g., UTF-8 character boundary integrity and digit place-value boundary alignment for mathematics. To validate whether these metrics are predictive of downstream model performance, we conduct controlled language model pretraining experiments, varying solely the tokenizers'training data mixture, pretokenization strategy, and training algorithm. We evaluate the resulting models on bits-per-byte (a tokenizer-agnostic version of perplexity) and several benchmarks, spanning linguistic understanding, mathematical reasoning, and code generation. Our experiments suggest that different intrinsic properties have different impacts on model abilities: information-theoretic metrics predict language modeling abilities (Spearman rho up to 0.80), while structure-sensitive metrics, such as those measuring digit and line-break handling, correlate with task accuracy. We hope TokEval enables more principled tokenizer evaluation, replacing pretraining sweeps with intrinsic measurement wherever the two agree.

Clara Meister · 0 citations
#machine learning Preprint Aug 2026

Primitive Representation Learning for Unsupervised Dynamic Contrast Enhanced MRI Reconstruction

This work proposes a multi-dimensional, primitive based framework for dynamic contrast-enhanced MRI reconstruction that disentangles the underlying anatomy, the dynamic contrast enhancement, and residual motion into separate temporal basis functions, thereby enabling a geometrical interpretation of the representation.

Veronika Spieker, Wenqi Huang, Cemre Ariyurek et al. · 0 citations
#machine learning Preprint Open access Aug 2026

Where A Small Language Model Helps in Invoice Categorisation, Understood Through Embedding Geometry

Analysis of the pre-trained embedding geometry of a small sentence transformer (SBERT) and classic SLM reveals that pre-trained embedding geometry is associated with classification performance and reveals a counterintuitive finding that a structured input that would help a human reader does not improve the SLM performance.

Emma Ceccherini, Daniel Lawson, Anjulika Salhan · 0 citations
#artificial intelligence Open access Aug 2026

Against Political Polarization: A Unified Framework for Tracing Evolving Political Ideologies on Social Media

TSN4PI, a unified framework for tracking the evolution of political ideologies on social media, includes two core modules: the PIDN and the PIPN, which employs temporal graph neural networks to predict future ideological shifts, enabling comprehensive analysis of ideology presence, intensity, and evolution.

Yijie Xu, Chao Wang, Hui Xiong · 0 citations
#artificial intelligence Preprint Aug 2026

Towards Zero-Shot Task Transfer with Neurosymbolic World Models

State-of-the-art model-based reinforcement learning methods learn neural world models that allow policy improvement by planning in a latent space, without assumptions on the structure of the underlying environment. While expressive, these models are generally task-dependent: they learn uninterpretable latent representations that are tied to the training task and thus hard to generalize to new tasks. In this work, we present a novel world model formulation where the reward prediction only depends on a subset of structured, symbolic components of the whole latent state. Decoupling observation reconstruction and reward prediction allows us to learn world models that can adapt zero-shot, i.e. without further environment interactions, to new reward functions defined over the same symbolic state space. We discuss the main advantages and challenges of learning these neurosymbolic world models and demonstrate the strong generalisation properties of our approach over purely neural methods.

Isidoro Tamassia, Lennert De Smet, Giuseppe Marra · 0 citations
#artificial intelligence Preprint Aug 2026

Procedural Content Metageneration via Program Search and Continual Abstraction Discovery

Large language models can generate executable programs, which makes it possible to search directly over procedural content generators rather than individual levels. We study this approach in Sokoban, Zelda, Dangerous Dave, and Lode Runner. Each run evolves complete Python generators through language-model mutation and crossover. We introduce Continual Abstraction Discovery, or CAD, which extracts reusable primitives from high-fitness programs into a run-specific helper module. A 2x2 experiment crosses CAD with access to a fixed hand-written domain API. The completed data set contains 160 complete runs, with at least ten 50-generation runs in every cell. CAD raises mean final best fitness in all eight domain and API comparisons. Across all CAD runs, learned libraries are adopted by most later programs and repeatedly rediscover validation, reachability, and structural utilities. These results support that discovering reusable primitives improves evolutionary program search for content generators.

Matthew Siper, A. Khalifa, Julian Togelius · 0 citations
#machine learning Preprint Aug 2026

AppendiGrade: An XAI-Enhanced Deep Learning Framework for Grading Appendicitis in Ultrasound with Gaussian Blur and Grad-CAM

An advanced system capable of automatically detecting complicated appendicitis from ultrasound images was developed and was explained with gradient-weighted class activation mapping (Grad-CAM), which creates a heatmap of the regions responsible for the model's prediction of the infected areas.

Fahad Ahammed, Omar Faruq Shikdar, Navid Zaman et al. · 0 citations

From tech blogs

See all →
MIT News · Artificial Intelligence Aug 27, 2026

Looking beyond natural sequences

A new machine-learning framework aims to improve the success rate of computational protein design while moving away from results that reproduce sequences found in nature.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.