Skip to content

Category

machine learning

3,367 papers

#artificial intelligence Open access May 2025

TabularQGAN: a quantum generative model for tabular data synthesis

A novel quantum generative model for synthesizing tabular data by proposing a quantum generative adversarial network architecture with flexible data encoding and a novel quantum circuit ansatz for effectively modeling tabular data is introduced.

P. Bhardwaj, Caitlin Jones, Lasse Dierich et al. · 2 citations
#artificial intelligence Preprint Apr 2025

LZ Penalty: An information-theoretic repetition penalty for autoregressive language models

The LZ penalty is introduced, a penalty specialized for reducing degenerate repetitions in autoregressive language models without loss of capability and without instances of degenerate repetition, and enables state-of-the-art open-source reasoning models to operate with greedy decoding without loss of capability and without instances of degenerate repetition.

Antonio A. Ginart, Naveen Kodali, Jason Lee et al. · 0 citations
#artificial intelligence Preprint Jan 2025

Gradient Heterogeneity Complements Hessian Heterogeneity in Transformer Optimization

This study provides a theoretical analysis showing that gradient heterogeneity, together with Hessian heterogeneity, degrades the convergence of gradient-based methods such as SGD, while sign-based methods are substantially less sensitive to this effect.

Akiyoshi Tomihari, Issei Sato · 6 citations · ⚡1

Diffusion Models for Smarter UAVs: Decision-Making and Modeling

Simulation results confirm the effectiveness and benefits of DMs in generating neighbor velocity estimates in a four-UAV swarm coordination task using Deep Reinforcement Learning (DRL), and explore the integration of DMs with RL and DT.

Yousef Emami, Hao Zhou, Luís Almeida et al. · 9 citations

HyPE-GT: where Graph Transformers meet Hyperbolic Positional Encodings

Comprehensive experiments on four molecular benchmarks, including the four large-scale Open Graph Benchmark datasets, substantiate the effectiveness of hyperbolic positional encodings in enhancing the performance of Graph Transformers and provide extensive theoretical underpinnings to offer insights into the working mechanism of the HyPE framework.

Kushal Bose, Swagatam Das · 2 citations
#machine learning Preprint Aug 2026

TokEval: A Tokenizer Evaluation Suite

Language model tokenizers are typically selected with minimal evaluation, despite the fact that their design choices directly impact model capabilities. This can be partly attributed to a limited understanding of which tokenizer properties affect which aspects of downstream performance. We introduce TokEval, a framework of tokenizer evaluation metrics that goes beyond standard measures like fertility and compression rate to capture linguistically and structurally meaningful properties, e.g., UTF-8 character boundary integrity and digit place-value boundary alignment for mathematics. To validate whether these metrics are predictive of downstream model performance, we conduct controlled language model pretraining experiments, varying solely the tokenizers'training data mixture, pretokenization strategy, and training algorithm. We evaluate the resulting models on bits-per-byte (a tokenizer-agnostic version of perplexity) and several benchmarks, spanning linguistic understanding, mathematical reasoning, and code generation. Our experiments suggest that different intrinsic properties have different impacts on model abilities: information-theoretic metrics predict language modeling abilities (Spearman rho up to 0.80), while structure-sensitive metrics, such as those measuring digit and line-break handling, correlate with task accuracy. We hope TokEval enables more principled tokenizer evaluation, replacing pretraining sweeps with intrinsic measurement wherever the two agree.

Clara Meister · 0 citations
#machine learning Preprint Aug 2026

Primitive Representation Learning for Unsupervised Dynamic Contrast Enhanced MRI Reconstruction

This work proposes a multi-dimensional, primitive based framework for dynamic contrast-enhanced MRI reconstruction that disentangles the underlying anatomy, the dynamic contrast enhancement, and residual motion into separate temporal basis functions, thereby enabling a geometrical interpretation of the representation.

Veronika Spieker, Wenqi Huang, Cemre Ariyurek et al. · 0 citations
#machine learning Preprint Open access Aug 2026

Where A Small Language Model Helps in Invoice Categorisation, Understood Through Embedding Geometry

Analysis of the pre-trained embedding geometry of a small sentence transformer (SBERT) and classic SLM reveals that pre-trained embedding geometry is associated with classification performance and reveals a counterintuitive finding that a structured input that would help a human reader does not improve the SLM performance.

Emma Ceccherini, Daniel Lawson, Anjulika Salhan · 0 citations
#artificial intelligence Open access Aug 2026

Against Political Polarization: A Unified Framework for Tracing Evolving Political Ideologies on Social Media

TSN4PI, a unified framework for tracking the evolution of political ideologies on social media, includes two core modules: the PIDN and the PIPN, which employs temporal graph neural networks to predict future ideological shifts, enabling comprehensive analysis of ideology presence, intensity, and evolution.

Yijie Xu, Chao Wang, Hui Xiong · 0 citations
#artificial intelligence Preprint Aug 2026

Towards Zero-Shot Task Transfer with Neurosymbolic World Models

This work presents a novel world model formulation where the reward prediction only depends on a subset of structured, symbolic components of the whole latent state, and demonstrates the strong generalisation properties of this approach over purely neural methods.

Isidoro Tamassia, Lennert De Smet, Giuseppe Marra · 0 citations

From tech blogs

See all →
MIT News · Artificial Intelligence Aug 27, 2026

Looking beyond natural sequences

A new machine-learning framework aims to improve the success rate of computational protein design while moving away from results that reproduce sequences found in nature.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.