Skip to content

Category

machine learning

2,173 papers

ToolSense: A Diagnostic Framework for Auditing Parametric Tool Knowledge in LLMs

This work introduces ToolSense, an open-source LLM-powered diagnostic framework that takes any tool catalog as input and automatically generates three benchmarks: a Realistic Retrieval Benchmark (RRB) with queries at three ambiguity tiers, an MCQ probing benchmark, and a QA probing benchmark.

Ashutosh Hathidara, Sai Shruthi Sistla, Sebastian Schreiber et al. · 2 citations
#artificial intelligence Review Apr 2026

D3-Gym: Constructing Real-World Verifiable Environments for Data-Driven Discovery

D3-Gym is introduced, the first automatically constructed dataset with verifiable environments for scientific Data-Driven Discovery, and how D3-Gym environments can serve as a testbed for studying agentic optimization loops such as Autoresearch on real scientific workflows.

Hanane Nour Moussa, Yi-Fei Li, Zhuo-Yang Li et al. · 0 citations
#artificial intelligence Review Mar 2026

Prompts Without Evidence: How Neuroimaging Mentions Shift Clinical Vision-Language Model Predictions

In this single-model intervention, preference alignment suppresses MRI-referencing behavior but reduces the augmented-condition advantage, leaving the underlying issue unresolved, and results caution against reading surface metric gains as evidence of true multimodal integration.

Doan Nam Long Vu, Simone Balloccu · 0 citations
#artificial intelligence Review Aug 2026

Learning a Size-Weight Frontier for Synthetic-Augmented Inference

A general framework for synthetic-augmented inference across a population of related tasks is developed, which characterizes synthetic augmentation by the number of synthetic observations and their weight and specifies a size-weight frontier that specifies, for each weight, the largest synthetic sample size for which all smaller sizes attain the target task-marginal coverage.

Chengpiao Huang, Kaizheng Wang · 0 citations
#artificial intelligence Review Aug 2026

Blog: Survey of Optimizers

This survey organizes recent optimizers and training optimization methods along four largely independent axes: temporal estimation, update geometry, horizon management, and representation and systems.

Ruo-Ran Xu · 0 citations
#artificial intelligence Preprint Aug 2026

An Enclosed Mode Is a Gauge Choice: Topology Relative to Reach in Certified Code World Models

A code world model accepted by a sampling gate can be exactly right on everything the gate can see and arbitrarily wrong beyond it. We characterize what a certified model can know, and what its errors can cost, when the omission is an annular freeze mode enclosing an unreachable interior. The gate quotient makes the question precise: acceptance-with-certainty determines the model exactly on the reachable query set; beyond reach is gauge. On a minimal ring instrument we prove the extreme case (a wrong-topology filled-disc artifact unfalsifiable by any sampling gate and bitwise harmless at play) and measure, with LLM synthesis across three model families, how one knob (a channel of width gamma) walks the same artifact through three regimes: unfalsifiable-and-harmless, falsifiable-and-costly, and instantly falsified. Three principles organize the empirics. First, danger is topology relative to reach: a channel the planner can use collapses the blind model's exploitation (play cost 1.09 to ~0 over a knee at gamma ~ 0.1), while a hidden channel with the same first Betti number keeps it at full strength (1.12). Second, repair is parameter-bound and sensor-bound: no family recovers the region from outside evidence; from inside, models pose the right topology but cannot pin its parameters, and the posed topology tracks the guiding persistent-homology summary's wrong beta_1 (a sensor with a measured geometric resolution limit), not the truth. Third, mitigation must match the error's dimension and direction: point fences fail against the one-dimensional boundary, a dimension-matched persisted fence collapses exploitation to a two-lesson transient (0.999 to 0.058), and the dual freedom certificate collapses the invented-mode failure symmetrically (1.769 to 0.029). In n dimensions the shell makes misidentification near-certain while the danger stays fully exploitable: the two axes are independent.

Javier Aguilar Martín · 0 citations
#artificial intelligence Preprint Aug 2026

How Proper Scoring Rules Shape LLM Forecasting

This paper evaluates how reward function choice shapes the performance and behavior of LLM forecasters. We compare five proper scoring rules as training objectives for binary forecasts of resolved real-world events. Although the rules share the same theoretical incentive for truthful probability reporting, the resulting models differ in calibration, probability use, and estimated profiles of bias, information, and noise, with smaller differences in aggregate accuracy and discrimination. The Brier-trained model has the lowest observed Brier score and highest AUC-ROC, while the log-trained model has the highest observed log score and lowest calibration error. Models with similar aggregate performance also reach that performance through different combinations of bias, information, and noise. Proper scoring rules therefore need not behave interchangeably as training objectives. Reward choice may shape not only how well an LLM forecasts, but how its forecasting errors are structured. Each condition uses a single seed, so some differences may reflect training stochasticity.

Benjamin Turtel, Paul Wilczewski, Kris Skotheim et al. · 0 citations
#artificial intelligence Preprint Aug 2026

Deriving Scaling Laws for OpenEuroLLM Models: Learning Rate, Batch Size and Loss

This study establishes a first baseline and scaling procedure for the development of future OpenEuroLLM models, and characterize the dependence of loss on model capacity and dataset size, evaluating recently proposed scaling forms that explicitly model their interaction.

Niccolò Ajroldi, Diana-Alexandra Onutu, Haider Al-Tahan et al. · 0 citations
#artificial intelligence Preprint Aug 2026

VISTA: Verifier-Informed Student-to-Teacher Adaptation for On-Policy Self-Distillation

Verifier-Informed Student-to-Teacher Adaptation (VISTA), which preserves the standard OPSD student update while using outcome-verified rollouts to adapt the teacher toward the student distribution and demonstrates the value of student supervision from outcome-verified rollouts.

Ze-Wen Ding, Ze-Zhong Wu, Zhou Tao et al. · 0 citations
#artificial intelligence Preprint Aug 2026

Performative Privacy: When Differential Privacy Maximizes Utility

It is shown, through a theoretical study of the dynamics and numerical experiments, that a finite privacy budget can outperform non-private estimation in the long term when the feedback loop between leakage and participation is sufficiently strong.

Uddalak Mukherjee, Edwige Cyffers, Y. Chevaleyre · 0 citations

From tech blogs

See all →
MIT News · Artificial Intelligence Aug 27, 2026

Looking beyond natural sequences

A new machine-learning framework aims to improve the success rate of computational protein design while moving away from results that reproduce sequences found in nature.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.