Skip to content

Category

machine learning

3,367 papers

The Discrete-Log Clock: How a Transformer Learns Modular Multiplication

The transformer reduces multiplication to addition in discrete-log space, implementing a "Discrete-Log Clock" algorithm analogous to Nanda et al.'s Clock algorithm for addition, which generalizes: matching the analysis basis to the algebraic structure of the task reveals interpretable structure where standard tools see noise.

Huu-Tuan Nguyen · 0 citations

TokenPilot: Cache-Efficient Context Management for LLM Agents

TokenPilot is presented, a dual-granularity context management framework that reduces costs by 61% and 56% in isolated mode, and 61% and 87% in continuous mode, while maintaining competitive performance compared to prior systems.

Buqiang Xu, Z. Xue, Dian Chen et al. · 1 citation

LongDS-Bench: On the Failure of Long-Horizon Agentic Data Analysis

LongDS is introduced, a benchmark for long-horizon, multi-turn data analysis where agents must maintain, update, restore, and compose evolving analytical states, suggesting that the key bottleneck is maintaining a correct analytical state rather than increasing interaction budget.

Kewei Xu, Xiaobe Lu, Shuofei Qiao et al. · 1 citation

Negligible in Size, Significant in Effect: On Scale Vectors in Large Language Models

This work presents a systematic study of scale vectors in LLMs from the perspectives of expressivity, optimization, and architectural structure, and proposes three lightweight and complementary improvements to scale vectors: branch-specific heterogeneity, improved placement around linear mappings, and magnitude-direction reparameterization.

Mingze Wang, Shuchen Zhu, Yuxin Fang et al. · 3 citations

More Expressive Feedforward Layers: Part I. Token-Adaptive Mixing of Activations

This work proposes Mixture of Activations (MoA), a token-adaptive FFN design that mixes a dictionary of activation functions using lightweight input-dependent gates while sharing the same linear projections, suggesting that token-adaptive activation mixing is a simple and effective mechanism for improving FFN expressivity in LLMs.

Mingze Wang, Jinbo Wang, Yikuan Xia et al. · 3 citations

SkillSafetyBench: Evaluating Agent Safety under Skill-Facing Attack Surfaces

This work presents SkillSafetyBench, a runnable benchmark for evaluating skill-facing safety failures, and suggests that agent safety depends not only on model-level alignment, but also on how agents interpret skills, trust workflow context, and act through executable environments.

Chang Jin, Anr'an W'ang, Zeming Wei et al. · 11 citations · ⚡1

ABC: Any-Subset Autoregression via Non-Markovian Diffusion Bridges in Continuous Time and Space

This work proposes ABC: Any-Subset Autoregressive Models via Non-Markovian Diffusion Bridges in Continuous Time and Space, and derives SDE dynamics via changes-of-measure on path space, yielding another advantage: path-dependent conditioning on arbitrary subsets of the state history and/or future.

Gabriel Guo, Thanawat Sornwanee, L. Hao et al. · 0 citations

G-Loss: Graph-Guided Fine-Tuning of Language Models

G-Loss is presented, a graph-guided loss function that incorporates semi-supervised label propagation to use structural relationships within the embedding manifold to build a document-similarity graph that captures global semantic relationships.

Aditya Sharma, Vinti Agarwal, Rajesh Kumar · 0 citations

Beyond Output Correctness: Benchmarking and Evaluating Large Language Model Reasoning in Coding Tasks

This work introduces CodeRQ-Bench, the first benchmark for evaluating LLM reasoning quality across three coding task categories: generation, summarization, and classification, and proposes VERA, a two-stage evaluator that combines evidence-grounded verification with ambiguity-aware score correction.

Yuangang Li, Justin Tian Jin Chen, Ethan Yu et al. · 1 citation

PolicyLong: Towards On-Policy Context Extension

PolicyLong is proposed, shifting data construction towards a dynamic on-policy paradigm, by iteratively re-executing data screening (entropy computation, retrieval, and verification) using the current model, which ensures the training distribution tracks evolving capabilities, yielding an emergent self-curriculum.

Junlong Jia, Jiangnan Zhou, Ziyang Chen et al. · 0 citations

Camera-Agnostic Pruning of 3D Gaussian Splats via Descriptor-Based Beta Evidence

This paper proposes a camera-agnostic, one-shot, post-training pruning method for 3D Gaussian splats that relies solely on attribute-derived neighbourhood descriptors, and introduces a hybrid descriptor framework that captures structural and appearance consistency directly from the splat representation.

Peter O. Fasogbon, Ugurcan Budak, P. R. Alface et al. · 0 citations

Var-JEPA: A Variational Formulation of the Joint-Embedding Predictive Architecture - Bridging Predictive and Generative Self-Supervised Learning

The Variational JEPA (Var-JEPA), which makes the latent generative structure explicit by optimizing a single Evidence Lower Bound (ELBO) and yields meaningful representations without ad-hoc anti-collapse regularizers and allows principled uncertainty quantification in the latent space.

Moritz Gögl, Christopher Yau · 3 citations · ⚡1

From tech blogs

See all →
MIT News · Artificial Intelligence Aug 27, 2026

Looking beyond natural sequences

A new machine-learning framework aims to improve the success rate of computational protein design while moving away from results that reproduce sequences found in nature.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.