Skip to content

Category

machine learning

3,367 papers

Recirculation

This work describes an inference-time architectural enhancement for off-the-shelf foundation models that markedly reduces perplexity and boosts accuracy across generation and reasoning tasks, and proposes and evaluates an adaptive variant of recirculation which requires only light tuning of hyperparameters while freezing the original model weights.

Michael C. Mozer, Shoaib Ahmed Siddiqui, Danny Sawyer et al. · 0 citations
#artificial intelligence Preprint Aug 2026

Too Sure to Be Safe: Model Calibration for Reliable Log Anomaly Detection

Log Reconstruction and Distance (LoRD), a lightweight post-hoc calibration framework for reliable log anomaly detection, is proposed and demonstrates that LoRD consistently improves confidence reliability and substantially reduces overconfident anomaly-related errors without sacrificing anomaly detection performance.

Bin Li, Dongdong Wang, Siyang Lu · 0 citations
#machine learning Preprint Aug 2026

Understanding the Surprising Generalization Properties of Tabular Foundation Models

A task-centric, retrieval-based perspective is offered for how TFMs generalize: it is believed that tabular in-context generalization is largely retrieval-based, and good models are those that learn to identify relevant examples in the provided context and aggregate them well.

Nour Shaheen, Junwei Ma, Alex Labach et al. · 1 citation
#artificial intelligence Preprint Aug 2026

SIGMA: SHAP-Guided Implicit-Trajectory Generation for Metadata-Free LLM-Based AutoFE

This work proposes a SHAP-enhanced Implicit-trajectory Generation for Metadata-free AutoFE (SIGMA), a scalable constant-context optimization framework that leverages SHAP values to provide task-aware signals for guiding group feature generation instead of semantic information.

Xu Zheng, Kento Uchida, Shinichi Shirakawa · 0 citations
#machine learning Preprint Aug 2026

Dynamic Compression in Recurrent Networks

Recurrent models process long contexts efficiently by compressing their history into a fixed-size state, but modern architectures typically do so in a single causal pass over the sequence. Each input must therefore be compressed before the model knows how it will later be used, forcing a limited state to compromise across possible future demands. We introduce dynamic compression, which allows a recurrent model to selectively revisit past tokens and revise its fixed-size state through additional recurrent updates. The model need not preserve every part of the history at uniformly high fidelity in its recurrent state, because lower-fidelity information can be revisited from the retained raw sequence when it becomes relevant. We study this in a controlled setting where the model first learns multiple functions in-context and, later in the same sequence, encounters a series of few-shot tasks that each require it to identify and reuse one of those functions. A single-pass model must preserve every function at sufficient fidelity for any future task, whereas selective re-scanning allows the model to revisit and refine only the function currently needed. We find that dynamic compression substantially reduces the recurrent state required for accurate reuse and scales more favorably as the number of stored functions grows. These results demonstrate a computation--memory tradeoff in which recurrent models can spend more computation revisiting their history to make more effective use of a fixed-size state.

Jyothish Pari, Ryan Bahlous-Boldi, Pulkit Agrawal · 0 citations
#machine learning Preprint Aug 2026

Efficient Resource Optimization for Split Federated Learning

This work establishes an efficient optimization framework for SFL under resource-constrained networks that jointly optimizes model splitting and resource allocation to minimize training cost, which is defined as the weighted sum of latency and energy costs.

Wei Wei, Xianhao Chen · 0 citations
#machine learning Preprint Aug 2026

MoRAX: Mobility-based Representation Augmentation for Geospatial Foundation Models

MoRAX, a lightweight framework for augmenting geospatial embeddings with functional structure derived from human mobility, is introduced and transfer results across countries further demonstrate that modulation conditioned on mobility flows provides a general mechanism for grounding geospatial foundations in the human dimension of cities.

Ya Wen, Jixuan Cai, Yulun Zhou et al. · 0 citations
#machine learning Preprint Aug 2026

Leveraging Association Context Retrieval in Knowledge Edit- ing to Build White-Box Attacks on LLMs

A novel white-box attack inspired by locate-then-edit approaches from the field of Knowledge Editing is proposed, modifying the editing framework by incorporating as- sociative knowledge retrieved from the model, thereby extending constraint removal to an entire thematic category rather than being limited to prompts from a predefined dataset.

R. Maksimov, Vladimir Aletov, V. Solodkin et al. · 0 citations

From tech blogs

See all →
MIT News · Artificial Intelligence Aug 27, 2026

Looking beyond natural sequences

A new machine-learning framework aims to improve the success rate of computational protein design while moving away from results that reproduce sequences found in nature.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.