This work describes an inference-time architectural enhancement for off-the-shelf foundation models that markedly reduces perplexity and boosts accuracy across generation and reasoning tasks, and proposes and evaluates an adaptive variant of recirculation which requires only light tuning of hyperparameters while freezing the original model weights.
Michael C. Mozer, Shoaib Ahmed Siddiqui, Danny Sawyer et al.· Rhinology· 0 citations
Light is shed on the effect of dissimilarity between train and test feature distributions on forecasting models, compares deep learning versus non-deep learning models, and introduces modifications that are effective for non-deep learning models.
Log Reconstruction and Distance (LoRD), a lightweight post-hoc calibration framework for reliable log anomaly detection, is proposed and demonstrates that LoRD consistently improves confidence reliability and substantially reduces overconfident anomaly-related errors without sacrificing anomaly detection performance.
A task-centric, retrieval-based perspective is offered for how TFMs generalize: it is believed that tabular in-context generalization is largely retrieval-based, and good models are those that learn to identify relevant examples in the provided context and aggregate them well.
Nour Shaheen, Junwei Ma, Alex Labach et al.· 1 citation
Reach audiences
Advertise in front of researchers, engineers, and readers.
An LLM synthesizes an executable world model that a classical planner searches, and the model is accepted when it reproduces sampled transitions, and it is asked what that acceptance certifies in continuous control.
This work proposes a SHAP-enhanced Implicit-trajectory Generation for Metadata-free AutoFE (SIGMA), a scalable constant-context optimization framework that leverages SHAP values to provide task-aware signals for guiding group feature generation instead of semantic information.
A plug-and-play graph-based online difficulty estimator that shares rollout feedback across related samples and continuously updates their difficulty estimates, mitigating cold start and staleness without dedicated probing is proposed.
Zhizhao Liu, Zhiliang Tian, Xi Wang et al.· 0 citations
This work presents a hybrid and light-weight machine learning (ML) based approach that combines a decision tree with linear regression to improve pre-routing delay estimations generated by the open-source RTL-to-GDSII tool OpenLane.
Recurrent models process long contexts efficiently by compressing their history into a fixed-size state, but modern architectures typically do so in a single causal pass over the sequence. Each input must therefore be compressed before the model knows how it will later be used, forcing a limited state to compromise across possible future demands. We introduce dynamic compression, which allows a recurrent model to selectively revisit past tokens and revise its fixed-size state through additional recurrent updates. The model need not preserve every part of the history at uniformly high fidelity in its recurrent state, because lower-fidelity information can be revisited from the retained raw sequence when it becomes relevant. We study this in a controlled setting where the model first learns multiple functions in-context and, later in the same sequence, encounters a series of few-shot tasks that each require it to identify and reuse one of those functions. A single-pass model must preserve every function at sufficient fidelity for any future task, whereas selective re-scanning allows the model to revisit and refine only the function currently needed. We find that dynamic compression substantially reduces the recurrent state required for accurate reuse and scales more favorably as the number of stored functions grows. These results demonstrate a computation--memory tradeoff in which recurrent models can spend more computation revisiting their history to make more effective use of a fixed-size state.
Jyothish Pari, Ryan Bahlous-Boldi, Pulkit Agrawal· 0 citations
This work establishes an efficient optimization framework for SFL under resource-constrained networks that jointly optimizes model splitting and resource allocation to minimize training cost, which is defined as the weighted sum of latency and energy costs.
MoRAX, a lightweight framework for augmenting geospatial embeddings with functional structure derived from human mobility, is introduced and transfer results across countries further demonstrate that modulation conditioned on mobility flows provides a general mechanism for grounding geospatial foundations in the human dimension of cities.
Ya Wen, Jixuan Cai, Yulun Zhou et al.· 0 citations
A novel white-box attack inspired by locate-then-edit approaches from the field of Knowledge Editing is proposed, modifying the editing framework by incorporating as- sociative knowledge retrieved from the model, thereby extending constraint removal to an entire thematic category rather than being limited to prompts from a predefined dataset.
R. Maksimov, Vladimir Aletov, V. Solodkin et al.· 0 citations
A weeklong summer workshop brought higher education faculty to campus to explore how AI and machine learning materials can be adapted for their classrooms.
A new machine-learning framework aims to improve the success rate of computational protein design while moving away from results that reproduce sequences found in nature.
MIT News · Artificial Intelligence· news.mit.eduAug 24, 2026