Skip to content

Category

natural language processing

1,585 papers

#artificial intelligence Preprint Aug 2026

MoNe: Modular Neural Memory for Efficient Long Context Inference

MoNe is a lightweight modular neural memory that attaches to any frozen pretrained Transformer to enable long-context inference without retraining, achieving strong performance on needle-in-a-haystack and word extraction benchmarks from RULER, where ICL degrades sharply.

Won-Yong Cho, Kyubyung Chae, Tribhuvanesh Orekondy et al. · 0 citations
#machine learning Preprint Aug 2026

Reflex-Guard: A Low-Latency Guardrail for LLM Prompt Safety Using Dense Semantic Embeddings

Reflex-Guard is introduced, a lightweight guardrail that runs locally that uses jailbreak-aware preprocessing, compact sentence-transformer embeddings, and seven fast binary classifiers that enable high-accuracy prompt safety filtering with much lower latency than existing solutions.

Istiaque Ahmed, Afia Anjum Borsha, Ranat Das Prangon et al. · 0 citations
#machine learning Preprint Aug 2026

Temporal Leakage in Financial News NLP: A Multi-Architecture Audit with a Regime-Specific M&A Signal

This work audits financial-news direction prediction dependence on a 49,799-article corpus across 16 feature-model combinations spanning TF-IDF, MiniLM, FinBERT, and fine-tuned RoBERTa-large / DeBERTa-v3-large, plus separate zero/few-shot and LoRA probes of Llama-3 and Qwen2.

Chenhao Xue, Raslen Guesmi, Siwei Feng et al. · 0 citations
#machine learning Preprint Aug 2026

Margin-Regularized Structured Semantic Alignment for Brain-Language Correspondence

This work proposes a margin-regularized structured semantic alignment framework that directly aligns brain embeddings with text embeddings in a shared semantic space, enabling retrieval-based decoding and enables explicit modeling of the correspondence between neural representations and language semantics.

Jiaqi Wang, Huawen Hu, Shu Zhang · 0 citations

Intent-Driven Dynamic Chunking: Segmenting Documents to Reflect Predicted Information Needs

Intent-Driven Dynamic Chunking (IDC) is introduced, a novel approach that uses predicted user queries to guide document segmentation and aligning document structure with anticipated information needs significantly boosts retrieval performance, particularly for long and heterogeneous documents.

Christos Koutsiaris · 0 citations
#machine learning Preprint Aug 2026

An Empirical Study of Reward Specification and Benchmark Reliability in GRPO-based LLM Unlearning

This work compares four reward designs that span lexical suppression, anti-refusal shaping, rubric-based broad answering, and an explicit refusal contrast, and shows that optimization success is not equivalent to behavioral unlearning: RWKU forget scores, held-out completion audits, terminal training-rollout audits, and training dynamics can point to different conclusions.

Rubén Balbastre, J. Orduña, M. Perez · 0 citations
#artificial intelligence Preprint Aug 2026

Domain-Adapted Molecular Language Models for Efficient Search of Make-on-Demand Libraries

Pretrained molecular language models are increasingly used as molecular encoders for learning structure-property relationships. However, their practical suitability for molecular discovery within and beyond their pretraining domain remains unclear. Herein, we systematically benchmark four molecular language models across six virtual molecular libraries spanning drug discovery, organic materials, and catalysis. Native molecular language model embeddings show substantial variation in discovery performance across libraries, whereas molecular fingerprints provide a consistently strong and robust baseline. Consistent with a potential domain-representation mismatch, we show that explicit domain adaptation substantially improves representation performance. Fine-tuning molecular language model encoders on structures from the target virtual library consistently improves sample efficiency, with several adapted encoders emerging as the top-performing representations across the benchmark tasks. These results show that molecular representation quality depends strongly on the target domain and that explicit adaptation can improve the practical utility of molecular foundation models. More broadly, our findings establish domain-adapted molecular representations as a promising strategy for sample-efficient adaptive decision making in virtual screening and self-driving laboratories.

Henrik Wille, Luis-Finley Schütz, Felix Strieth-Kalthoff · 0 citations
#machine learning Preprint Aug 2026

J-Miner: Recovering Executable Decision Knowledge from Language-Model Classifiers

J-Miner is introduced, which mines text-level named concepts by aggregating vocabulary-aligned internal signals across layers and token positions, and uses the classifier's own predictions to learn executable decision rules over them, and shows that task-specific decision knowledge can be faithfully represented in an explicit, executable form and reused beyond the classifier in which it was learned.

Yunfan Gao, Xinyi Huang, Tao Sheng et al. · 0 citations
#artificial intelligence Preprint Aug 2026

Cross-Model Memory Transfer via Target-Side Reader Adaptation

The results suggest that Engram can serve as a reusable external knowledge artifact, provided that the target has access to a compatible reader interface and target-side adaptation can further improve alignment when direct reader reuse is insufficient.

Mingyuan Li, Guangsheng Yu, Xu Wang et al. · 0 citations
#artificial intelligence Preprint Aug 2026

Neurosymbolic Embodied Agents

A neurosymbolic agent that factors long-horizon household tasks into task-directed visual exploration and constrained symbolic planning and evaluates executable continuations using a domain-independent planning heuristic is presented.

Mohammad Albinhassan, Yuming Feng, Alessandra Russo et al. · 0 citations

From tech blogs

See all →
MIT News · Artificial Intelligence Aug 27, 2026

Looking beyond natural sequences

A new machine-learning framework aims to improve the success rate of computational protein design while moving away from results that reproduce sequences found in nature.

MIT News · Artificial Intelligence Aug 20, 2026

Paving the way for greener ammonia production

New MIT research could lead to better materials for a fossil-fuel-free process for making the chemical that's essential to fertilizer and other products.