Semi-structured information extraction (IE) from OCR-derived clinical reports is crucial for efficiently reconstructing patients' longitudinal medical histories. In practice, this scenario commonly involves three tasks: (i) field-header (key) discovery, (ii) key-conditioned question answering (QA), and (iii) end-to-end key-value pair extraction. However, existing evaluations often under-model two factors: heterogeneous and incompletely known key representations, and OCR-induced noise. This makes it difficult to assess model robustness in real-world settings.
We present MedStruct-S, a benchmark specifically designed to evaluate these tasks under unknown keys and OCR noise. MedStruct-S contains 3,582 annotated real-world clinical report pages. Using MedStruct-S, we benchmark two representative paradigms: encoder-only sequence labeling with post-processing and decoder-only structured generation, covering four encoder-only and five decoder-only models spanning 0.11B to 103B parameters. Our results show that encoder-only models achieve the best performance for non-null-value key-conditioned QA despite being substantially smaller than decoder-only models. When comparing models of similar order of magnitude, encoder-only models still perform better overall. Without controlling for model scale, fine-tuned decoder-only models deliver the strongest overall results. These findings show that the benchmark provides a reliable and practical basis for selecting and comparing models across different semi-structured IE settings.
This work presents AutoOR, a scalable synthetic data generation and reinforcement learning pipeline that trains LLMs to autoformulate optimization problems specified in natural language across linear, mixed-integer, and non-linear categories and introduces a curriculum RL strategy that bootstraps from limited initial training data to make this class tractable for post-training.
S. Motwani, Chuan Du, A. Petrov et al.· 2 citations
Findings show that preserving the temporal extent of recurrent history is important for efficient whole-piece modeling, and that memory cost can instead be reduced through KV representation compression.
Yungang Yi, Weihua Li, Matthew Kuo et al.· 0 citations
FiLoRA is introduced, an instruction-conditioned, parameter-efficient adaptation framework that enables controllable modulation of feature reliance while keeping the task and predictive objective fixed and suggests that instruction-conditioned parameter adaptation can serve as a practical mechanism for intervening on internal model behavior.
Hyunsuk Chung, Caren Han, Yerin Choi et al.· arXiv.org· 1 citation
Reach audiences
Advertise in front of researchers, engineers, and readers.
This work offers the first random generator of synthetic causal benchmarks with coverage guarantees and transparent assumptions operating on the three levels of causal reasoning: observation, intervention, and counterfactual, and demonstrates its utility by evaluating several state-of-the-art methods under diverse conditions and assumptions.
Panayiotis N. Panayiotou, Audrey Poinsot, A. Leite et al.· arXiv.org· 0 citations
NINJA (short for Needle-in-haystack jailbreak attack), a method that jailbreaks aligned LMs by appending benign, model-generated content to harmful user goals to reveal fundamental vulnerabilities in modern LMs.
R. Shah, C. Wu, Shashwat Saxena et al.· arXiv.org· 4 citations
This work explores image generation using flow matching using flow matching and proposes an iterative process that can be integrated into virtually any generative modeling technique, thereby enhancing the performance and robustness of image synthesis systems.
Eldad Haber, Shadab Ahamed, Md Shahriar Rahim Siddiqui et al.· SIAM Journal on Scientific C...· 3 citations
This work evaluates Constrained Bayesian Optimization with the primary objective of minimizing energy consumption and subject to the constraint that the generalization performance is above some threshold and demonstrates that CBO achieves lower energy consumption without compromising the predictive performance of ML models.
Pallavi Mitra, F. Biessmann· arXiv.org· 0 citations
Mechanist is an agentic system that uses AI as a scientific instrument for the autonomous discovery of mechanisms underlying AI intelligence, and develops a mechanism theory of belief, revealing how models represent world knowledge, form beliefs, infer the beliefs of others, and how these mechanisms emerge during pretraining.
Mengru Wang, Junfeng Fang, Shuofei Qiao et al.· 0 citations
A unified benchmark for comprehensive, instruction-conditioned EEG understanding is introduced and results vary substantially across models, subsets, difficulty levels, and execution paradigms, showing that EEG competence depends on the model and its operationalization.
Yangxuan Zhou, Sha Zhao, Yuning Chen et al.· 0 citations
This paper introduces an LLM-Augmented Reinforcement Learning Agent that integrates LLM-driven planning with RL-based action optimization, and highlights a promising direction for building more capable autonomous systems.
Christophe D. Hounwanou, John Emeka Eze, Yaé Ulrich Gaba· 0 citations
SkillBoost is proposed, a three-stage framework that mitigates both risks: structured exploitation localizes observed failures to editable skill components, prior-guided exploration draws on prior knowledge in the LLM to generate diverse repair candidates, and verified acceptance commits a candidate only when it improves performance within a regression bound.
Hongqiang Lin, Chao Liu, Xiaofan Bai et al.· arXiv.org· 1 citation
A weeklong summer workshop brought higher education faculty to campus to explore how AI and machine learning materials can be adapted for their classrooms.
A new machine-learning framework aims to improve the success rate of computational protein design while moving away from results that reproduce sequences found in nature.
MIT News · Artificial Intelligence· news.mit.eduAug 24, 2026