Marginal Coverage Credit for PGPSE (MCC-PGPSE), which combines leave-one-policy-out coverage with state-owner specialization to estimate policy-specific credit, is introduced and supports contribution-conditioned auxiliary reward allocation as an interpretable approach to improving complementary coverage among parallel policies in discrete state spaces.
Junhao Cao, Hong-Yi Xia, Jianian Wu et al.· 0 citations
The proposed DMRA, a deficit-based diagnostic framework that quantifies the contribution of these components to identify the primary cause of unsuccessful cases, reveals that relational reasoning is the primary source of error across all models, followed by memory limitations.
N. Yilmaz, Naga Sai Abhiram Kusumba, Stella Wenxing Liu et al.· 0 citations
This study evaluates survival models that predict time-to-repurchase directly, and adopts Exponential AFT for probability-consuming surfaces and Log-Normal for pure ranking, exposing a principled calibration-ranking trade-off within a single AFT family.
Akshay Kekuda, Shreeranjani Srirangamsridharan, Ishan Bhatt et al.· 0 citations
GRACE, a gradient-guided coreset selection method that constructs both forget and retain sets for LLM unlearning, improves model utility while maintaining comparable forget quality, with particularly consistent gains over prior gradient-based selection methods.
Praveen Bushipaka, Andrea D'Angelo, Lucia C. Passaro et al.· 0 citations
Reach audiences
Advertise in front of researchers, engineers, and readers.
This work proposes a lookahead-guided decoding framework for context-free grammars based on pushdown automata based on bounded pushdown summaries with reachability labels and upper-bound distances to acceptance.
Vincenzo Collura, Karim Tit, Eleonora Giunchiglia et al.· 0 citations
This patient-grouped benchmark identifies contactless overnight sensing as a promising biomedical engineering direction for agitation-risk research in hospitalized dementia cohorts.
Zhen Liu, Marta Bono, R. Decloedt et al.· 0 citations
A six-category information taxonomy and four dimensions of linguistic style are defined and applied to real user prompts from SWE-chat and problem statements from SWE-bench Verified and Pro, finding that explicitly stating the desired behavior and motivation, which most real prompts omit, substantially improves the LLM's software engineering performance.
Gyuhyeong Kim, Hyojung Gwon, Jeonghyeon Kim et al.· 0 citations
This work introduces CURA (Certified Runtime Alarms for Computer-Use Agents), an external monitor that reads only harness-visible telemetry, with no model internals, extra LLM calls, or prompt changes, and turns the running trajectory into a sequential test with certified false-alarm control.
Divake Kumar, Sina Tayebati, Devashri Naik et al.· 0 citations
Tea plantations are vulnerable to Postelectrotermes militaris, commonly known as the Upcountry Live Wood Termite (ULWT), which can cause substantial damage when infestations remain undetected. This study proposes an IoT-enabled acoustic monitoring framework integrated with deep learning for early detection and severity assessment of ULWT infestations in tea plantations. Research Method: Audio signals were captured non-invasively from tea trunks using a high-sensitivity microphone connected to a Raspberry Pi-based IoT device, with geographic coordinates recorded for spatial tracking. After trimming, resampling, and segmentation, 2,000 ten-second samples were obtained, comprising 1,000 healthy and 1,000 infested samples, and divided into 1,600 training, 200 validation, and 200 test samples. The dataset used in this study is publicly available on Kaggle (Senevirathna et al. 2026). Fourier-derived spectrograms trained a CNN for infestation classification and probability estimation. A weighted severity model combined CNN probability, mean acoustic amplitude, and nearby infested plants within 5 m, with geospatial mapping used to visualize infestation distribution. Findings and Values: Field trials in a ULWT-affected tea plantation in Pundaluoya demonstrated feasibility under realistic environmental noise. On the held-out test set, the CNN achieved 81.5% accuracy, 80.6% precision, 83.0% recall, 81.8% F1-score, and 0.819 ROC-AUC. Beyond binary infestation detection, the framework introduced quantitative severity assessment using infestation probability, acoustic amplitude, and nearby infested plants. The resulting severity and geospatial outputs can support plantation managers in identifying high-risk areas, prioritizing field inspections, and implementing more timely and targeted control measures.
D. K. C. Senevirathna, A. Nanayakkara, H. M. C. K. Kulathunga et al.· 0 citations
Results show that agent-guided hypothesis refinement can recover heterogeneous governing laws without prescribing a parametric form for their spatial coefficients.
Yu-Jie Huang, Wenwu He, Zhuoxiao Lin et al.· 0 citations
To the authors' knowledge, JIT-Agent is the first model purpose-built for just-in-time harness generation, establishing harness intelligence as a trainable, transferable, and compounding dimension of agent capability orthogonal to model scaling.
Guibin Zhang, Leo Lu, Fang-Zhou Xie et al.· 0 citations
This method learns permanent handoff policies from accumulated trajectory evidence and Teacher-Annotated Censored Intervention Times (TACIT) and represents each annotation as an interval-censored observation on a cumulative-risk scale and achieves the highest held-out success among learned policies on both ALFWorld and DABench.
A new machine-learning framework aims to improve the success rate of computational protein design while moving away from results that reproduce sequences found in nature.
MIT News · Artificial Intelligence· news.mit.eduAug 24, 2026
A new method for surgically removing training examples from a model reveals that as datasets grow, the link between what a model learns and what it produces dissolves.