This work proposes a unified multi-domain and multi-task training framework that models heterogeneous event schemas within a single model, enabling dynamic adaptation to dataset-specific schemas without requiring complete event label sets at inference time.
Siting Liang, Omar Adjali, Daniel Sonntag· 0 citations
Experiments on long-term conversational memory show that mixed observed-and-estimated reranking improves answer accuracy over semantic retrieval by up to 6.62% and remains effective when only 17.5% of candidates receive direct LLM relevance scores, thereby substantially reducing the inference overhead of LLM reranking.
An LLM router picks which model should answer each query. The appeal is that models fail on different questions. Whatever single model is best overall still gets some wrong, and another model in the pool gets many of those right. Getting that choice right every time is the ceiling, and a router is an attempt to approach it. However, recent work reports that routers do not get close. Across 21 routing methods on five benchmarks, sharply different designs land within a fraction of a point of each other, and all of them stay far below that ceiling. Learned routers often fail to beat simply always calling the strongest model. We ask what those missed questions have in common. We set fourteen models to answer all 294 questions, with 7 task types across 3 languages: Korean, English and Hindi. We ran the whole matrix twice, changing nothing, but 5.37% of the 4,116 model-question pairs came out scored differently anyway. Run-to-run movement like that is normal, and we argue that a small win does not show that routing did anything, ours or anyone else's. Counting an answer correct only when the model got it right in both runs, 29 questions on this matrix can be improved with routing. Every correct-answer count here is on that rule. Task type accounts for most of them: assigning each task type one model in advance, chosen once and never updated, improves 21 of the 29. Splitting each task type by language improves 2 more and leaves 6 of 294 unoptimized. That handful is what a learned router would have been built for, and it is smaller than the run-to-run movement above, which is a share of pairs rather than of questions. The static table we adopted answers 262 of 294 questions at $3.33 per run, against the best single model's 245 at $7.69. All of this is fitted and scored on the same 294 questions with no holdout.
MIND was able to answer many questions about escitalopram in a manner deemed accurate, complete, and safe by psychiatrists the majority of the time, and represents a step towards building safe LLM systems to enhance patient education in psychiatry.
Alexander J. Hish, A. Nagendran, S. Compton· 0 citations
Reach audiences
Advertise in front of researchers, engineers, and readers.
This work introduces future querying, a paradigm that probes whether large language models can function as implicit medical world models by evaluating their ability to answer time-indexed clinical queries about a patient's future, and shows that small, locally fine-tuned open-weight models can match or approach larger proprietary systems, making the framework suitable for privacy-preserving, on-premise deployment.
Siri Willems, James Butterworth, L. Goetschalckx et al.· 0 citations
This work compares data-driven knowledge distillation (DDKD) against zero-shot inference and fine-tuning on out-of-domain D2T data, and introduces structure-preserving augmentation via structural subsampling and perturbation in cross-domain D2T generation.
Yifei Song, Kun Efimov-Zhang, Claire Gardent· 0 citations
This work proposes ProxyFormer, a general dual-stream architecture built upon proxy tokens, and introduces factorized multi-level compression/decompression, layer-wise dynamic compression ratios, asymmetric dual embeddings, and a proxy-only KV-cache inference scheme.
This work investigates LLM-based forecasting agents, meaning systems in which a language model contributes to a scored prediction about a future or currently unobserved target, and organizes architectures into three groups.
Xiaogang Xu, Jiaqi Tang, Jianmin Chen et al.· 0 citations
This work proposes StrategyBench, which selects strategy-inducible tasks from BIG-Bench, constructs reference strategies, and defines evaluation metrics along two dimensions: strategy quality and downstream utility, and experiments show that explicit strategy utility differs substantially across task categories and depends on both strategy generation and execution conditions.
Jinghan Tan, Yuanzheng Wang, Lu Chen et al.· 0 citations
Green-patent indicators based on Cooperative Patent Classification Y02 tags increasingly inform research, industrial policy, and climate-oriented investment, yet their construct validity has not been evaluated at corpus scale. We ask whether Y02 classification errors are random measurement noise or systematic, direction-specific bias. We introduce an Error-as-Signal framework in which disagreement between an administrative label and an independent model is treated as evidence of potential measurement error. Screening 9,075,421 USPTO granted patents from 1962-2024 with a fine-tuned domain model identifies 517,772 disagreements. Two independent open-weight large language models then assess whether each flagged invention has a direct climate-mitigation or adaptation function. Cross-model consensus identifies 180,384 administrative Type I errors (False Green) and 29,465 Type II errors (Silent Green). Correcting these errors reduces the measured green-patent population by 25.5%, from 592,387 to 441,468 patents. Misclassification is systematic rather than random. Atypicality predicts Silent Green in an inverted-U pattern, while reflection complexity independently increases under-recognition: controlling for atypicality and filing year, a one-standard-deviation increase is associated with 1.61 times the odds of Silent Green. Structural complexity has the opposite association. Among consensus-attributed errors, the same increase in reflection complexity is associated with 2.45 times the odds that an error is Silent Green rather than False Green. Event tests show no discrete rise in misclassification when green classification became more salient and only limited evidence of increased explicit green framing after the 2013 CPC launch. The evidence is more consistent with bounded classification capacity than with applicant gaming.
Hamid Bekamiri, Jan Auernhammer, Milad Abbasiharofteh et al.· 0 citations
We present HelaBERT, a family of two BERT-based masked language models pre-trained from scratch on approximately 1 billion tokens of Sinhala text sourced from MADLAD-400, CulturaX, and a custom corpus comprising news articles, Sinhala Wikipedia, and web crawl data. HelaBERT-Small (~23.3M parameters, 6 layers) and HelaBERT-Large (~110M parameters, 12 layers) both use a SentencePiece Unigram tokenizer (vocabulary size 32,000) tailored to Sinhala's agglutinative morphology and complex script. We evaluate both models on four downstream Sinhala text classification tasks: news category classification, news source classification, sentiment analysis, and writing style classification, using 5 independent seed runs with stratified 80/20 train/test splits. We additionally propose a dual pooling classification head and evaluate it systematically across all four tasks, finding consistent improvements on sentiment analysis and a moderate gain on news category classification for HelaBERT-Small, while the standard [CLS]-linear head remains competitive on news source classification, a headline-level task with short average input length. We release both models to support further research in Sinhala NLP.
An audit is introduced that predicts a mechanism's fate as a population grows, asking how often the mechanism can act, whether agents use the information it supplies, and whether the measurement itself creates apparent scale effects.
What if pathology foundation models could do more with less? GigaPath-Flash and GigaTIME-Flash cut computational demands while maintaining strong performance, opening the door to larger studies and broader exploration. The post GigaPath-Flash and GigaTIME-Flash: Toward population-scale discovery with efficient pathology foundation models appeared first on Microsoft Research.
MIT News · Artificial Intelligence· news.mit.eduAug 31, 2026
With millions of users across the world, Julia has been used to conduct cutting-edge research and to design new drugs, jet engines, heat pumps, and more.