Skip to content

Category

large language models

328 papers

#large language models Open access Aug 2026

A multi-task video reasoning dataset across semantic, spatial, and temporal categories with four difficulty levels

Visual reasoning, the capability to interpret visual input in response to implicit text query through multi-step reasoning, remains a challenge for deep learning models due to the lack of relevant benchmarks. Previous work in visual reasoning has primarily focused on reasoning segmentation, where models aim to segment objects based on implicit text queries. This paper introduces reasoning visual tasks (RVTs), a unified formulation that extends beyond traditional video reasoning segmentation to a diverse family of visual language reasoning problems, which can therefore accommodate multiple output formats including bounding boxes, natural language descriptions, and question-answer pairs. Correspondingly, we identify the limitations in current benchmark construction methods that rely solely on large language models (LLMs), which inadequately capture complex spatial-temporal relationships and multi-step reasoning chains in video due to their reliance on token representation, resulting in benchmarks with artificially limited reasoning complexity. To address this limitation, we propose a novel automated RVT benchmark construction pipeline that leverages digital twin (DT) representations as structured intermediaries between perception and the generation of implicit text queries. Based on this method, we construct RVTBench, a RVT benchmark containing 3,896 queries of over 1.2 million tokens across four types of RVT (segmentation, grounding, VQA and summary), three reasoning categories (semantic, spatial, and temporal), and four increasing difficulty levels, derived from 200 video sequences. Finally, we propose RVTagent, an agent framework for RVT that allows for zero-shot generalization across various types of RVT without task-specific fine-tuning. Dataset and code are available at https://doi.org/10.5281/zenodo.19697191 and https://github.com/yiqings/rvt .

Yiqing Shen, Chenjia Li, Chenxiao Fan et al. · 0 citations
#large language models Open access Aug 2026

Two Americas of Well-Being: Divergent Rural–Urban Patterns of Life Satisfaction and Happiness from 2.6 B Social Media Posts

Abstract Using 2.6 billion geolocated tweets (2014–2022) and a fine-tuned generative language model, we construct county-level indicators of life satisfaction and happiness for the United States. We document an apparent rural–urban paradox : even in unadjusted county-level means, rural counties express higher life satisfaction while urban counties exhibit greater happiness . This opposite gradient persists and is further characterized once the two are treated as distinct layers of subjective well-being, evaluative vs. hedonic, showing that each maps differently onto place, politics, and time. Democratic-leaning areas show a suggestive negative association with evaluative well-being, conditional on structural and temporal controls, but this effect is modest and does not extend to happiness, where no meaningful partisan gradient emerges. Temporal shocks dominate the hedonic layer: happiness falls sharply during 2020–2022, whereas life satisfaction moves more modestly. These patterns are robust across logistic and OLS specifications with clustered standard errors and align with well-being theory. Interpreted as associations for the population of geolocated tweets, the results show that large-scale, language-based indicators can help clarify why prior findings about the rural–urban divide may differ by distinguishing the type of well-being expressed, offering a transparent, reproducible complement to traditional surveys.

Stefano M. Iacus, Giuseppe Porro · 0 citations
#large language models Open access Aug 2026

Impact of Intermediate-Task Language Count on Zero-Shot Cross-Lingual Transfer Performance

Intermediate-task training---fine-tuning a pretrained model on an intermediate task before fine-tuning again on the target task---often improves model performance substantially on language understanding tasks in monolingual English settings. We investigate whether English intermediate-task training is still helpful on non-English target tasks. Using nine intermediate language-understanding tasks, we evaluate intermediate-task transfer in a zero-shot cross-lingual setting on the XTREME benchmark. We see large improvements from intermediate training on the BUCC and Tatoeba sentence retrieval tas Research goal: What is the impact of varying the number of languages in intermediate-task training on zero-shot cross-lingual transfer performance, measured by XGLUE accuracy and F1 scores? Autonomous synthesis report generated by Assignee Research. Tribunal consensus score: 9.2/10.

Assignee Research · 0 citations
#large language models Open access Aug 2026

mT5 Efficiency in Low-Resource Languages via Intermediate-Task Training

Intermediate-task training---fine-tuning a pretrained model on an intermediate task before fine-tuning again on the target task---often improves model performance substantially on language understanding tasks in monolingual English settings. We investigate whether English intermediate-task training is still helpful on non-English target tasks. Using nine intermediate language-understanding tasks, we evaluate intermediate-task transfer in a zero-shot cross-lingual setting on the XTREME benchmark. We see large improvements from intermediate training on the BUCC and Tatoeba sentence retrieval tas Research goal: How does intermediate-task training on English data influence the efficiency of mT5 in low-resource languages (e.g., inference speed, memory usage) while maintaining zero-shot cross-lingual reasoning performance on XTREME-R, measured by throughput (tokens/sec) and model latency? Autonomous synthesis report generated by Assignee Research. Tribunal consensus score: 9.2/10.

Assignee Research · 0 citations
#large language models Open access Aug 2026

Impact of Intermediate-Task Language Count on Zero-Shot Cross-Lingual Transfer Performance

Intermediate-task training---fine-tuning a pretrained model on an intermediate task before fine-tuning again on the target task---often improves model performance substantially on language understanding tasks in monolingual English settings. We investigate whether English intermediate-task training is still helpful on non-English target tasks. Using nine intermediate language-understanding tasks, we evaluate intermediate-task transfer in a zero-shot cross-lingual setting on the XTREME benchmark. We see large improvements from intermediate training on the BUCC and Tatoeba sentence retrieval tas Research goal: What is the impact of varying the number of languages in intermediate-task training on zero-shot cross-lingual transfer performance, measured by XGLUE accuracy and F1 scores? Autonomous synthesis report generated by Assignee Research. Tribunal consensus score: 9.2/10.

Assignee Research · 0 citations
#large language models Open access Aug 2026

mT5 Efficiency in Low-Resource Languages via Intermediate-Task Training

Intermediate-task training---fine-tuning a pretrained model on an intermediate task before fine-tuning again on the target task---often improves model performance substantially on language understanding tasks in monolingual English settings. We investigate whether English intermediate-task training is still helpful on non-English target tasks. Using nine intermediate language-understanding tasks, we evaluate intermediate-task transfer in a zero-shot cross-lingual setting on the XTREME benchmark. We see large improvements from intermediate training on the BUCC and Tatoeba sentence retrieval tas Research goal: How does intermediate-task training on English data influence the efficiency of mT5 in low-resource languages (e.g., inference speed, memory usage) while maintaining zero-shot cross-lingual reasoning performance on XTREME-R, measured by throughput (tokens/sec) and model latency? Autonomous synthesis report generated by Assignee Research. Tribunal consensus score: 9.2/10.

Assignee Research · 0 citations
#large language models Open access Aug 2026

Moral Concern Is a Response Regime: A Preregistered Cross-Model Test of Persona Framing in Large Language Models

Four large language models from four providers judged the same 20 ambiguous scenarios under a neutral frame and six loaded personas, across 11,200 main trials. The prediction came from Motivated Violation Construction (MVC), a theory of human moral cognition on which a loaded actor identity can move ambiguous behaviour into the violation category. A registered test-retest gate fixed the confirmatory panel, and semantically inert stimuli tested for frame-driven responding with no behaviour to classify. The models occupied sharply different operating ranges. Two classified almost every scenario as concerning under the neutral frame, leaving little headroom for the predicted increase, and one failed the correlation-based stability gate under near-ceiling responding. Across the retained three-model panel, none of four MVC predictions was supported, and the registered adjudication favoured valence priming over construction. On inert stimuli, three models almost never expressed concern, while one did so on roughly a third of trials and more often under loaded personas. The transfer test was therefore not supported. Whether that is also a refutation turns on detectability, which differed sharply across the retained panel. The durable result is a measurement result. Baseline, headroom, retest behaviour, sensitivity to prompt position and wording, refusal and missingness, and false entry on inert stimuli all differed by model. We call that configuration a model’s response regime, and it meant the same forced question carried different information in each system. A shared prompt is not a shared instrument, and a forced moral verdict measures the regime before it measures anything moral.

Emile Boullineau, José Daniel Muñoz Arciniegas · 0 citations

From tech blogs

See all →
Microsoft Research Blog Aug 31, 2026

GigaPath-Flash and GigaTIME-Flash: Toward population-scale discovery with efficient pathology foundation models

What if pathology foundation models could do more with less? GigaPath-Flash and GigaTIME-Flash cut computational demands while maintaining strong performance, opening the door to larger studies and broader exploration. The post GigaPath-Flash and GigaTIME-Flash: Toward population-scale discovery with efficient pathology foundation models appeared first on Microsoft Research.