Visual reasoning, the capability to interpret visual input in response to implicit text query through multi-step reasoning, remains a challenge for deep learning models due to the lack of relevant benchmarks. Previous work in visual reasoning has primarily focused on reasoning segmentation, where models aim to segment objects based on implicit text queries. This paper introduces reasoning visual tasks (RVTs), a unified formulation that extends beyond traditional video reasoning segmentation to a diverse family of visual language reasoning problems, which can therefore accommodate multiple output formats including bounding boxes, natural language descriptions, and question-answer pairs. Correspondingly, we identify the limitations in current benchmark construction methods that rely solely on large language models (LLMs), which inadequately capture complex spatial-temporal relationships and multi-step reasoning chains in video due to their reliance on token representation, resulting in benchmarks with artificially limited reasoning complexity. To address this limitation, we propose a novel automated RVT benchmark construction pipeline that leverages digital twin (DT) representations as structured intermediaries between perception and the generation of implicit text queries. Based on this method, we construct RVTBench, a RVT benchmark containing 3,896 queries of over 1.2 million tokens across four types of RVT (segmentation, grounding, VQA and summary), three reasoning categories (semantic, spatial, and temporal), and four increasing difficulty levels, derived from 200 video sequences. Finally, we propose RVTagent, an agent framework for RVT that allows for zero-shot generalization across various types of RVT without task-specific fine-tuning. Dataset and code are available at https://doi.org/10.5281/zenodo.19697191 and https://github.com/yiqings/rvt .
Yiqing Shen, Chenjia Li, Chenxiao Fan et al.· Scientific Data· 0 citations
Abstract Using 2.6 billion geolocated tweets (2014–2022) and a fine-tuned generative language model, we construct county-level indicators of life satisfaction and happiness for the United States. We document an apparent rural–urban paradox : even in unadjusted county-level means, rural counties express higher life satisfaction while urban counties exhibit greater happiness . This opposite gradient persists and is further characterized once the two are treated as distinct layers of subjective well-being, evaluative vs. hedonic, showing that each maps differently onto place, politics, and time. Democratic-leaning areas show a suggestive negative association with evaluative well-being, conditional on structural and temporal controls, but this effect is modest and does not extend to happiness, where no meaningful partisan gradient emerges. Temporal shocks dominate the hedonic layer: happiness falls sharply during 2020–2022, whereas life satisfaction moves more modestly. These patterns are robust across logistic and OLS specifications with clustered standard errors and align with well-being theory. Interpreted as associations for the population of geolocated tweets, the results show that large-scale, language-based indicators can help clarify why prior findings about the rural–urban divide may differ by distinguishing the type of well-being expressed, offering a transparent, reproducible complement to traditional surveys.
Stefano M. Iacus, Giuseppe Porro· Journal of Happiness Studies· 0 citations
Intermediate-task training---fine-tuning a pretrained model on an intermediate task before fine-tuning again on the target task---often improves model performance substantially on language understanding tasks in monolingual English settings. We investigate whether English intermediate-task training is still helpful on non-English target tasks. Using nine intermediate language-understanding tasks, we evaluate intermediate-task transfer in a zero-shot cross-lingual setting on the XTREME benchmark. We see large improvements from intermediate training on the BUCC and Tatoeba sentence retrieval tas Research goal: What is the impact of varying the number of languages in intermediate-task training on zero-shot cross-lingual transfer performance, measured by XGLUE accuracy and F1 scores? Autonomous synthesis report generated by Assignee Research. Tribunal consensus score: 9.2/10.
Assignee Research· Zenodo (CERN European Organi...· 0 citations
Intermediate-task training---fine-tuning a pretrained model on an intermediate task before fine-tuning again on the target task---often improves model performance substantially on language understanding tasks in monolingual English settings. We investigate whether English intermediate-task training is still helpful on non-English target tasks. Using nine intermediate language-understanding tasks, we evaluate intermediate-task transfer in a zero-shot cross-lingual setting on the XTREME benchmark. We see large improvements from intermediate training on the BUCC and Tatoeba sentence retrieval tas Research goal: How does intermediate-task training on English data influence the efficiency of mT5 in low-resource languages (e.g., inference speed, memory usage) while maintaining zero-shot cross-lingual reasoning performance on XTREME-R, measured by throughput (tokens/sec) and model latency? Autonomous synthesis report generated by Assignee Research. Tribunal consensus score: 9.2/10.
Assignee Research· Zenodo (CERN European Organi...· 0 citations
Reach audiences
Advertise in front of researchers, engineers, and readers.
Intermediate-task training---fine-tuning a pretrained model on an intermediate task before fine-tuning again on the target task---often improves model performance substantially on language understanding tasks in monolingual English settings. We investigate whether English intermediate-task training is still helpful on non-English target tasks. Using nine intermediate language-understanding tasks, we evaluate intermediate-task transfer in a zero-shot cross-lingual setting on the XTREME benchmark. We see large improvements from intermediate training on the BUCC and Tatoeba sentence retrieval tas Research goal: What is the impact of varying the number of languages in intermediate-task training on zero-shot cross-lingual transfer performance, measured by XGLUE accuracy and F1 scores? Autonomous synthesis report generated by Assignee Research. Tribunal consensus score: 9.2/10.
Assignee Research· Zenodo (CERN European Organi...· 0 citations
Intermediate-task training---fine-tuning a pretrained model on an intermediate task before fine-tuning again on the target task---often improves model performance substantially on language understanding tasks in monolingual English settings. We investigate whether English intermediate-task training is still helpful on non-English target tasks. Using nine intermediate language-understanding tasks, we evaluate intermediate-task transfer in a zero-shot cross-lingual setting on the XTREME benchmark. We see large improvements from intermediate training on the BUCC and Tatoeba sentence retrieval tas Research goal: How does intermediate-task training on English data influence the efficiency of mT5 in low-resource languages (e.g., inference speed, memory usage) while maintaining zero-shot cross-lingual reasoning performance on XTREME-R, measured by throughput (tokens/sec) and model latency? Autonomous synthesis report generated by Assignee Research. Tribunal consensus score: 9.2/10.
Assignee Research· Zenodo (CERN European Organi...· 0 citations
Four large language models from four providers judged the same 20 ambiguous scenarios under a neutral frame and six loaded personas, across 11,200 main trials. The prediction came from Motivated Violation Construction (MVC), a theory of human moral cognition on which a loaded actor identity can move ambiguous behaviour into the violation category. A registered test-retest gate fixed the confirmatory panel, and semantically inert stimuli tested for frame-driven responding with no behaviour to classify. The models occupied sharply different operating ranges. Two classified almost every scenario as concerning under the neutral frame, leaving little headroom for the predicted increase, and one failed the correlation-based stability gate under near-ceiling responding. Across the retained three-model panel, none of four MVC predictions was supported, and the registered adjudication favoured valence priming over construction. On inert stimuli, three models almost never expressed concern, while one did so on roughly a third of trials and more often under loaded personas. The transfer test was therefore not supported. Whether that is also a refutation turns on detectability, which differed sharply across the retained panel. The durable result is a measurement result. Baseline, headroom, retest behaviour, sensitivity to prompt position and wording, refusal and missingness, and false entry on inert stimuli all differed by model. We call that configuration a model’s response regime, and it meant the same forced question carried different information in each system. A shared prompt is not a shared instrument, and a forced moral verdict measures the regime before it measures anything moral.
Emile Boullineau, José Daniel Muñoz Arciniegas· Zenodo (CERN European Organi...· 0 citations
What if pathology foundation models could do more with less? GigaPath-Flash and GigaTIME-Flash cut computational demands while maintaining strong performance, opening the door to larger studies and broader exploration. The post GigaPath-Flash and GigaTIME-Flash: Toward population-scale discovery with efficient pathology foundation models appeared first on Microsoft Research.
MIT News · Artificial Intelligence· news.mit.eduAug 31, 2026
With millions of users across the world, Julia has been used to conduct cutting-edge research and to design new drugs, jet engines, heat pumps, and more.