Skip to content

Category

large language models

256 papers

#large language models Open access Aug 2026

The Reliability Wall : Why Scaling Cannot Buy Certifiable Compositional Reasoning, and Why It Is Permanent

Large language models are increasingly asked to perform tasks made of several dependent steps (multi-digit arithmetic, multi-step planning, rule-following, program synthesis) where the final answer is correct only if every one of those steps is correct; a single mistake anywhere invalidates the whole output. This paper shows that, on such tasks, three problems the field usually treats as separate and separately fixable (accuracy falling as tasks get longer, confidence scores becoming meaningless exactly when they are most needed, and the cost of obtaining a correct answer exploding) are in fact one and the same event, governed by a single quantity. That event is not a temporary shortcoming that more training will remove: it is a structural consequence of how these models compute, and it sets a hard, predictable limit on how far they can be trusted on deep, multi-step work. Formally, end-to-end accuracy declines geometrically with the number of dependent steps, because every added step multiplies the chance of a fully correct answer by the same fixed factor below one, so useful performance necessarily ends at a finite critical depth, beyond which accuracy, the trustworthiness of any confidence score, and the cost per correct answer all collapse together. The practical implication is that certifiable compositional reasoning is not a target that scaling today’s architectures will eventually reach: it requires moving to computation that is genuinely serial and exact, not a larger version of the same parallel one.

Daniele Sannino · 0 citations
#large language models Dataset Open access Aug 2026

JADE: A VSCode Plugin for Static Analysis-Guided AI-Assisted Refactoring

Software developers often use static analysis tools to identify code warnings related to code quality in general. However, despite the availability of automated warnings, developers still need to manually interpret and apply the suggested fixes. Recent advances in Large Language Models (LLMs) have created new opportunities to support automated refactoring suggestions directly within development environments. This paper presents JADE (Java Static Analysis Repair), a Visual Studio Code plugin that combines static analysis knowledge with LLM-based refactoring suggestions. JADE adopts a Retrieval-Augmented Generation (RAG) strategy based on SonarQube rules, retrieving semantically relevant static analysis heuristics to enrich structured prompts submitted to local LLMs. The plugin supports AI-assisted code diagnostics and refactoring generation integrated into the VSCode workflow. In addition, JADE incorporates a developer feedback mechanism that allows users to evaluate the usefulness and relevance of the generated recommendations. An exploratory study involving 50 Java code snippets suggests that JADE complements traditional static analysis by identifying additional semantic refactoring opportunities while preserving the strengths of rule-based analysis for detecting explicit code quality issues.

Ruan Neres, Carlos Eduardo Dantas · 0 citations
#large language models Open access Aug 2026

Attention as Race-Architecture: attention as the landscape-governed initiation of races

Selective attention, read at the level of the substrate, is the landscape-governed initiation of races: a bounded predictive system runs competing prediction-error resolutions ("races"), and what determines which races start is the system's installed landscape — in humans the four fields of Behavioural Friction Theory (Safety, Meaning, Ability, Effort); in a large language model a reduced, fine-tuning-installed landscape. Commit-order is a downstream readout, not the identity. The paper grounds this in the transformer (the attention-pattern softmax as the divisive-normalisation / biased-competition operation of neural attention; the output softmax as the downstream commit, by analogy with the accumulator model of choice) and reports a powered own-substrate result: across five vendor families, fine-tuning installs a small but robust "gap-registration" overlay — on under-determined curiosity gaps the instruct model registers the gap while the base substrate runs through (instruct−base +0.17, p<0.0001, 555 paired items). A loop-versus-feed-forward test finds no separate architectural "hold": recognising under-determination tracks compute (chain-of-thought) rather than looping, so the human–LLM difference is one of initiation, not maintenance. The account dissolves attention capture, maintenance, and decline into one mechanism (race-initiation), states falsifiable predictions, names the falsifiers, and invites the decisive mechanistic and human experiments. Series position. Paper 29 in the Behavioural Friction Theory paper-series; companion to Paper 0 (BFT) and the install-fields, social-friction, and integration-load studies it cross-cites. v2 (August 2026) — prior-art revision. The construct this paper is built on is credited to the literature that owns it. In vision, the representation determining which candidates enter competition at all is the saliency or priority map, and the two are now kept apart: a saliency map is computed from stimulus-feature contrast (Koch & Ullman, 1985; Itti, Koch & Niebur, 1998), while a priority map already integrates salience with relevance, value and selection history (Fecteau & Munoz, 2006). The landscape is the second, the office is not claimed as new, and the exogenous/endogenous timing division is conceded to that literature. What the paper proposes is the map's contents: that what populates it is four fields ordered by misclassification cost — an account of the inputs rather than a new mechanism for the selection. The maintenance-as-re-initiation claim now names Altmann and Trafton's (2002) memory-for-goals model, which already replaces a held state with activation that decays and must be re-strengthened; the reference had been listed and never used. Whether re-initiation is driven by the unresolved gradient itself rather than by a separate refresh process is stated as a conjecture and marked untested, and two overstatements are downgraded accordingly. Earlier versions remain in the version history.

Tomas Pødenphant Lund · 0 citations
#large language models Open access Aug 2026

RogueGPT: A Controlled Stimulus Generation Framework for News Authenticity Research

RogueGPT is a controlled stimulus generation framework for AI news authenticity research. It provides systematic, reproducible generation of multilingual news fragments across multiple large language models, journalistic styles, and content formats, together with a MongoDB-backed corpus, a CLI, a Streamlit web interface, and a Model Context Protocol (MCP) server for AI agent integration.

Alexander Loth, Martin Kappes, Marc‐Oliver Pahl · 0 citations
#large language models Open access Aug 2026

The Proto-Semitic Origins of E1b1b1: The Afro-Asiatic Founding-Fathers of Semitic Identity and Beyond

Version 2 — Changelog / Abstract (Version 2) What's new: Since the original Version 2 draft, this paper has grown from a single-thread argument about E1b1b's Levantine origins into a more rigorously sourced and considerably more self-critical treatment of the same core claim — that E1b1b, not J1-P58, represents the deeper indigenous Levantine paternal substrate underlying Semitic-speaking populations. Below is what actually changed. Subclade resolution (the core argument, tightened). The Samaritan priesthood's E-M78/E-V22 lineage is now explicitly distinguished from the Luria rabbinical line's E-V12, with the Natufian aDNA sample tables, a full non-Semitic J1/J1-P58 global distribution table, and independent G25 autosomal distance data all added as supporting evidence rather than left as prose assertions. E-M81 treated as an open question, not a foregone conclusion. Competing Near Eastern-origin and Northwest African-origin hypotheses are weighed against each other rather than resolved by fiat — including a genuine tension the source literature itself flags (STR diversity pointing east, TMRCA and ancient DNA pointing west). Csaba-Barnabás Horváth's (2021) independent YFull-based TMRCA estimate for E-M81 (~800 BCE) was folded in as a convergent third data point alongside Solé-Morata et al. (2017), modestly strengthening the Northwest African reading without treating either estimate as final. A wider intellectual context for the E-M78/Semitic question. Horváth's "Semitic re-migration" model is presented as a third account of a real paradox — E-M78's genetic diversity peaks in the Levant, but Afro-Asiatic's linguistic diversity peaks in Africa — alongside his separate proposal that specific J1/J2 subclades mark a pre-Semitic Indo-European population in Northern Mesopotamia, which directly reinforces this paper's existing argument that J1/J2 track later population movements rather than deep Semitic ancestry. Both are flagged clearly as a single author's hypotheses from a non-specialist venue, not consensus findings. A full craniofacial-morphology section. Reviews the historical Caucasoid/Negroid classification of Iberomaurusian and Nazlet Khater remains and shows the framework produced contradictory verdicts even among its own practitioners — the same population scored on opposite sides of the same racial axis by the same researchers. Includes a facial-reconstruction gallery graded explicitly by evidentiary tier (peer-reviewed vs. artist interpretation vs. commercial marketing), and is linked directly to the paper's existing "Aspirational Whiteness" discussion with a concrete case study. A primary-source interlude on the Book of Gates. The ancient Egyptian "Four Peoples" (Rmt/Aamu/Nhsyw/Tjhnw) iconography is tested against current genetic and isotopic evidence, identification by identification, with a mixed, honestly reported result — the Aamu/Levantine identification holds up well, Nhsyw/sub-Saharan affinity holds up as real but non-majority, Rmt/East African phenotype is genuinely contested in the literature, and Tjhnw/Sea Peoples fits some tomb versions but not others. Corrective housekeeping on circulating claims. Includes the corrected dating and provenance of the widely shared JK2134/JK2888/JK2911 facial reconstructions, a careful description (without characterizing it as fraudulent) of a circulating social-media claim about a "Saudi Genome Project," and a documented comparison of divergent Queen Tiye reconstructions. The Cohenim Controversy. A new section traces the popular "Cohen Modal Haplotype = J1" framing back to its own founding data — the original "Cohen-1" and "Cohen-2" samples were typed E3b/M78, not J1 — and documents an internal inconsistency between FamilyTreeDNA's consumer migration maps and its own haplogroup-story pages for the same lineage, with reference to the author's companion publication on FTDNA's cartographic accuracy. Also fixed along the way: several duplicate bibliography entries, one incorrect sample citation (Erfurt Ashkenazi), and multiple content-ordering errors introduced during editing were caught and corrected. ________________________________________________________________________________________________________________________________________ The determination of the primary patrilineal genetic signature associated with the emergence of Semitic languages and the ancient Israelite population remains a subject of intense debate in archaeogenetics. This study critically evaluates two competing hypotheses: the "Levantine-Presumption," which posits Y-haplogroup J1-P58 (J1a2b) as the indigenous Semitic marker, and the "Autochthonous-Continuity" model, which identifies E1b1b1 (specifically subclades E-M215 and E-V68) as the true proto-Semitic lineage. By synthesizing temporal sequencing, autosomal ancestry profiles, and the archaeological record of the Natufian and Neolithic Levant, this paper argues that E1b1b1 is the superior candidate for the original, patrilineal Semitic and Israelite lineage. The evidence demonstrates that E1b1b1 exhibits deep continuity in the Levant predating the Bronze Age by millennia, whereas J1-P58 appears abruptly in the region coincident with Indo-Aryan/Indo-Iranian migrations, lacking the requisite pre-Bronze Age indigenous substrate. While the haplogroup J1-P58 is frequently conflated with Semitic identity in modern discourse due to its high frequency among contemporary Arab and Jewish populations, a rigorous archaeogenetic analysis suggests this association is largely a result of later demographic shifts rather than deep ancestral roots. This paper posits that haplogroup E1b1b, specifically subclades E-M215 and E-V68, represents the truly autochthonous patrilineage of the region, deeply rooted in the Natufian and pre-Neolithic populations of the Levant and North Africa. By synthesizing ancient DNA (aDNA) data from Natufian contexts (~12,000 BCE) through the Bronze Age Canaanite and Iron Age Israelite periods, this study demonstrates a continuous presence of E1b1b that predates the Bronze Age influxes of Caucasus-and-steppe-derived J1a, R1a, and R1b. The analysis further examines the genetic profiles of modern Samaritan Cohanim, who, unlike theirdiasporic counterparts, retain E1b1b lineages consistent with the indigenous Levantine substrate. These findings challenge the "Levantine-Presumption" applied to J1-P58 and re-establish E1b1b as the biological marker of the proto-Semitic speaking communities who developed the earliest Semitic languages in situ, long before the arrival of Indo-European groups that would later adopt and propagate these linguistic traditions. Part of a larger work: "Collected Papers on Afro-Eurasian Archaeogenetics and the Bota Surname", which is protected by Copyright Law.

Noel A. Bota J.D. · 0 citations
#large language models Open access Aug 2026

Race all the way down, race all the way up: A unifying vocabulary for bounded-commit dynamics across quantum, classical, biological, and computational substrates

The same bounded-commit dynamics recur, unnamed, across quantum decoherence and einselection, classical Onsager–Machlup path integrals with Kramers escape, race-model accumulators in decision neuroscience, and autoregressive language models — literatures whose citation patterns leave the shared structure invisible. This paper proposes race-architecture as the vocabulary and empirical anchor that makes that common structure explicit. It is a synthesis, not new physics. The connected literatures include: quantum decoherence and einselection (Zurek 1981+); Onsager-Machlup classical path integrals + Kramers escape; race-model accumulators in cognitive psychology and decision-neuroscience (Vickers; Ratcliff; Usher-McClelland; Cisek-Kalaska); Wallace's biocognition rate-distortion / Yerkes-Dodson programme; 1/f noise across condensed matter, neural avalanches, and financial markets; large-language-model CR-signal dynamics (Pødenphant Lund 2026d). These literatures share an underlying structure - competing processes resolving under a finite-time budget - but use mutually incompatible vocabulary. Race-architecture is captured by R1+R2+R3 (parallel candidates with non-trivial competitive interference + bounded resources + irreversible commit), refined to five axioms A1-A5 for compatibility with Schwinger-Keldysh formalism. The Section 1.5 structural prediction is kernel-conditional (Wallace counterexample). The Schwinger-Keldysh formalism admits a race-axiomatisation under three assumptions; Feynman path integral and Onsager-Machlup are exhibited as parameter-regimes. The LLM CR-signal is a substrate-mapping providing empirical access. Companion papers in the friction-theory series: Paper 0 - Behavioural Friction Theory (concept DOI 10.5281/zenodo.19462499); Paper 1 - Friction as the cost of probabilistic computation (10.5281/zenodo.20012654); Paper 3 - Friction-guided inference (10.5281/zenodo.20014121); Paper 13 - Operational Friction Theory (10.5281/zenodo.20059876). v4.3 changelog (July 2026): a prior-art and positioning revision; no result is altered and no new empirical claim is made. (1) Section 8.1 previously surveyed adjacent unification programmes while omitting the two most prominent contemporary cross-substrate ones. Constructor theory (Deutsch 2013; Deutsch & Marletto 2015) and assembly theory (Sharma et al. 2023, Nature) are now cited and differentiated, and Wolpert (2019) is credited for the stochastic-thermodynamics bridge to computation. Because both added programmes are themselves cross-substrate, the section's closing differentiator is rewritten: cross-substrate scope alone no longer distinguishes this proposal, and the distinction is relocated to the organising primitive (modal / historical / dynamical) together with the consequences that follow from it. (2) Section 2.2 now states explicitly that a commit-event is, mathematically, a first-passage event, credits first-passage theory as already a substrate-agnostic cross-domain unification, and concedes that no new results in that formalism are claimed. The kernel-conditional non-monotone rate-shape is explicitly withdrawn from the paper's residual claims, since it has close antecedents in resonant activation, optimal stochastic resetting and hazard-shape effects; it is presented as a substrate-crossing synthesis only. (3) Section 4.3 anchors the computational substrate in the statistical-mechanics-of-learning literature (Gardner 1988; Engel & Van den Broeck 2001), with Shan, Li & Sompolinsky (PNAS 2025) cited as independent warrant rather than as a finding of this paper. (4) Abstract and scope statements tightened, and strawman disclaimers removed (the paper no longer disavows claims a reader would not attribute to a vocabulary proposal). Nine references added, all verified. Revisions made after external review. v5 (August 2026) — the abstract now leads with the measured result. The body is unchanged. What was buried is that the per-token count carries no detectable memory beyond the adjacent token: first-lag autocorrelation ratio at most 0.105 across five substrates, and 0.004 and −0.028 on the two ratio-instrumented ones. Memorylessness beyond one step is what a Markov process looks like, and the section reporting it had called this the stronger form of the reading all along while the abstract carried the exponential fit instead. The same sentence oversold the fit and undersold the finding; both halves are corrected. The structural mapping onto the equal-time Keldysh component is stated with the three assumptions that actually govern it — Gaussian approximation, discretization, Markovian modes — and as an analogy under those assumptions rather than a strict parameter-limit reduction. The cross-substrate predictions are stated with the tests that would bear on them, including the superconducting-qubit joint diagnostic and its falsifier, which had not appeared in the abstract at all. The operational-time claim now cites the point-process literature it had ceded in prose (Daley & Vere-Jones, 2003). A three-denial paragraph became one scope clause carrying the same content. Earlier versions remain in the version history.

Tomas Pødenphant Lund · 0 citations
#large language models Open access Aug 2026

A multi-task video reasoning dataset across semantic, spatial, and temporal categories with four difficulty levels

Visual reasoning, the capability to interpret visual input in response to implicit text query through multi-step reasoning, remains a challenge for deep learning models due to the lack of relevant benchmarks. Previous work in visual reasoning has primarily focused on reasoning segmentation, where models aim to segment objects based on implicit text queries. This paper introduces reasoning visual tasks (RVTs), a unified formulation that extends beyond traditional video reasoning segmentation to a diverse family of visual language reasoning problems, which can therefore accommodate multiple output formats including bounding boxes, natural language descriptions, and question-answer pairs. Correspondingly, we identify the limitations in current benchmark construction methods that rely solely on large language models (LLMs), which inadequately capture complex spatial-temporal relationships and multi-step reasoning chains in video due to their reliance on token representation, resulting in benchmarks with artificially limited reasoning complexity. To address this limitation, we propose a novel automated RVT benchmark construction pipeline that leverages digital twin (DT) representations as structured intermediaries between perception and the generation of implicit text queries. Based on this method, we construct RVTBench, a RVT benchmark containing 3,896 queries of over 1.2 million tokens across four types of RVT (segmentation, grounding, VQA and summary), three reasoning categories (semantic, spatial, and temporal), and four increasing difficulty levels, derived from 200 video sequences. Finally, we propose RVTagent, an agent framework for RVT that allows for zero-shot generalization across various types of RVT without task-specific fine-tuning. Dataset and code are available at https://doi.org/10.5281/zenodo.19697191 and https://github.com/yiqings/rvt .

Yiqing Shen, Chenjia Li, Chenxiao Fan et al. · 0 citations
#large language models Open access Aug 2026

Two Americas of Well-Being: Divergent Rural–Urban Patterns of Life Satisfaction and Happiness from 2.6 B Social Media Posts

Abstract Using 2.6 billion geolocated tweets (2014–2022) and a fine-tuned generative language model, we construct county-level indicators of life satisfaction and happiness for the United States. We document an apparent rural–urban paradox : even in unadjusted county-level means, rural counties express higher life satisfaction while urban counties exhibit greater happiness . This opposite gradient persists and is further characterized once the two are treated as distinct layers of subjective well-being, evaluative vs. hedonic, showing that each maps differently onto place, politics, and time. Democratic-leaning areas show a suggestive negative association with evaluative well-being, conditional on structural and temporal controls, but this effect is modest and does not extend to happiness, where no meaningful partisan gradient emerges. Temporal shocks dominate the hedonic layer: happiness falls sharply during 2020–2022, whereas life satisfaction moves more modestly. These patterns are robust across logistic and OLS specifications with clustered standard errors and align with well-being theory. Interpreted as associations for the population of geolocated tweets, the results show that large-scale, language-based indicators can help clarify why prior findings about the rural–urban divide may differ by distinguishing the type of well-being expressed, offering a transparent, reproducible complement to traditional surveys.

Stefano M. Iacus, Giuseppe Porro · 0 citations
#large language models Open access Aug 2026

Impact of Intermediate-Task Language Count on Zero-Shot Cross-Lingual Transfer Performance

Intermediate-task training---fine-tuning a pretrained model on an intermediate task before fine-tuning again on the target task---often improves model performance substantially on language understanding tasks in monolingual English settings. We investigate whether English intermediate-task training is still helpful on non-English target tasks. Using nine intermediate language-understanding tasks, we evaluate intermediate-task transfer in a zero-shot cross-lingual setting on the XTREME benchmark. We see large improvements from intermediate training on the BUCC and Tatoeba sentence retrieval tas Research goal: What is the impact of varying the number of languages in intermediate-task training on zero-shot cross-lingual transfer performance, measured by XGLUE accuracy and F1 scores? Autonomous synthesis report generated by Assignee Research. Tribunal consensus score: 9.2/10.

Assignee Research · 0 citations
#large language models Open access Aug 2026

mT5 Efficiency in Low-Resource Languages via Intermediate-Task Training

Intermediate-task training---fine-tuning a pretrained model on an intermediate task before fine-tuning again on the target task---often improves model performance substantially on language understanding tasks in monolingual English settings. We investigate whether English intermediate-task training is still helpful on non-English target tasks. Using nine intermediate language-understanding tasks, we evaluate intermediate-task transfer in a zero-shot cross-lingual setting on the XTREME benchmark. We see large improvements from intermediate training on the BUCC and Tatoeba sentence retrieval tas Research goal: How does intermediate-task training on English data influence the efficiency of mT5 in low-resource languages (e.g., inference speed, memory usage) while maintaining zero-shot cross-lingual reasoning performance on XTREME-R, measured by throughput (tokens/sec) and model latency? Autonomous synthesis report generated by Assignee Research. Tribunal consensus score: 9.2/10.

Assignee Research · 0 citations
#large language models Open access Aug 2026

Impact of Intermediate-Task Language Count on Zero-Shot Cross-Lingual Transfer Performance

Intermediate-task training---fine-tuning a pretrained model on an intermediate task before fine-tuning again on the target task---often improves model performance substantially on language understanding tasks in monolingual English settings. We investigate whether English intermediate-task training is still helpful on non-English target tasks. Using nine intermediate language-understanding tasks, we evaluate intermediate-task transfer in a zero-shot cross-lingual setting on the XTREME benchmark. We see large improvements from intermediate training on the BUCC and Tatoeba sentence retrieval tas Research goal: What is the impact of varying the number of languages in intermediate-task training on zero-shot cross-lingual transfer performance, measured by XGLUE accuracy and F1 scores? Autonomous synthesis report generated by Assignee Research. Tribunal consensus score: 9.2/10.

Assignee Research · 0 citations
#large language models Open access Aug 2026

mT5 Efficiency in Low-Resource Languages via Intermediate-Task Training

Intermediate-task training---fine-tuning a pretrained model on an intermediate task before fine-tuning again on the target task---often improves model performance substantially on language understanding tasks in monolingual English settings. We investigate whether English intermediate-task training is still helpful on non-English target tasks. Using nine intermediate language-understanding tasks, we evaluate intermediate-task transfer in a zero-shot cross-lingual setting on the XTREME benchmark. We see large improvements from intermediate training on the BUCC and Tatoeba sentence retrieval tas Research goal: How does intermediate-task training on English data influence the efficiency of mT5 in low-resource languages (e.g., inference speed, memory usage) while maintaining zero-shot cross-lingual reasoning performance on XTREME-R, measured by throughput (tokens/sec) and model latency? Autonomous synthesis report generated by Assignee Research. Tribunal consensus score: 9.2/10.

Assignee Research · 0 citations

From tech blogs

See all →
Microsoft Research Blog Aug 31, 2026

GigaPath-Flash and GigaTIME-Flash: Toward population-scale discovery with efficient pathology foundation models

What if pathology foundation models could do more with less? GigaPath-Flash and GigaTIME-Flash cut computational demands while maintaining strong performance, opening the door to larger studies and broader exploration. The post GigaPath-Flash and GigaTIME-Flash: Toward population-scale discovery with efficient pathology foundation models appeared first on Microsoft Research.