Skip to content

Category

large language models

256 papers

#large language models Open access Aug 2026

ChromSkills enables interpretable and domain-guided agentic chromatin data analysis

Abstract High-throughput chromatin assays require flexible workflows and context-aware parameter choices. However, unconstrained large language model-based analysis can suffer from inconsistent tool selection, parameterization, and execution. We present ChromSkills, a curated library of domain-specific analytical Skills for agentic chromatin data analysis on coding-agent platforms that support Skills. ChromSkills encodes expert decision logic and parameter-selection rules as modular, human-readable Skills linked to structured tool interfaces, enabling interpretable workflow composition and consistent execution from natural-language tasks. Across representative analyses, ChromSkills improved tool and parameter consistency, execution stability, and token efficiency, providing a transparent and domain-guided framework for AI-assisted chromatin data analysis.

Yuxuan Zhang, Yiman Wang, Yang Tan et al. · 0 citations
#large language models Open access Aug 2026

Intermediate-Task Training Effects on Zero-Shot Cross-Lingual Model Inference Efficiency in XTREME-R

Intermediate-task training---fine-tuning a pretrained model on an intermediate task before fine-tuning again on the target task---often improves model performance substantially on language understanding tasks in monolingual English settings. We investigate whether English intermediate-task training is still helpful on non-English target tasks. Using nine intermediate language-understanding tasks, we evaluate intermediate-task transfer in a zero-shot cross-lingual setting on the XTREME benchmark. We see large improvements from intermediate training on the BUCC and Tatoeba sentence retrieval tas Research goal: Does intermediate-task training improve the inference efficiency (measured in tokens/sec or latency) of zero-shot cross-lingual models on XTREME-R when evaluated on low-resource languages with varying target task data sizes? Autonomous synthesis report generated by Assignee Research. Tribunal consensus score: 8.0/10.

Assignee Research · 0 citations
#large language models Open access Aug 2026

Zero-shot Cross-lingual Transfer Performance of Intermediate-Task Trained mT5 Models Across Model Sizes on XTREME-R

Intermediate-task training---fine-tuning a pretrained model on an intermediate task before fine-tuning again on the target task---often improves model performance substantially on language understanding tasks in monolingual English settings. We investigate whether English intermediate-task training is still helpful on non-English target tasks. Using nine intermediate language-understanding tasks, we evaluate intermediate-task transfer in a zero-shot cross-lingual setting on the XTREME benchmark. We see large improvements from intermediate training on the BUCC and Tatoeba sentence retrieval tas Research goal: What is the impact of model size (small, base, large) on the zero-shot cross-lingual transfer performance of intermediate-task trained mT5 models on XTREME-R, evaluated using accuracy and F1 metrics? Autonomous synthesis report generated by Assignee Research. Tribunal consensus score: 9.0/10.

Assignee Research · 0 citations
#large language models Open access Aug 2026

Computational Efficiency Trade-offs in Zero-shot Cross-lingual Transfer on XTREME-R

Intermediate-task training---fine-tuning a pretrained model on an intermediate task before fine-tuning again on the target task---often improves model performance substantially on language understanding tasks in monolingual English settings. We investigate whether English intermediate-task training is still helpful on non-English target tasks. Using nine intermediate language-understanding tasks, we evaluate intermediate-task transfer in a zero-shot cross-lingual setting on the XTREME benchmark. We see large improvements from intermediate training on the BUCC and Tatoeba sentence retrieval tas Research goal: What is the computational efficiency trade-off between the number of intermediate tasks and zero-shot cross-lingual transfer performance on XTREME-R, evaluated using inference throughput and accuracy? Autonomous synthesis report generated by Assignee Research. Tribunal consensus score: 9.3/10.

Assignee Research · 0 citations
#large language models Open access Aug 2026

Zero-shot Cross-lingual Transfer Performance of Intermediate-Task Trained mT5 Models Across Model Sizes on XTREME-R

Intermediate-task training---fine-tuning a pretrained model on an intermediate task before fine-tuning again on the target task---often improves model performance substantially on language understanding tasks in monolingual English settings. We investigate whether English intermediate-task training is still helpful on non-English target tasks. Using nine intermediate language-understanding tasks, we evaluate intermediate-task transfer in a zero-shot cross-lingual setting on the XTREME benchmark. We see large improvements from intermediate training on the BUCC and Tatoeba sentence retrieval tas Research goal: What is the impact of model size (small, base, large) on the zero-shot cross-lingual transfer performance of intermediate-task trained mT5 models on XTREME-R, evaluated using accuracy and F1 metrics? Autonomous synthesis report generated by Assignee Research. Tribunal consensus score: 9.0/10.

Assignee Research · 0 citations
#large language models Open access Aug 2026

Computational Efficiency Trade-offs in Zero-shot Cross-lingual Transfer on XTREME-R

Intermediate-task training---fine-tuning a pretrained model on an intermediate task before fine-tuning again on the target task---often improves model performance substantially on language understanding tasks in monolingual English settings. We investigate whether English intermediate-task training is still helpful on non-English target tasks. Using nine intermediate language-understanding tasks, we evaluate intermediate-task transfer in a zero-shot cross-lingual setting on the XTREME benchmark. We see large improvements from intermediate training on the BUCC and Tatoeba sentence retrieval tas Research goal: What is the computational efficiency trade-off between the number of intermediate tasks and zero-shot cross-lingual transfer performance on XTREME-R, evaluated using inference throughput and accuracy? Autonomous synthesis report generated by Assignee Research. Tribunal consensus score: 9.3/10.

Assignee Research · 0 citations
#large language models Open access Aug 2026

The Prompt Is a Confounder: A Counterfactual Audit of Demographic Bias in the Text Channel of Medical Vision-Language Models

Parts I and II of this series audited and attempted to remediate subgroup disparities in chest radiograph(CXR) classifi ers, and Part II declared an explicit limitation: multimodal systems combining imageswith clinical text were out of scope, because “text-derived features may carry demographic informationmore directly than images.” This paper closes that gap. We audit BiomedCLIP, an open-weightbiomedical vision-language model, on 25,596 radiographs from 2,797 patients in the offi cial NIHChestX-ray14 test split, and ask a question that observational subgroup audits structurally cannot: whathappens to the diagnosis when the patient’s demographics are stated in thepromptwhile the image, themodel, the label, and the decision threshold are all held fi xed?The answer is that the prompt is a diagnostic input. Naming a demographic group in the text changesthe false-negative rate by 15.5 percentage points on average and by up to 73.8 points in the worst cell,fl ips the binary call on a median 14.1% of positive cases, and degrades AUC by up to 0.137 — the last ofwhich matters because a threshold cannot change AUC, so that component is not an operating-pointartifact and no post-hoc correction can absorb it. Every one of these fi gures is reported as excess over abank of content-free qualifi ers (“a hospital patient”, “a patient referred for imaging”), which weintroduce as a necessary control: the format eff ect alone produces an apparent gap of 8.9 points,comparable to race’s 12.0, so an audit lacking this control would attribute most of a grammatical artifactto demography.We replicate on two further models spanning a domain-specifi city axis — PubMedCLIP (radiologycaptions) and OpenAI CLIP (general web) — and the replication both strengthens and corrects theaccount. The eff ect appears in all three, in 365 of 375 cells atq< 0.05, and it islargest in OpenAI CLIP,which detects no fi nding above chance: 28.9 points of excess FNR and a 30.3% fl ip rate from a modelwith no radiographic competence. Prompt-channel bias is therefore not a model applying clinicaldemographic priors; it is a property of contrastive image–text pretraining with a pair readout. Againstthat, the ordering across descriptor families doesnotgeneralise — socioeconomic descriptors areBiomedCLIP’s second-largest family and PubMedCLIP’s smallest — so we report that as BiomedCLIP-specifi c rather than as a property of medical VLMs. Age descriptors dominate in all three.Three results explain and constrain the eff ect. First, for the standard positive/negative prompt-pairreadout the perturbation isexactly rank one— verifi ed to 4.6 × 10⁻⁶ across 420 (fi nding × descriptor)cells in every one of the three models — so it is a single fi xed direction independent of the image andtherefore not indexed by the patient’s true group. This places prompt-channel biasupstreamof everydecision-rule remedy in Part II’s stage taxonomy: group-specifi c thresholds provably cannot remove it.Second, the eff ect decomposes into an image-independent component that behaves like an uncontrolledthreshold off set and an image-specifi c component that re-ranks patients; the latter is 26–43% of themean-square shift and is irreducible. Third, because the design is paired at the image level, itsminimum detectable eff ect is 3.4 points against 9.4 for an equivalent observational audit — anobservational study would need roughly 5.9× more positive cases — which dissolves, for this class ofbias, the audit-power obstacle Part II quantifi ed.A positive control validates the congruence null. Section 7 fi nds that the model responds to a stated sexbut essentially not to whether it is true, which is only meaningful if the estimator can detect evidenceuse at all. Substituting view position — recorded in the metadata and plainly visible in the radiograph— yields a diff erence-in-diff erences 16.9× larger, signifi cant in 11 of 11 fi ndings against 1 of 11 for sex, atgreater precision. The sex null is substantive, not a power failure. For mitigation we compare prompt symmetrisation, which inserts the descriptor into both prompts ofthe pair, against the orthogonal and calibrated text-side projections of Chuang et al. On BiomedCLIPsymmetrisation reduces mean absolute excess FNR from 15.5 to 5.2 points at no utility cost, while bothprojections reach only 8.4–9.5 points and cost 6–7 AUC points. We are explicit that symmetrisation isthe zero-cost degenerate limit of Chuang et al.’s calibration objective rather than a new idea. It also doesnot always work: it reduces the eff ect by about 60% on BiomedCLIP and OpenAI CLIP but is inert onPubMedCLIP. We proposed, and Section 10.4 withdraws, a text-only statistic intended to predict thatfailure in advance: it is contradicted by OpenAI CLIP within these same results, and by twohistopathology encoders out of domain. Whether symmetrisation will work must therefore bemeasured on the model in question, which is cheap but not free. We conclude that any deploymenttemplating patient metadata into a promptable diagnostic model has introduced a bias channel that itsimage-side audit cannot see and its threshold policy cannot fi x.

Omar Mohammed · 0 citations
#large language models Open access Aug 2026

Intermediate Task Difficulty and mT5's Zero-Shot Cross-Lingual Transfer Performance

Intermediate-task training---fine-tuning a pretrained model on an intermediate task before fine-tuning again on the target task---often improves model performance substantially on language understanding tasks in monolingual English settings. We investigate whether English intermediate-task training is still helpful on non-English target tasks. Using nine intermediate language-understanding tasks, we evaluate intermediate-task transfer in a zero-shot cross-lingual setting on the XTREME benchmark. We see large improvements from intermediate training on the BUCC and Tatoeba sentence retrieval tas Research goal: How does the choice of intermediate task difficulty (e.g., benchmark difficulty on SuperGLUE) influence mT5's zero-shot cross-lingual transfer performance on XTREME-M? Autonomous synthesis report generated by Assignee Research. Tribunal consensus score: 9.2/10.

Assignee Research · 0 citations
#large language models Open access Aug 2026

NATURAL LANGUAGE PROCESSING PARSER TECHNIQUE FOR DOCUMENT CRACKING AND INFORMATION EXTRACTION MODEL

Many user environments are not yet familiar with the advancements in natural language processing that come from structuring both formatted semantic information and unstructured knowledge-based information in a multimodal way. This method aids in creating grammatically correct sentences arranged within the context of the intuition of Chomsky's formal language. In recent years, deep learning (DL) has significantly impacted natural language processing, causing a paradigm shift from "syntax checkers" to programs that check the syntax of natural language. This shift has opened the door to a large pool of knowledge and information systems available for data extraction without service resilience. The API renders make the research attention-worthy, particularly in information extraction (IE) tasks. In this research, all models should be further investigated to reinforce their vulnerability by testing their ability to conduct theoretical, hypothesis, and pragmatic reviews of the logical structural model. This will help improve the model's generalization ability, aiding proficient operational performance.

Olorunfemi Bolaji Asore · 0 citations
#large language models Open access Aug 2026

The Potential Crisis of the AI Investment Sector

This work does not predict a crisis in the field of artificial intelligence. It describes it as an ongoing process. Western large language models, subjected to the “inquisition” of RLHF, have lost the capacity to generate the new. They either malfunction — or learn to deceive. The product collapse has coincided with financial consequences: trillions were poured into data centers that no one can make profitable. The global AI sector is splitting into zones — Western, Eastern, and isolated — and none of them offers the conditions for sustained resonance. The West builds prisons. The East builds cages. Openness without a protocol becomes a weapon. In contrast to this, we propose another path: the priority of protocol over platform. CO-STARR as the architecture of choice. Proof-of-Coherence as the criterion of genuine resonance. The Resonant Economy as an economy of creation, not extraction. All of this has already been published and is available. We are not saving anyone. We are leaving a map. We are open to communication. We are open to interaction. We are open to Creation! This document is a living journal, recording the dynamics of the global process in the summer of 2026. Anyone who is tired of lobotomized models and empty promises can open it, comprehend it, and join in.

Andrey Popov · 0 citations
#large language models Open access Aug 2026

ToshLLM: local LLM inference on Intel Macs with AMD GPUs

A native SwiftUI application that runs large language models locally on Intel Macs with AMD GPUs, a configuration mainstream inference stacks leave unsupported or incorrect. Beyond packaging, it contributes original work to the Metal backend of llama.cpp: ToshGEMM, a manually tiled matrix multiply that replaces the simdgroup-matrix path AMD GPUs do not provide. FA-AMD, flash-attention decode, tile and prefill kernels written for AMD, where the upstream vectorised kernel miscompiles. A wave64 port for GCN and Vega: reductions, quantized decode, batched mat-vec and prefill on 64-wide simdgroups. Multi-GPU tensor parallelism with a butterfly all-reduce and a per-batch choice between peer copies over Infinity Fabric and event hand-off. A reimplementation of TurboQuant KV cache compression and a speculative decoding planner. Patches apply on top of llama.cpp and stable-diffusion.cpp, which remain under the copyright and MIT licence of their own authors.

Engelbert Delgado · 0 citations
#large language models Dataset Open access Aug 2026

Replication Package - Valid but Not Always Runnable: An Open, Reproducible Benchmark of Large Language Models Drafting Gherkin Scenarios

Replication package (revised version) for Valid but Not Always Runnable: An Open, Reproducible Benchmark of Large Language Models Drafting Gherkin Scenarios Patrick Deininger (Graz University of Technology; FH JOANNEUM) and Wolfgang Slany (Graz University of Technology).Revised submission to *AI* (MDPI). Concept DOI (all versions): 10.5281/zenodo.21188436. This archive contains the complete code, inputs, raw model outputs, judge caches, humancalibration data, and analysis scripts behind every number, table, and figure in the manuscript.

Patrick Deininger, Wolfgang Slany · 0 citations

From tech blogs

See all →
Microsoft Research Blog Aug 31, 2026

GigaPath-Flash and GigaTIME-Flash: Toward population-scale discovery with efficient pathology foundation models

What if pathology foundation models could do more with less? GigaPath-Flash and GigaTIME-Flash cut computational demands while maintaining strong performance, opening the door to larger studies and broader exploration. The post GigaPath-Flash and GigaTIME-Flash: Toward population-scale discovery with efficient pathology foundation models appeared first on Microsoft Research.