Skip to content

Category

artificial intelligence

4,654 papers

Learning to Trace Seiberg Dualities

This paper uses machine learning methods to address the question of how to efficiently establish dualities of supersymmetric quiver gauge theories for Seiberg dualities of supersymmetric quiver gauge theories and finds that for quivers with a modest number of quiver nodes, different network architectures tend to outperform deterministic algorithms.

J. Heckman, S. Meynet, Alessandro Mininno et al. · 1 citation

WorldCupArena: Fine-Grained Evaluation of Language Models and Deep-Research Agents on Football Forecasting

WorldCupArena is presented, a dynamic benchmark for language models and deep-research agents that can be reused for future leagues and cups, and shows only small gains in result and exact-score accuracy, but a clearer gain in Scoreline.

Zhaokai Wang, T. Gui, Jiayuan Rao et al. · 1 citation · ⚡1
#artificial intelligence Preprint Jul 2026

Cultural Bias Without a Cultural Self:A Disassociation Study of LLM's Persona and Bias

Language models prompted with cultural personas increasingly stand in for human respondents in cross-cultural research. Their responses separate personas cleanly, and that separation is read as evidence of a cultural point of view. We show that the separation is real, that the point of view is not, and that one criterion tells them apart. A trait is structure internal to one respondent that survives a change of measurement frame; a bias needs only group-specific item means. To test for the first, we represent a single response set as an Item--Dimension matrix and treat its correlation matrix as a point on the manifold of symmetric positive definite matrices. In humans this carries what a trait should: it reproduces across test--retest sessions sharing no items, order or context ($r=0.77$, $N=89$); on public NEO-PI-R data it identifies individuals at up to $76\%$ against a $0.4\%$ chance level ($N=263$); and it predicts GPA ($R^2=0.281$, $p=0.003$) where BigFive aggregates from the same responses predict nothing ($R^2=0.018$). In four frontier LLMs it returns nothing. Persona structure is readable only while every instance shares one item order: give each its own order and separation falls from $94.7\%$ to chance, while realigning instances to \emph{any} shared random order restores it to $82$--$84\%$. Responses generated independently item by item, with no latent structure, reproduce the entire pattern. The cultural signal is a group template, not a property of any instance, and alignment regimes differ only in which stereotype survives on the surface.

Yu Yuan · 0 citations

Self-Organized Conformal Prediction: Reducing Regional Coverage Gaps with Unsupervised Group Discovery

Self-Organized Conformal Prediction (SOCP), a calibration scheme that discovers input-space groups with an unsupervised Self-Organizing Map (SOM) trained without calibration labels, is introduced, providing a concise route to group-local calibration without supervised partitions or predictor retraining with a diagnostic toolkit.

Louis Berthier, A. Shokry, M. Moreaud et al. · 0 citations

MultiHashFormer: Hash-based Generative Language Models

This paper proposes MultiHashFormer, a new framework that allows hash-based autoregression that consistently outperforms standard Transformer LMs across multiple benchmarks and shows that the model handles multilingual vocabulary expansion with a constant parameter footprint without any modifications.

Hui Xue, Atsuki Yamaguchi, Nikolaos Aletras · 0 citations

In LLM Reasoning, there is Irrationality on top of Value Misalignment

It is argued that, even when an LLM has been well aligned in (post-)training, it may still fail to maximise the aligned value in reasoning, and the utility discrepancy between a model's deployed reasoning strategy and its rational counterpart whose responses maximise utility in the steepest direction is formalised.

Kejiang Qian, Feng-Xiang He · 0 citations

PEAR: Permutation-Equivariant Adaptive Routing Multi-Agent Debate

Comprehensive empirical evaluations demonstrate PEAR significantly improves average accuracy over the strongest debate baselines, and theoretically characterize PEAR as an equivariant sparse router: it preserves accuracy under agent relabeling while reducing routing complexity and improving generalization.

Yang Feng, Ziwei Xu, Xia Hu et al. · 0 citations

From AGI to ASI

How AI itself might continue to develop in a post-AGI world along the continuum of machine intelligence is investigated, which can intuitively be understood as a system that is more intelligent and cognitively capable than large organisations of humans.

Tim Genewein, Matija Franklin, Alexander Lerchner et al. · 4 citations

Twelve quick tips for designing AI-driven HPC workflows

This article offers a framework for transitioning from rigid execution pipelines to adaptive, intelligent computational environments, broadly applicable across distributed environments, they are particularly tailored to the resource-intensive throughput demands of modern computational biology.

J. Alnasir · 0 citations
#artificial intelligence Review Jun 2026

CrowdMath: A Dataset of Crowdsourced Mathematical Research Discussions

A dataset of 164 expert-annotated progress chains from the MIT PRIMES--Art of Problem Solving CrowdMath program (2016-2025), a collaborative research initiative whose discussions have led to peer-reviewed publications, is introduced.

Sherin Muckatira, Jesse Geneson, Slava Gerovitch et al. · 0 citations

From tech blogs

See all →

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.