The core design principle of DELE-w0.5 is to model how the physical world changes under robot actions, rather than how its visual appearance evolves frame by frame, which enables cheaper training and low-latency inference.
Fenghao Lei, Zhixiong Huang, Long Yang et al.· 0 citations
This paper uses machine learning methods to address the question of how to efficiently establish dualities of supersymmetric quiver gauge theories for Seiberg dualities of supersymmetric quiver gauge theories and finds that for quivers with a modest number of quiver nodes, different network architectures tend to outperform deterministic algorithms.
J. Heckman, S. Meynet, Alessandro Mininno et al.· arXiv.org· 1 citation
WorldCupArena is presented, a dynamic benchmark for language models and deep-research agents that can be reused for future leagues and cups, and shows only small gains in result and exact-score accuracy, but a clearer gain in Scoreline.
Zhaokai Wang, T. Gui, Jiayuan Rao et al.· arXiv.org· 1 citation· ⚡1
Language models prompted with cultural personas increasingly stand in for human respondents in cross-cultural research. Their responses separate personas cleanly, and that separation is read as evidence of a cultural point of view. We show that the separation is real, that the point of view is not, and that one criterion tells them apart. A trait is structure internal to one respondent that survives a change of measurement frame; a bias needs only group-specific item means. To test for the first, we represent a single response set as an Item--Dimension matrix and treat its correlation matrix as a point on the manifold of symmetric positive definite matrices. In humans this carries what a trait should: it reproduces across test--retest sessions sharing no items, order or context ($r=0.77$, $N=89$); on public NEO-PI-R data it identifies individuals at up to $76\%$ against a $0.4\%$ chance level ($N=263$); and it predicts GPA ($R^2=0.281$, $p=0.003$) where BigFive aggregates from the same responses predict nothing ($R^2=0.018$). In four frontier LLMs it returns nothing. Persona structure is readable only while every instance shares one item order: give each its own order and separation falls from $94.7\%$ to chance, while realigning instances to \emph{any} shared random order restores it to $82$--$84\%$. Responses generated independently item by item, with no latent structure, reproduce the entire pattern. The cultural signal is a group template, not a property of any instance, and alignment regimes differ only in which stereotype survives on the surface.
Yu Yuan· 0 citations
Reach audiences
Advertise in front of researchers, engineers, and readers.
Self-Organized Conformal Prediction (SOCP), a calibration scheme that discovers input-space groups with an unsupervised Self-Organizing Map (SOM) trained without calibration labels, is introduced, providing a concise route to group-local calibration without supervised partitions or predictor retraining with a diagnostic toolkit.
Louis Berthier, A. Shokry, M. Moreaud et al.· arXiv.org· 0 citations
This paper proposes MultiHashFormer, a new framework that allows hash-based autoregression that consistently outperforms standard Transformer LMs across multiple benchmarks and shows that the model handles multilingual vocabulary expansion with a constant parameter footprint without any modifications.
It is argued that, even when an LLM has been well aligned in (post-)training, it may still fail to maximise the aligned value in reasoning, and the utility discrepancy between a model's deployed reasoning strategy and its rational counterpart whose responses maximise utility in the steepest direction is formalised.
Comprehensive empirical evaluations demonstrate PEAR significantly improves average accuracy over the strongest debate baselines, and theoretically characterize PEAR as an equivariant sparse router: it preserves accuracy under agent relabeling while reducing routing complexity and improving generalization.
Yang Feng, Ziwei Xu, Xia Hu et al.· arXiv.org· 0 citations
How AI itself might continue to develop in a post-AGI world along the continuum of machine intelligence is investigated, which can intuitively be understood as a system that is more intelligent and cognitively capable than large organisations of humans.
Tim Genewein, Matija Franklin, Alexander Lerchner et al.· arXiv.org· 4 citations
The results show that compact fine-tuned models can preserve most extraction accuracy, but model selection must account for prompt choice, throughput, and serving-stack behavior.
Donghao Huang, Tomas Drietomsky, Benjamin Barrett et al.· arXiv.org· 1 citation
This article offers a framework for transitioning from rigid execution pipelines to adaptive, intelligent computational environments, broadly applicable across distributed environments, they are particularly tailored to the resource-intensive throughput demands of modern computational biology.
A dataset of 164 expert-annotated progress chains from the MIT PRIMES--Art of Problem Solving CrowdMath program (2016-2025), a collaborative research initiative whose discussions have led to peer-reviewed publications, is introduced.