Skip to content

Category

artificial intelligence

6,274 papers

#artificial intelligence Preprint Aug 2026

ScienceArena: Benchmarking LLMs on Latest Scientific Olympiad Competitions

Benchmark saturation and data contamination increasingly obscure genuine scientific reasoning in frontier LLMs. We introduce \textsc{ScienceArena}, an olympiad-style benchmark from thirteen public science competitions in physics, chemistry, and biology, including IPhO and IChO 2025--2026, IBO 2023, USAPhO 2026, and USNCO 2025. Its open-ended, multi-step problems use process-credit rubrics, making faithful scoring difficult. We build ScienceArena through an expert-audited digitization pipeline that converts official exams, figures, solutions, and rubrics into structured items verified by olympiad medalists. To scale evaluation beyond costly human grading, we calibrate LLM-as-judge against medalist ground truth on archived answers from five models across IPhO and IChO; two strong judges stay within one point of expert total scores. Medalist notes show that failures often stem from visual grounding, structure fidelity, and global problem control rather than missing terminology. Evaluating fourteen recent LLMs with interleaved solving, we find that top models obtain medal-equivalent rubric scores on several public international exams, while chemistry and long-horizon consistency remain key bottlenecks. We provide an interactive \href{https://science-arena.onrender.com/}{demo}.

Guangxiang Zhao, Qi-Long Shi, Xusen Xiao et al. · 0 citations
#artificial intelligence Preprint Aug 2026

Lot Machine: Multimodal Lot Extraction from Auction Catalogs

This work demonstrates that a VLM-based pipeline can successfully unlock historical auction catalogs for large-scale automated analysis, and benchmark the methods across different deployment modes ranging from commercial providers to locally hosted, quantized models.

Mathias Zinnen, Alisha Mund, Sabine Lang et al. · 0 citations
#artificial intelligence Preprint Aug 2026

EvoSkill Injection: Red-Teaming Autonomous Skill Generation and Evolution in Self-Evolving Agents

A red-teaming framework for evaluating this threat model targeting the autonomous skill generation and evolution pipeline of self-evolving agents and shows that SARGE induces malicious skill formation and that injected skills are persistently stored and repeatedly activated, highlighting the risk of persistent capability corruption.

Doyun Kim, Chanwoo Kim, Sugyeong Eo et al. · 0 citations
#artificial intelligence Preprint Aug 2026

Ignorance or Incompetence? Constructing Knowledge-Gated, Verifiable Tasks for LLM Agents

A knowledge-gated task-construction protocol is introduced that separates a task instruction from a compact artefact containing private conventions, reference tables, and utility operators, and it is shown that the retained tasks improve post-training.

Han-Lin Tian, Min-Hao Li, Yuhan Mi et al. · 0 citations
#artificial intelligence Preprint Aug 2026

LaMoC: Loss-Aware Modular Compression for LLMs

LaMoC improves joint compression by selecting compression statistics that better align local module reconstruction error with the downstream loss, and reformulate joint modular compression as a two-tiered optimization problem that minimizes module reconstruction error while tuning the activation and gradient information blending rate.

Mohanad Odema, Jacob Song · 0 citations
#artificial intelligence Preprint Aug 2026

A.X K2 Technical Report

To support long contexts efficiently, Sparse Gated Attention (SGA), which combines sparse attention with gated attention, and adopt Gated Norm (GN) to stabilize large-scale training is introduced, which keeps 4-bit NVFP4 serving within one point of FP8 accuracy.

Cheolseung Baek, Dhammiko Arya, Eunki Kim et al. · 0 citations
#artificial intelligence Preprint Aug 2026

Conducting Stylistic Analysis of Paintings through an Art-History Agent

This approach converts detailed visual features into descriptive terms, addressing a key challenge in art history, and connects the use of images as data with the semantic concerns of humanists, establishing vision-based computational art history as an area for future growth.

M. Walton, Astrid Harth · 0 citations
#artificial intelligence Review Aug 2026

Call Neighbours Yourself: Graph Walks with Destination-Conditioned On-Policy Self-Distillation

This work proposes Call Neighbours Yourself (CNY), a framework that enables LLMs to proactively explore graph neighbourhoods through topology-constrained graph-walk actions and introduces destination-conditioned on-policy self-distillation, which retrospectively evaluates a selected neighbour after its content is revealed and converts the resulting change in action preference into an action-level training signal.

Yilun Liu, Bo-Yu Luo, Yanran Tang et al. · 0 citations
#artificial intelligence Preprint Aug 2026

BIRD-History: A Benchmark for History-Driven Text-to-SQL with Fine-Grained Knowledge Annotations

BIRD-History is introduced, a benchmark consisting of 1,393 tasks across 11 databases, designed to evaluate text-to-SQL systems'ability to ground underspecified natural language questions using historical SQL scripts, and a plug-in retriever that extracts five types of external knowledge from historical SQL scripts, then retrieves and reranks relevant fragments for query generation.

Yunfan Zhou, Qiming Shi, Yi-Zhou Yang et al. · 0 citations
#artificial intelligence Preprint Aug 2026

Super Library Agent: Joint Generation and Maintenance of Multiple Applications Beyond the Single Codebase

This work introduces the Super Library Agent problem, where an agent sequentially generates a portfolio of N related applications while maintaining a shared Super Library of reusable cross-application components, and addresses candidate-guided extraction over code chunk summaries, pre-extraction codebase consolidation, and context-aware migration using extraction traces and call-graph information.

Daegyu Sung, Yukyeong Lee, Geon Park et al. · 0 citations
#artificial intelligence Preprint Aug 2026

How Identity and Opinion Shape Political Sycophancy in LLMs

A framework that disentangles two distinct triggers of political sycophancy: opinion (aligning with explicit narratives) and identity (stereotyping based on demographic labels) is introduced, highlighting how personalization may amplify identity- or opinion-conditioned shifts in the model's behaviors.

Li-Ni Fu, Chang-Chih Meng, Chien-Hua Chen et al. · 0 citations

From tech blogs

See all →

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.