Skip to content

Category

artificial intelligence

4,637 papers

#artificial intelligence Preprint Dec 2024

Learning Personalized Prompts for Healthcare Guidance

The results show that the proposed personalized prompt learning (PPL) approach produces more personalized healthcare guidance and wins 97 out of 100 comparisons in expert evaluation, demonstrating its potential for broader healthcare applications.

Ruize Shi, Hong Huang, Wei Zhou et al. · 6 citations
#artificial intelligence Conference Open access Nov 2023

General Phrase Debiaser: Debiasing Masked Language Models at a Multi-Token Level

An automatic multi-token debiasing pipeline called General Phrase Debiaser, which is capable of mitigating phrase-level biases in masked language models, and can significantly reduce gender biases on both career and multiple disciplines, across models with varying parameter sizes.

Bingkang Shi, Xiaodan Zhang, Dehan Kong et al. · 4 citations
#artificial intelligence Preprint Aug 2026

BLOOM-WILT: Logit Tilting for Behaviour Elicitation in Automated LLM Auditing

BLOOM-WILT is introduced, a full auditing pipeline that elicits natural multi-turn instances of rare behaviours, without training cost or access beyond the target's next-token distribution and raises average behaviour presence from 51% to 100% when eliciting self-harm encouragement from Qwen3.5-4B.

Adrians Skapars, Edoardo Manino · 0 citations
#artificial intelligence Preprint Aug 2026

Agentic Context Cracking: Token-Efficient Data Reasoning Agents via Adaptive Structuring of Unstructured Data

This work proposes agentic data cracking, a method that structures unstructured data adaptively and speculatively as a byproduct of reasoning itself, a first step toward next-generation data infrastructure for agentic reasoning over unstructured data.

Milad Rezaei Hajidehi, Qitong Wang, Stratos Idreos · 0 citations
#artificial intelligence Review Aug 2026

Responsible Integration of AI in Cancer Genomics: Barriers, Risks, and Pathways to Trustworthy Clinical Translation

A conceptual framework and roadmap for addressing four interrelated translational failure domains through rigorous validation, uncertainty-aware methods, interoperable infrastructures, regulatory alignment, and human oversight across the AI lifecycle is proposed.

B. Ilgen, Yiannos S. Tolias, Denise Kühnert et al. · 0 citations
#artificial intelligence Preprint Aug 2026

HSRM: Hidden-State Reward Models for Test-Time Verification

HSRM is introduced, a lightweight hidden-state reward model that verifies candidate solutions by directly reading the generator's internal representations rather than re-processing its text, providing an efficient alternative to text-only verification by reusing representations already computed during generation.

Xianzhi Li, Xiao-Dan Zhu · 0 citations
#artificial intelligence Preprint Aug 2026

Geometry of Divergence: Tracking Hidden-State Trajectories for Adaptive Multi-Turn Reasoning

This work formulate multi-turn reasoning as a hidden-state trajectory of the underlying LLM that is characterized via two complementary signals: temporal curvature that captures the directional consistency of turn-to-turn updates, and variance slope which measures the expansion or contraction of the exploration space.

Jie Liang, Zhengxin Yu, H. Nasiri et al. · 0 citations
#artificial intelligence Preprint Aug 2026

ScienceArena: Benchmarking LLMs on Latest Scientific Olympiad Competitions

Benchmark saturation and data contamination increasingly obscure genuine scientific reasoning in frontier LLMs. We introduce \textsc{ScienceArena}, an olympiad-style benchmark from thirteen public science competitions in physics, chemistry, and biology, including IPhO and IChO 2025--2026, IBO 2023, USAPhO 2026, and USNCO 2025. Its open-ended, multi-step problems use process-credit rubrics, making faithful scoring difficult. We build ScienceArena through an expert-audited digitization pipeline that converts official exams, figures, solutions, and rubrics into structured items verified by olympiad medalists. To scale evaluation beyond costly human grading, we calibrate LLM-as-judge against medalist ground truth on archived answers from five models across IPhO and IChO; two strong judges stay within one point of expert total scores. Medalist notes show that failures often stem from visual grounding, structure fidelity, and global problem control rather than missing terminology. Evaluating fourteen recent LLMs with interleaved solving, we find that top models obtain medal-equivalent rubric scores on several public international exams, while chemistry and long-horizon consistency remain key bottlenecks. We provide an interactive \href{https://science-arena.onrender.com/}{demo}.

Guangxiang Zhao, Qi-Long Shi, Xusen Xiao et al. · 0 citations
#artificial intelligence Preprint Aug 2026

Lot Machine: Multimodal Lot Extraction from Auction Catalogs

This work demonstrates that a VLM-based pipeline can successfully unlock historical auction catalogs for large-scale automated analysis, and benchmark the methods across different deployment modes ranging from commercial providers to locally hosted, quantized models.

Mathias Zinnen, Alisha Mund, Sabine Lang et al. · 0 citations
#artificial intelligence Preprint Aug 2026

EvoSkill Injection: Red-Teaming Autonomous Skill Generation and Evolution in Self-Evolving Agents

A red-teaming framework for evaluating this threat model targeting the autonomous skill generation and evolution pipeline of self-evolving agents and shows that SARGE induces malicious skill formation and that injected skills are persistently stored and repeatedly activated, highlighting the risk of persistent capability corruption.

Doyun Kim, Chanwoo Kim, Sugyeong Eo et al. · 0 citations
#artificial intelligence Preprint Aug 2026

Ignorance or Incompetence? Constructing Knowledge-Gated, Verifiable Tasks for LLM Agents

A knowledge-gated task-construction protocol is introduced that separates a task instruction from a compact artefact containing private conventions, reference tables, and utility operators, and it is shown that the retained tasks improve post-training.

Han-Lin Tian, Min-Hao Li, Yuhan Mi et al. · 0 citations

From tech blogs

See all →

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.