Retrospective Harness Optimization is introduced, a self-supervised method that optimizes the agent harness using only past trajectories and alters the agent's behavior patterns and sustains higher accuracy during long-horizon sessions.
Wenbo Pan, Shujie Liu, Chin-Yew Lin et al.· arXiv.org· 8 citations· ⚡1
DASH is introduced, which supervises the conditional and unconditional branches independently and an anchor term regularises the conditional prediction toward ground-truth noise, and the teacher's final learned per-timestep curriculum transfers into the student as a frozen prior.
A. Shafi, Kazi Saeed Alam, Sk. Imran Hossain et al.· arXiv.org· 1 citation
TIGER is presented, an inference-time framework that redesigns feedback for localized repair that reduces unsupported content while preserving task quality and a CrisisFACTS case study suggests that the same repair mechanism can improve grounding in multi-source settings.
Kaixiang Zhao, Tianrun Yu, Shawn Huang et al.· 0 citations
This study formalizes Autonomous Agentic Data Engineering, a novel task designed to evaluate LLMs as autonomous data engineers that drive model specialization through end-to-end data curation, and charts a path toward agent-driven model specialization.
To make CBM measurable, BeliefTrack is introduced, a closed-world benchmark spanning Rule Discovery and Circuit Diagnosis, where a finite belief space and symbolic verifiers enable exact turn-level evaluation.
Hao-Ming Xu, Weihong Xu, Zongrui Li et al.· arXiv.org· 2 citations
This paper presents a method for aggressively pruning experts from modern mixture-of-experts LLMs while incurring negligible degradation in translation quality, and shows that translation requires only a fraction of the LLM, enabling substantial compression of the MoE blocks that contain over 90% of parameters.
Liu O. Martin, Lucas Bandarkar, Nanyun Peng· arXiv.org· 2 citations· ⚡1
A taxonomy of CoT is proposed consisting of Explicit CoT, which outputs all operations without aggregation, Composed CoT, which combines multiple operations into a single step, and Implicit CoT, which omits intermediate operations.
Kohsei Matsutani, Gouki Minegishi, Takeshi Kojima et al.· arXiv.org· 1 citation
The Unlearning Depth Score (UDS), a metric that quantifies the mechanistic depth of unlearning via activation patching, is introduced, confirming the causal approach as the most reliable for unlearning evaluation.
Jaeung Lee, Dohyun Kim, Jaemin Jo· arXiv.org· 1 citation
FinCAD is proposed, an inference-time adaptation of Context-Aware Decoding that attenuates contributions from memorised historical outcomes without retraining and raises the subset-averaged in-sample/out-of-sample Spearman correlation on an eleven-model leaderboard.
SciAtlas is presented, a shared, machine-actionable cross-disciplinary scholarly knowledge infrastructure that integrates evidential, conceptual, disciplinary, expertise, and normative layers under a shared schema and achieves a unified neuro-symbolic retrieval mechanism that grounds heterogeneous research objects, propagates relevance across the scholarly topology, and projects the resulting relevance field into the context required by each scientific workflow.
Shuofei Qiao, Yun-Xiang Wei, Bu-Sheng Zhang et al.· 1 citation
Focusing on the autonomous driving safety-critical case of pedestrian detection in the dark, it is shown how synthetic low-light samples can be used to better characterize the performance of a state-of-the-art object detection model as a function of the scene illumination.
V. Pais, Malena Mendilaharzu, Daniele Faccio et al.· arXiv.org· 0 citations
Self-Anchored Consensus (SAC), a fully decentralized filter-and-refine protocol in which agents iteratively exchange responses, locally evaluate and filter unreliable messages, and refine their own outputs, is proposed.
Haejoon Lee, Vincent Yun, Hyeonho Oh et al.· arXiv.org· 3 citations