· System-2 Reasoning: From Semantic Anchoring to Causal Intelligence· 0 citations· 53 references
TL;DR
RAudit, a diagnostic protocol for auditing LLM reasoning without ground truth access, is presented and it is proved bounded correction and O ( log ( 1 /𝜀)) termination are proved.
Nudgeability offers a simple, post-training-free way to evaluate both sensitivity and targeting as endogenous self-reflection mechanisms mature as endogenous self-reflection mechanisms mature.
This case shows how evaluation-design validity can be checked structurally before model inference and why base correctness does not determine intervention-response fidelity.
Jun Luo, Ning Huang, Zi-Qi Sha et al.· 0 citations
This work uses synthetic multi-hop lookup tasks to measure faithfulness causally at the activation level, specifically on self-generated reasoning, and aims to measure faithfulness causally at the activation level, specifically on self-generated reasoning.
Large language models are often asked which input factors influenced their outputs. For structured inputs, such reports can be checked by counterfactual perturbation, but each factor must be queried multiple times to estimate its effect, so verification is usually budget-limited. We study how this limited-budget settin...
Tao-Lin Zhang, Han-Yu Wang, Jiu-Heng Wan et al.· 0 citations
Coding agents can return plausible patches that omit required behavior. These failures are hard to review because long traces and confident summaries often hide what was missed. We ask when a nominally weaker reviewer can reliably decide whether a patch solves its issue. We study 411 execution-labeled traces from three...
Jun-Yu Guo, Shangding Gu, Ming Jin et al.· 0 citations
A probe corpus of 42 retracted, fraudulent, and pseudoscientific papers is paired with a methodology for eliciting and scoring single-shot model engagement with each paper's framing, indicating an urgent need for guardrail infrastructure for scientific deployment of language models.
V. Rodionov, Shamil Assylbekov· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.