· System-2 Reasoning: From Semantic Anchoring to Causal Intelligence· 0 citations· 19 references
TL;DR
Preliminary experiments reveal a Four-Quadrant Control Landscape where static audit policies universally fail, a finding that demonstrates CausalT5k ’s value for advancing trustworthy reasoning systems.
DiagLoop is presented, a counterfactual data flywheel that converts codified physical relations or clinical guidelines, authored once per mechanism family, into training supervision beyond recorded cases, and improves strict path correctness over the strongest conventional baseline.
CausalArena is introduced, a unified and evolvable benchmark for causal discovery under a common protocol, and substantial ranking shifts across SCM families and protocols are revealed, showing that strong performance in one benchmark regime does not reliably transfer to others.
Zi-Rong Li, Si-Zhuang Liu, Tian-Zuo Wang et al.· 0 citations
A probe corpus of 42 retracted, fraudulent, and pseudoscientific papers is paired with a methodology for eliciting and scoring single-shot model engagement with each paper's framing, indicating an urgent need for guardrail infrastructure for scientific deployment of language models.
Preliminary results reveal substantial differences in safety-utility calibration across models and show that failures can emerge only after several initially safe interaction steps, motivating treating agent safety as a trajectory-level property rather than a single-turn or binary success criterion.
Sadia Asif, Mohammad Mohammadi Amiri, Momin Abbas et al.· 0 citations
A mechanistic analysis of the jailbreaking behavior in a large-scale, safety-aligned LLM, focusing on LLaMA-2-7B-chat-hf identifies computational circuits responsible for generating affirmative responses to jailbreak prompts and uncovers key attention heads and MLP pathways that mediate adversarial prompt exploitation.
Paria Mehrbod, Boris Knyazev, Guy Wolf et al.· 0 citations
This work uses synthetic multi-hop lookup tasks to measure faithfulness causally at the activation level, specifically on self-generated reasoning, and aims to measure faithfulness causally at the activation level, specifically on self-generated reasoning.
Abhiram Bhupatiraju, Rayan Nyaupane· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.