Skip to content

Author

Bo-Wen Liu

3 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Jul 2026

Context-Masked Truncated Reasoning Audits for Answer-Key Dependence in LLM Tutors

Context masking is established as necessary for attributing early answer availability to an explanation rather than its hidden input when early-prefix evidence disappears after masking.

Bo-Nan Shen, Ding-Yan Shang, You Wang et al. · 0 citations
Jul 2026

Safety, or Just Capability? A Validity Audit of Agent-Safety Benchmarks

Naming the benchmark, metric, target behavior, and model panel is the minimum a safety claim needs, and both instruments score harmful compliance, so this is evidence of convergent validity rather than general safety.

You Wang, Xiao Han, Ding-Yan Shang et al. · 1 citation
Preprint Aug 2026

$R^3$-Bench: LLMs Struggle with Resource-Rational Reasoning under Shared Budgets

In cognitive science, resource rationality asks how an agent should allocate limited computation to maximize expected value. Most reasoning and agent benchmarks use independent per-task budgets; existing shared-budget studies do not calibrate suite performance against the same model's demonstrated single-problem competence. We introduce $R^3$-Bench, which evaluates six-problem suites under shared budgets across mathematics, competitive programming, and abstract reasoning in tool-free and agentic settings. Matched single-problem response curves define an offline empirical oracle over observed successes. Across 72 main-table cells for six models, the oracle mean matches or exceeds the contest mean in all cells and is strictly higher in 71. Under moderate tool-free pressure, equal-allocation replay also exceeds contest performance for four of six models. Trajectory diagnostics reveal limited strategy updating and pressure-dependent failure patterns. In a three-model diagnostic under strong agentic pressure, at least one fixed scheduler exceeds the contest mean in six of nine cells, but no policy dominates across domains. These results expose a persistent gap between demonstrated competence and shared-budget realization.

Peisong Wang, Zhiwei Ma, Bo-Wen Liu et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.