Multi-path reasoning methods such as self-consistency (SC) sample $K$ reasoning paths and choose the most frequent answer. However, their gains quickly plateau as $K$ increases, and existing methods do not predict when this saturation will occur. We formalize multi-path LLM reasoning as a diversity combining problem fr...
Guang-Sheng Yu, Litianyi Zhang, Qin Wang et al.· 0 citations
Test-time reinforcement learning (TTRL) enables models to improve their reasoning without relying on labeled training data, but existing approaches typically optimize a large fraction of the model parameters. This raises a natural question: can effective test-time adaptation emerge when both the reward signal and the o...
PhysAssistBench is introduced, a benchmark for interactive doctor-patient-EHR assistance that uses a scalable pipeline to construct agentic patients: interactive, record-grounded agents that turn static EHR records into multi-turn clinical scenarios while preserving clinical factuality.
T. Du, Peijie Yu, Sihan Shang et al.· arXiv.org· 0 citations
This work proposes a pipeline for automatically generating step-by-step linguistic reasoning traces from Universal Dependencies treebanks, dictionaries, and grammar-rule banks and shows that linguistic reasoning traces are most effective as inference-time guidance in ICL, which substantially improve translation perform...
It is found that standard benchmarks do not adequately capture the strengths of the dataset, but expert judgment shows that SQPsych makes LLMs significantly better at therapist roleplaying.
Doan Nam Long Vu, Rui Tan, L. Mary Moench et al.· arXiv.org· 5 citations· ⚡2
The results suggest that Engram can serve as a reusable external knowledge artifact, provided that the target has access to a compatible reader interface and target-side adaptation can further improve alignment when direct reader reuse is insufficient.
Mingyuan Li, Guangsheng Yu, Xu Wang et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.