A tuned fixed weight is a difficult-to-beat default across the tested scale range; reliability estimation gave no deployable adaptive advantage.
Abstract
Objective: Published P300-speller fusion schemes fix prior trust regardless of trial reliability; we tested whether a reliability estimate improves on it. Methods: We reanalyzed 3,373 archived P300-speller selections from 47 people with ALS (BigP3BCI). A fair, matched-search-space comparison, tuning both a fixed weight and an adaptive policy out-of-fold, was evaluated across 22 evaluable language-model priors up to 46.7B parameters. Two representative priors, GPT-2 and a classical 5-gram, additionally received detailed naive and mechanistic analyses. Results: No prior's 95% CI favored adaptive fusion under the fair comparison, despite unexploited oracle headroom at every scale. Under GPT-2, the naive comparison was significantly worse for adaptive fusion; both anchors converged to a degenerate or near-degenerate fair-comparison solution. For the representative anchors, three further controllers failed to convert that headroom into benefit; the fixed-fused posterior's output probability outperformed the best controller for flagging errors (2.8- to 3.8-fold enrichment). Conclusion: A tuned fixed weight is a difficult-to-beat default across the tested scale range; reliability estimation gave no deployable adaptive advantage. Significance: Adaptive weighting should be validated against a fairly tuned baseline across model families and scales; in this dataset, the fused output's confidence identified high-risk selections better than the tested purpose-built ranker.
This retrospectively re-decoded 3,373 P300-speller selections from 47 people with amyotrophic lateral sclerosis, reconstructing a neural posterior and combining it with 25 language priors ranging from 5-grams to 46.7-billion-parameter models to reveal how strongly a fused brain-computer interface decision depends on it...
A. Gorenshtein, M. Omar, E. Jia et al.· medRxiv· 0 citations
People increasingly reason with large language models (LLMs), yet complementary capabilities do not guarantee outperforming both components. In a between-subjects study, participants (N=535) solved a 40-item battery of matrix reasoning, mental rotation, syllogisms, and letter-string analogies, unaided or with GPT-5.6-L...
Robin Welsch, Michelle Rausch, Pascal Knierim et al.· 0 citations
Objective Neural-only decoding performance does not identify how much prediction neural history adds beyond task structure or recent output. We tested how this increment changes when context and temporal information boundaries are made explicit. Approach We defined context-conditional neural gain as held-out error redu...
Zonghan Du, Zhong-Yuan Lai, Liang Hu et al.· bioRxiv· 0 citations
Language-model post-editing produced fluent semantic substitutions that rose with corruption, confidence did not reliably flag, and no interface policy removed, and this does not demonstrate clinical harm; prospective human-in-the-loop evaluation is needed.
A. Gorenshtein, M. Omar, E. Jia et al.· medRxiv· 0 citations
dLVLMs reverse the yes-bias of AR models in binary visual queries and collapse to near-zero accuracy on underrepresented racial groups with opposite-polarity gender bias, suggesting reliability is shaped by the generative paradigm together with training data.
Bounded margins mitigate confident hallucinations during post-training, implemented through an entropy-dependent margin bound in direct preference optimization (DPO) and shown to mitigate confident hallucinations during post-training.
Qing-Jia Huang, Ya-Kai Li, Jian-Guo Wu et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.