Whether independently implemented training stacks can serve as differential oracles for a whole fine-tuning pipeline, rather than the operators and inference paths that prior differential testing targets, is studied.
RatoGuide demonstrated favorable performance in typical cases, but accuracy declined in atypical cases with artifacts or altered anatomy, particularly for atypical cases and organs in high-dose gradient regions.
R. Tozuka, M. Saito, Masaki Matsuda et al.· Journal of Applied Clinical...· 0 citations
The PSI-based hybrid workflow resulted in statistically greater positioning accuracy compared to conventional splint-based fixation, however, the absolute improvements were small, likely of limited clinical relevance, and were not associated with reduced operating times.
Jonathan Bertram, H. Leonhardt, Fred Podmelle et al.· BMC Oral Health· 0 citations
These findings provide practical guidance for selecting LLM judges, designing role prompts, and employing multi-judge voting strategies in “automated software quality assurance”.
An automated, multi-dimensional evaluation framework for C# code generation, applying it to four state-of-the-art LLMs: GPT, Gemini, Claude, and Grok is presented and a substantial gap between correctness and quality attributes is revealed.
Seyed Mohammad Mahdi Ghalandarian, Majid Bazargani, Masoumeh Taromirad· 0 citations
The results recast package hallucination as both a measurement problem and a decoding-time control problem, and they demonstrate that the choice of defense must be matched to the threat model and recommendation utility.
Albérick Euraste Djiré, Iyiola E. Olatunji, Melissa Tessa et al.· 1 citation
This work introduces Risa (Routing-Informed Steering and Arbitration): within trajectories, routing encourages diverse exploration and controlled convergence during patch commitment; across separately sampled trajectories, agreement at informative patch positions selects a final candidate.
Kang Chen, Junjie Nian, Yixin Cao et al.· 0 citations
DPIAgent is proposed, a structured agentic framework built on three principles, Divide, Protocol, Isolate (DPI), that mitigates compound-objective ambiguity and goal drift, and shows that architectural structure and backbone capability are complementary axes rather than substitutes, demonstrating DPI's generalizability across model classes.
Hao Liu, Steven Liu, Xin Zhang et al.· 0 citations
SWE Refactor Bench is introduced, a benchmark comprising 20 whole-repository migrations, covering 4 kinds of technical debt, and SWE Refactor Bench is positioned as a rigorous testbed for developing coding agents for reliable whole-repository migrations.
Neuro-formal verification is introduced, which harnesses that automation for developers of mainstream programming languages and returns a Dafny proof of correctness or of a bug on 57% of the entries at 92% precision, and a CBMC counterexample for 63% of the buggy programs at 90% precision.
The first complete description of SweepLSD is given, a line segment detector that reads the image exactly once and emits each segment within a few rows of its last pixel passing the scan line, with the tightest frame-time distribution and the best per-segment direction accuracy of the four detectors.
The ezSRT test is capable of producing reliable SRT estimates in 6 min that are sensitive to different experimental conditions and listener groups and that result in different SiN performance levels when later tested in a fixed SNR configuration, and using the ezSRT test as part of the new i-RRT protocol to determine SNRs targeting specific intelligibility levels.
Christopher Slugocki, Francis Kuk, Petri Korhonen et al.· Ear and Hearing· 0 citations
A USAF cadet and a Lincoln Laboratory researcher found AI chatbots can help nontechnical service members produce viable software applications for their unique problems.