Skip to content

Category

software testing

441 papers

#software testing Review Open access Aug 2026

Clinical evaluation of novel deep learning‐based auto‐segmentation software: Utility and potential pitfalls

RatoGuide demonstrated favorable performance in typical cases, but accuracy declined in atypical cases with artifacts or altered anatomy, particularly for atypical cases and organs in high-dose gradient regions.

R. Tozuka, M. Saito, Masaki Matsuda et al. · 0 citations
#software testing Open access Aug 2026

Accuracy of PSI-based hybrid workflow using a temporary intermediate splint versus conventional splint-based maxillary positioning in orthognathic surgery: a retrospective cohort study

The PSI-based hybrid workflow resulted in statistically greater positioning accuracy compared to conventional splint-based fixation, however, the absolute improvements were small, likely of limited clinical relevance, and were not associated with reduced operating times.

Jonathan Bertram, H. Leonhardt, Fred Podmelle et al. · 0 citations
#software testing Preprint Aug 2026

Benchmarking the Titans: A Multi-Dimensional Empirical Evaluation of LLM Code Generation Quality in the .NET Ecosystem

An automated, multi-dimensional evaluation framework for C# code generation, applying it to four state-of-the-art LLMs: GPT, Gemini, Claude, and Grok is presented and a substantial gap between correctness and quality attributes is revealed.

Seyed Mohammad Mahdi Ghalandarian, Majid Bazargani, Masoumeh Taromirad · 0 citations
#software testing Preprint Aug 2026

Disagree to Explore, Agree to Commit: Routing-Guided Test-Time Scaling for Software Agents

This work introduces Risa (Routing-Informed Steering and Arbitration): within trajectories, routing encourages diverse exploration and controlled convergence during patch commitment; across separately sampled trajectories, agreement at informative patch positions selects a final candidate.

Kang Chen, Junjie Nian, Yixin Cao et al. · 0 citations
#software testing Preprint Aug 2026

DPIAgent: Divide, Protocol, Isolate for Agentic Reproduction Test Generation

DPIAgent is proposed, a structured agentic framework built on three principles, Divide, Protocol, Isolate (DPI), that mitigates compound-objective ambiguity and goal drift, and shows that architectural structure and backbone capability are complementary axes rather than substitutes, demonstrating DPI's generalizability across model classes.

Hao Liu, Steven Liu, Xin Zhang et al. · 0 citations
#software testing Preprint Aug 2026

SWE Refactor Bench: Can Coding Agents Complete a Long-Horizon, Whole-Repository Stack Migration?

SWE Refactor Bench is introduced, a benchmark comprising 20 whole-repository migrations, covering 4 kinds of technical debt, and SWE Refactor Bench is positioned as a rigorous testbed for developing coding agents for reliable whole-repository migrations.

Deyao Hong, Y. Chi, Wenyi Li et al. · 0 citations
#software testing Review Aug 2026

Neuro-Formal Verification: Agentic Language-Agnostic Formal Program Reasoning

Neuro-formal verification is introduced, which harnesses that automation for developers of mainstream programming languages and returns a Dafny proof of correctness or of a bug on 57% of the entries at 92% precision, and a CBMC counterexample for 63% of the buggy programs at 90% precision.

Shuvendu K. Lahiri · 0 citations
#software testing Preprint Aug 2026

SweepLSD: A One-Pass, O(width)-Memory Line Segment Detector with an Integer-Only Streaming Core and a Real-Time FPGA Realization

The first complete description of SweepLSD is given, a line segment detector that reads the image exactly once and emits each segment within a few rows of its last pixel passing the scan line, with the tightest frame-time distribution and the best per-segment direction accuracy of the four detectors.

Yoshiyasu Shimizu · 0 citations

The Intelligibility-Based Repeat-Recall Test: I. Bayesian-Guided Estimation of Multiple Speech Reception Thresholds.

The ezSRT test is capable of producing reliable SRT estimates in 6 min that are sensitive to different experimental conditions and listener groups and that result in different SiN performance levels when later tested in a fixed SNR configuration, and using the ezSRT test as part of the new i-RRT protocol to determine SNRs targeting specific intelligibility levels.

Christopher Slugocki, Francis Kuk, Petri Korhonen et al. · 0 citations

From tech blogs

See all →
MIT News · Artificial Intelligence Aug 17, 2026

Q&A: Rethinking how innovation happens

In his latest book, Professor Eugene Fitzgerald examines the forces that turn breakthroughs into value — and why innovation resists simple formulas.