Oct 2026· Journal of the American College of Radiology· 0 citations
Medicine
Abstract
Purpose
Federated learning enables breast-imaging sites to jointly train mammography artificial intelligence (AI) without sharing images, but radiologists at each site must still annotate selected images during active-learning rounds. We developed and evaluated a client-adaptive vision-language gatekeeper that withholds radiopaque artifact-containing images considered unsuitable for a breast-density annotation objective and quantified its effect on annotation time, queue quality, and label reliability.
Methods
In this retrospective, scanner-partitioned study, we used FedEMBED, derived from the public Emory Breast Imaging Dataset (222,700 training and 65,891 test images; four scanners; artifact prevalence 4.8%-13.1% per scanner). PromptGate adapted a biomedical vision-language model with site-specific learnable prompts and classified each image as relevant in-distribution (ID) or irrelevant out-of-distribution (OOD) using the class with the highest softmax probability. A single radiologist who is a member of the study team and has 2 years of breast imaging experience independently read a patient-separated, artifact-enriched cohort of 500 single mammographic views (ID, n=346; OOD, n=154). Cases were randomized, ID and OOD images were intermixed, and the reader was blinded to reference labels, PromptGate decisions, enrichment status, and scanner metadata.
Results
PromptGate withheld 101 of 154 artifact-containing images and 54 of 346 ID images, corresponding to 65.6% artifact sensitivity, 84.4% ID retention, and a 15.6% false-removal rate among ID images. Gross annotation time decreased by ∼30% (3.82 to 2.68 hours; 1.14 hours) in the enriched cohort; an analysis crediting only correctly removed artifacts yielded an estimated 20.0% reduction. A simple prevalence-based extrapolation suggests an approximate 20- to 65-min saving per 500 images at the observed scanner-specific prevalence, although the absolute benefit is expected to be smaller than in the enriched cohort. Artifact discrimination reached an area under the receiver operating characteristic curve of 0.85. Reader agreement with the manual artifact reference was 93.6% (κ=0.85), whereas agreement with the clinical-record density label was 60.9% exact (κ=0.48). Queue purity reached 92.5%, compared with 88.5% without filtering and 87.0% with a static vision-language model filter, while downstream density accuracy remained similar to the unfiltered condition.
Conclusion
In this retrospective proof-of-concept study, a privacy-oriented, scanner-adaptive vision-language gate improved the concentration of task-relevant images in federated breast-density annotation queues and reduced gross reading time. Because some ID images were incorrectly withheld and the reader study used an artifact-enriched cohort, a prospective multireader and multi-institutional evaluation is needed before operational deployment.
The results indicate that software engineering work practices are chosen opportunistically, adapted and configured to provide value under the constrains imposed by the startup context.
Nicolò Paternoster, Carmine Giardino, M. Unterkalmsteiner et al.· Information and Software Tec...· 394 citations· ⚡54
The results are packaged in the Greenfield Startup Model (GSM), which explains the priority of startups to release the product as quickly as possible, and the need to shorten time-to-market, by speeding up the development through low-precision engineering activities.
Carmine Giardino, Nicolò Paternoster, M. Unterkalmsteiner et al.· IEEE Transactions on Softwar...· 178 citations· ⚡14
This state-of-practice investigation was performed using a literature review followed by a multiple-case study approach and presents how inconsistency between managerial strategies and execution can lead to failure by means of a behavioral framework.
Carmine Giardino, Xiaofeng Wang, P. Abrahamsson· International Conference on...· 175 citations· ⚡19
Software startup companies develop innovative, software-intensive products within limited timeframes and with few resources, searching for sustainable and scalable business models.
M. Unterkalmsteiner, P. Abrahamsson, Xiaofeng Wang et al.· e-Informatica Software Engin...· 157 citations· ⚡17
This study conducts a case survey study based on the secondary data of the major pivots happened in 49 software startups, and demonstrates that customer need pivot is the most common among all pivot types.
Sohaib Shahid Bajwa, Xiaofeng Wang, Anh Nguyen-Duc et al.· Empirical Software Engineeri...· 127 citations· ⚡15
The comparison of adopter and non-adopter sample reveals three potential adoption inhibitor, security, data privacy, and portability, which underlines the importance of the technical and security perspectives for research investigating the adoption of technology.
Nattakarn Phaphoom, Xiaofeng Wang, S. Samuel et al.· Journal of Systems and Softw...· 111 citations· ⚡8
What does it take to trust AI-driven HVAC optimization? Our AI Model Factory combines agents, machine learning, reinforcement learning and deterministic checks in a governed workflow designed for messy, real-world building data. The post We built an AI factory for HVAC control appeared first on GPT-Lab.
MIT News · Artificial Intelligence· news.mit.eduAug 18, 2026
A new method for surgically removing training examples from a model reveals that as datasets grow, the link between what a model learns and what it produces dissolves.
Microsoft Research Blog· microsoft.comJul 30, 2026
LLMs do not get smarter just by remembering more. EvoLib turns experience into evolving knowledge, taking reusable skills and insights that help models learn and adapt across tasks long after deployment. The post EvoLib: Turning experience into evolving knowledge appeared first on Microsoft Research.
MIT News · Artificial Intelligence· news.mit.eduMay 20, 2026