Cura 1T, a healthcare-specialized LLM built on the open-weight Kimi-K2.6 and trained through a human-gated recursive self-improvement (RSI) loop, ranks at or near the top among frontier baselines while remaining competitive on out-of-domain reasoning and agentic benchmarks.
Abstract
Healthcare AI agents handle patient consultation, clinical reasoning over text and images, interactive diagnosis, and electronic health record (EHR) tool use, yet specialized agentic models that cover these use cases together remain limited. These capabilities fail in different ways, and a narrow update for one task can degrade another. We present Cura 1T, a healthcare-specialized LLM built on the open-weight Kimi-K2.6 and trained through a human-gated recursive self-improvement (RSI) loop. Specifically, in each round, the RSI harness plans a target capability, trains the model, evaluates benchmark trajectories, and refines the data mixture from observed failures with targeted synthetic and curated examples rather than a single generic medical-data update. Across the healthcare evaluation suite, Cura 1T ranks at or near the top among frontier baselines while remaining competitive on out-of-domain reasoning and agentic benchmarks.
A scoping review with systematic evidence mapping across five electronic sources, screened 1,649 exportable records, and provisionally included 557 unique studies that met predefined criteria for goal-directed task execution, tool use, interaction with external resources, feedback-based refinement, or multi-agent colla...
Zheng Tong, Yang Liu, Wan-Shu Fan et al.· arXiv.org· 0 citations
The PatientAgentBench framework is released as a reproducible, clinician-validated evaluation standard to help the field close this gap in healthcare agentic healthcare, and is validated by licensed clinicians annotated shared conversations.
Korosh Vatanparvar, Ashutosh Joshi, Maria Xenochristou et al.· arXiv.org· 0 citations
General-purpose language models generate fluent health reports that can fabricate derived clinical metrics. In an illustrative comparison on identical two-week CGM and meal data, leading foundation models produced reports with invented MAGE values, inflated meal counts, and unreferenced complication-risk projections: f...
A. Diament, G. Sapir, M. Gorodetski et al.· medRxiv· 0 citations
Electronic health records support a wide spectrum of clinical prediction and decision-support studies, but reproducible EHR research now requires more than training a single predictive model. As the field expands from machine learning and deep learning to LLM-based and agentic AI, differences in cohort construction, te...
Yinghao Zhu, Zi-Xiang Wang, Lei Gu et al.· Proceedings of the 32nd ACM...· 0 citations
A role-specialized Mixture-of-Agents (MoA) that combines medical knowledge retrieval with contrastive similar-patient reasoning is studied, placing role design as a key factor in privacy-constrained, training-free clinical LLM prediction.
This entry-level tutorial aims to equip healthcare professionals with the tools necessary to effectively integrate LLMs into clinical practice, ensuring that these powerful technologies are applied in a safe, reliable, and impactful manner.
Qiao Jin, Nicholas Wan, Robert Leaman et al.· Nature Protocols· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.