Background Drug–drug interactions (DDIs) remain a major contributor to preventable patient harm, particularly in the context of polypharmacy. Over the past two decades, interventions to mitigate inappropriate prescribing have evolved from deterministic, rule-based clinical decision support toward increasingly complex data-driven and stochastic models. However, the extent to which these methodological advances translate into improved clinical safety remains unclear. Methods We conducted a systematic review in accordance with PRISMA 2020 guidelines, guided by the SPIDER framework. PubMed, Scopus, ScienceDirect, and IEEE Xplore were searched from inception to October 2025 for primary studies evaluating computational or clinical decision support interventions aimed at reducing DDIs or inappropriate prescribing. Eligible studies included deterministic rule-based systems, ontological frameworks, and artificial intelligence-driven predictive models. Risk of bias was assessed using the Prediction Model Risk of Bias Assessment Tool, extended with artificial intelligence-specific considerations (PROBAST+AI). Due to heterogeneity in study designs and outcome measures, findings were synthesized narratively. Results Ten studies met the inclusion criteria. Earlier interventions predominantly employed deterministic approaches focused on workflow optimization, alert management, and policy enforcement, demonstrating modest improvements in prescribing processes but inconsistent links to patient-level outcomes. More recent studies applied stochastic and generative models using high-dimensional clinical datasets to predict DDIs, reporting strong internal performance metrics. However, PROBAST+AI assessment identified a consistently high risk of bias in the analysis domain for AI-driven studies, primarily due to limited external validation, insufficient calibration reporting, and unclear handling of overfitting and data leakage. Conclusions While stochastic and generative models offer enhanced predictive capacity for DDI detection, current evidence does not demonstrate a proportional improvement in clinically reliable decision support. Deterministic systems provide transparency and safety constraints but lack adaptability to patient-specific contexts. Future interventions must prioritize hybrid architectures that integrate explainable rule-based guardrails with rigorously validated stochastic models to ensure that methodological complexity yields reproducible gains in patient safety.
GAOKAO-Bench is introduced, an intuitive benchmark that employs questions from the Chinese GAOKAO examination as test samples, including both subjective and objective questions that contribute a robust evaluation benchmark for future large language models and offers valuable insights into the advantages and limitations...
Xiaotian Zhang, Chun-yan Li, Yi Zong et al.· arXiv.org· 216 citations· ⚡17
This work investigates the possibilities of using LLMs in a resume screening setting via a document retrieval framework that simulates job candidate selection and finds that the MTEs are biased, significantly favoring White-associated names in 85% of cases and female-associated names in only 11.1% of cases.
Empirically, PRISM reduces the end-to-end time for data selection and model tuning to just 30% of conventional pipelines, and achieves this efficiency while simultaneously enhancing performance, surpassing models fine-tuned on the full dataset across eight multimodal and three language understanding benchmarks.
Jinhe Bi, Yifan Wang, Danqi Yan et al.· arXiv.org· 73 citations· ⚡4
This paper proposes adaptive sampling with approximate expected futures (ASAp), a decoding algorithm that guarantees the output to be grammatical while provably producing outputs that match the conditional probability of the LLM's distribution conditioned on the given grammar constraint.
Kanghee Park, Jiayu Wang, Taylor Berg-Kirkpatrick et al.· Neural Information Processin...· 70 citations· ⚡5
The method, ECCOLA, is presented, which aims at making the high-level AI ethics principles more practical, making it possible for developers to more easily implement them in practice.
Ville Vakkuri, Kai-Kristian Kemell, P. Abrahamsson· EUROMICRO Conference on Soft...· 64 citations· ⚡6
The goal is to not only refine the accuracy of the LLM-based tool but also to underscore its potential in streamlining the software development lifecycle through proactive code improvement and education.
Z. Rasheed, Malik Abdul Sami, Muhammad Waseem et al.· arXiv.org· 62 citations· ⚡3
Related blog posts
MIT News · Artificial Intelligence· news.mit.eduSep 30, 2026
With $2.1 million funding from Google.org, the open-source Public Transit Intelligence Hub will unify public transit monitoring, operations, and passenger communication.