Jul 2026· International Conference on Knowledge-Based Intelligent Information & Engineering Systems· Vol 270, pp. 592-601· 0 citations· 11 references
Computer Science
TL;DR
Proprietary LLMs, especially in few-shot (O1) or fine-tuned (Gemini 2.0 Pro) settings, significantly outperformed other models and confirm the power of modern LLMs for genomic knowledge extraction.
Abstract
Extracting SNP-phenotype associations from biomedical literature is vital but challenging. We benchmarked diverse NLP models, including MLMs, hybrid architectures, and state-of-the-art LLMs (Gemini 2.0, OpenAI O-series, Qwen, Mistral), on the SNPPhenA corpus across three tasks: sentence-level, abstract-level, and association strength classification. OpenAI O1 achieved state-of-the-art (SOTA) results using few-shot learning for non-finetuned sentence-level classification (F1 0.89) and established a new SOTA for abstract-level classification (F1 0.82). Association strength classification proved difficult, though fine-tuned Gemini 2.0 Pro performed best (F1 0.60) in the first LLM evaluation of this task. Proprietary LLMs, especially in few-shot (O1) or fine-tuned (Gemini 2.0 Pro) settings, significantly outperformed other models. These findings confirm the power of modern LLMs for genomic knowledge extraction.
Findings demonstrate that compact models can achieve strong biomedical classification performance through KD under compatible teacher–student pairings, while also highlighting that KD effectiveness varies substantially depending on the specific model combination.
It is demonstrated that modern LLMs with reasoning capabilities can effectively support real-time clinical evidence synthesis without task-specific fine-tuning, offering a pathway toward scalable automated systems for clinical trial interpretation that could substantially reduce the evidence-to-practice gap in medical...
Shobanapriyan Chandrasegaran, Amal Htait· International Conference on...· 0 citations
The results show domain-specific changes associated with using recursive synthetic-data training in biomedical QA, but do not establish clinical hallucination rates or universal model collapse.
A critic-based verification mechanism where a second “critic” prompt reviews and verifies extracted relations is investigated, demonstrating that this approach is highly effective, reducing false positives by 58.5% and achieving 74.1% precision with top models, competitive with supervised methods.
A. Assi, Nour El Islem Karabadji, M. Elati et al.· 1 citation
VITA's advantages in accuracy and completeness persisted under the neutral judge; its communication scores were lower, and this results indicate that a purpose-built clinical RAG system remains competitive with frontier LLMs on an open benchmark, consistent with corpus specificity as a design variable that improves gro...
P. Reddy, C. Mandke, Suvrankar Datta et al.· 0 citations
In this study, a comprehensive evaluation of abstractive and extractive summarization performance across three prominent large language models (LLMs): ChatGPT, DeepSeek, and Gemini is presented. A total of 8,000 cardiovascular-related research abstracts were collected from PubMed and summarized using two distinct promp...
Burcu Baştürk, Aytuğ Onan· Sakarya University Journal o...· 0 citations
Related blog posts
MIT News · Artificial Intelligence· news.mit.eduSep 24, 2026
A new method, called CW-Net, translates the reasoning process of an autonomous vehicle’s AI system into understandable concepts that explain its behavior.
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.