Large Language Models (LLMs) are proposed as tools for high-throughput, deep phenotyping of psychiatric disorders. Applied to electronic health records, LLMs could in principle extract patient symptoms, outcome trajectories, risk factors, and treatment history at scale and these, when combined with increasingly available biological data, such as genomics or neuroimaging, could provide powerful resources for health research. Although proof-of-principle LLM-based symptom extractions have been carried out for some medical conditions, the heterogeneous nature of psychiatric disorders, including schizophrenia, requires extensive domain-specific evaluations. Here, we evaluate 14 general-purpose LLMs for extracting eight symptoms of schizophrenia from clinical summaries in a severe mental illness cohort (N = 704). Performance across symptoms was poor to moderate (macro F1 = 0.500 - 0.647), with positive symptoms more accurately extracted than negative symptoms. Few-shot prompting, a common strategy for improving LLM task performance, did not significantly improve these results. Nevertheless, LLM-predicted and gold-standard positive symptoms demonstrated comparable associations with clinical variables in regression analyses. Individual-level extraction errors attenuated group-level associations but did so unevenly across symptom domains. This indicates that general-purpose LLMs may not be able to extract psychosis-related phenotypes from clinical summaries with the quality required for clinical use but might nevertheless be useful for exploratory research on large cohorts. However, closing the performance gap between positive and negative symptom extractions seems essential groundwork to prevent a systematic bias in any LLM-derived characterisations of psychotic symptoms.
S. Lock, J. Boisson, L. M. Evans et al.· medRxiv· 0 citations
Word-in-Context (WiC) remains challenging for language models, despite recent progress on lexical-semantic tasks. We hypothesise that this difficulty arises not only from comparing two contextual uses of a word, but also from the absence of an explicit sense inventory that specifies the relevant level of semantic granularity. We evaluate open LLMs on WiC and traditional Word Sense Disambiguation (WSD) under similar settings. We find that providing candidate senses, similar to what is done in traditional WSD, improves WiC performance in all settings. In general, explicit sense information helps models make more consistent and targeted judgements. Human evaluation further shows that many apparent WiC errors reflect label ambiguity or mismatches between model and annotator sense boundaries rather than simple failures of lexical understanding. In particular, results show that LLMs overthink the sense distinction often leading to errors based on overly fine-grained distinctions.
Yi Zhou, Kiamehr Rezaee, D. Bollegala et al.· 0 citations
A data-centric analysis of semantic knowledge acquisition in word embeddings, focusing on word analogy and semantic similarity shows that, for relational semantics, training-data quality outweighs quantity, and that simple proxy models remain a practical, interpretable tool for efficient data selection.
Aishwarya Jadhav, Mark Anderson, J. Camacho-Collados et al.· Neural computing & applicati...· 0 citations
It is found that linguistic proximity itself introduces errors: closely related language pairs tend to perform worse, reflecting the challenge of semantic discrimination due to lexical overlap, and unlike other tasks where language distance poses additional challenges, it is found that linguistic proximity itself introduces errors.
Marta Vázquez Abuín, José Camacho-Collados, Marcos García· Annual Meeting of the Associ...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.