<link rel="stylesheet" href="styles.f3b1fba60ec7970c.css">

Real-world evaluation of large language models in detecting drug-related problems: A clinical pharmacist-AI concordance study in hematology care

dc.contributor.authorDemirpolat, Eren
dc.contributor.authorTuncel, Mahmud Sami
dc.contributor.authorGurer, Feyza
dc.contributor.authorSanli, Neslihan Mandaci
dc.date.accessioned2026-10-09T21:50:35Z
dc.date.issued2026
dc.departmentYüksek İhtisas Üniversitesi
dc.description.abstractIntroduction Large language models (LLMs) offer potential as clinical decision support systems (CDSS) for detecting drug-related problems (DRPs), yet their real-world performance compared to clinical pharmacists (CPs) remains unclear, especially in complex hematology care. We aimed to evaluate the concordance between a clinical pharmacist and three LLMs in identifying DRPs within a Bone Marrow Transplantation unit.Methods This prospective observational study evaluated the concordance between a CP and three LLMs (ChatGPT-4o, Grok-3, DeepSeek-v3) in a Bone Marrow Transplantation unit. Eighty-three anonymized patient cases encompassing 210 CP-identified DRPs, classified via the PCNE v9.1 system, were presented using a standardized CDSS-simulating prompt. Performance was assessed based on direct detection, prompted detection after structured follow-up, and the clinical relevance of AI-generated therapeutic recommendations against the CP's gold-standard assessments.Results Direct detection of intervention-requiring DRPs was limited (51.4%-60.5% across models), with nearly half missed initially. Guided prompting significantly improved overall detection rates to 93.8%-98.1%, with ChatGPT achieving the highest accuracy. All models produced hallucinations. Recommendation concordance with the CP exceeded 70% in most DRP categories. DeepSeek and ChatGPT showed more consistent performance in context-dependent evaluations, whereas Grok demonstrated higher direct detection but lower recommendation alignment. LLMs demonstrate meaningful potential to assist in DRP detection but are not sufficiently reliable as standalone tools. Expert-guided interaction substantially enhanced their performance, underscoring the critical value of hybrid pharmacist-AI workflows.Conclusion Future research should validate these findings across broader populations with multiple expert evaluators and integrate next-generation AI architectures for safer CDSS implementation.
dc.identifier.doi10.1177/10781552261418957
dc.identifier.issn1078-1552
dc.identifier.issn1477-092X
dc.identifier.orcid0000-0003-4405-4660
dc.identifier.orcid0000-0002-8539-607X
dc.identifier.pmid41662281
dc.identifier.scopus2-s2.0-105029562174
dc.identifier.scopusqualityQ3
dc.identifier.urihttps://doi.org/10.1177/10781552261418957
dc.identifier.urihttps://hdl.handle.net/20.500.12794/3762
dc.identifier.wosWOS:001685365800001
dc.identifier.wosqualityQ4
dc.indekslendigikaynakWeb of Science
dc.indekslendigikaynakScopus
dc.indekslendigikaynakPubMed
dc.indekslendigikaynak.digerScience Citation Index Expanded (SCI-EXPANDED)
dc.language.isoen
dc.publisherSage Publications Ltd
dc.relation.ispartofJournal of Oncology Pharmacy Practice
dc.relation.publicationcategoryMakale - Uluslararası Hakemli Dergi - Kurum Öğretim Elemanı
dc.rightsinfo:eu-repo/semantics/openAccess
dc.snmzKA_WoS_20260922
dc.subjectDrug-Related Problems
dc.subjectClinical Decision Support Systems
dc.subjectArtificial Intelligence
dc.subjectClinical Pharmacy
dc.subjectHematology Patients
dc.titleReal-world evaluation of large language models in detecting drug-related problems: A clinical pharmacist-AI concordance study in hematology care
dc.typeArticle

Files