Real-world evaluation of large language models in detecting drug-related problems: A clinical pharmacist-AI concordance study in hematology care
| dc.contributor.author | Demirpolat, Eren | |
| dc.contributor.author | Tuncel, Mahmud Sami | |
| dc.contributor.author | Gurer, Feyza | |
| dc.contributor.author | Sanli, Neslihan Mandaci | |
| dc.date.accessioned | 2026-10-09T21:50:35Z | |
| dc.date.issued | 2026 | |
| dc.department | Yüksek İhtisas Üniversitesi | |
| dc.description.abstract | Introduction Large language models (LLMs) offer potential as clinical decision support systems (CDSS) for detecting drug-related problems (DRPs), yet their real-world performance compared to clinical pharmacists (CPs) remains unclear, especially in complex hematology care. We aimed to evaluate the concordance between a clinical pharmacist and three LLMs in identifying DRPs within a Bone Marrow Transplantation unit.Methods This prospective observational study evaluated the concordance between a CP and three LLMs (ChatGPT-4o, Grok-3, DeepSeek-v3) in a Bone Marrow Transplantation unit. Eighty-three anonymized patient cases encompassing 210 CP-identified DRPs, classified via the PCNE v9.1 system, were presented using a standardized CDSS-simulating prompt. Performance was assessed based on direct detection, prompted detection after structured follow-up, and the clinical relevance of AI-generated therapeutic recommendations against the CP's gold-standard assessments.Results Direct detection of intervention-requiring DRPs was limited (51.4%-60.5% across models), with nearly half missed initially. Guided prompting significantly improved overall detection rates to 93.8%-98.1%, with ChatGPT achieving the highest accuracy. All models produced hallucinations. Recommendation concordance with the CP exceeded 70% in most DRP categories. DeepSeek and ChatGPT showed more consistent performance in context-dependent evaluations, whereas Grok demonstrated higher direct detection but lower recommendation alignment. LLMs demonstrate meaningful potential to assist in DRP detection but are not sufficiently reliable as standalone tools. Expert-guided interaction substantially enhanced their performance, underscoring the critical value of hybrid pharmacist-AI workflows.Conclusion Future research should validate these findings across broader populations with multiple expert evaluators and integrate next-generation AI architectures for safer CDSS implementation. | |
| dc.identifier.doi | 10.1177/10781552261418957 | |
| dc.identifier.issn | 1078-1552 | |
| dc.identifier.issn | 1477-092X | |
| dc.identifier.orcid | 0000-0003-4405-4660 | |
| dc.identifier.orcid | 0000-0002-8539-607X | |
| dc.identifier.pmid | 41662281 | |
| dc.identifier.scopus | 2-s2.0-105029562174 | |
| dc.identifier.scopusquality | Q3 | |
| dc.identifier.uri | https://doi.org/10.1177/10781552261418957 | |
| dc.identifier.uri | https://hdl.handle.net/20.500.12794/3762 | |
| dc.identifier.wos | WOS:001685365800001 | |
| dc.identifier.wosquality | Q4 | |
| dc.indekslendigikaynak | Web of Science | |
| dc.indekslendigikaynak | Scopus | |
| dc.indekslendigikaynak | PubMed | |
| dc.indekslendigikaynak.diger | Science Citation Index Expanded (SCI-EXPANDED) | |
| dc.language.iso | en | |
| dc.publisher | Sage Publications Ltd | |
| dc.relation.ispartof | Journal of Oncology Pharmacy Practice | |
| dc.relation.publicationcategory | Makale - Uluslararası Hakemli Dergi - Kurum Öğretim Elemanı | |
| dc.rights | info:eu-repo/semantics/openAccess | |
| dc.snmz | KA_WoS_20260922 | |
| dc.subject | Drug-Related Problems | |
| dc.subject | Clinical Decision Support Systems | |
| dc.subject | Artificial Intelligence | |
| dc.subject | Clinical Pharmacy | |
| dc.subject | Hematology Patients | |
| dc.title | Real-world evaluation of large language models in detecting drug-related problems: A clinical pharmacist-AI concordance study in hematology care | |
| dc.type | Article |







