<link rel="stylesheet" href="styles.f3b1fba60ec7970c.css">

Comparative quality, accuracy, and readability of large language model responses to patient questions about robotic-assisted total knee arthroplasty

dc.contributor.authorKolac, Ulas Can
dc.contributor.authorSili, Mazlum Veysel
dc.contributor.authorKarademir, Orhan Mete
dc.contributor.authorAyik, Gokhan
dc.contributor.authorAkkaya, Mustafa
dc.contributor.authorÇakmak, Gökhan
dc.date.accessioned2026-10-09T21:43:56Z
dc.date.issued2026
dc.departmentYüksek İhtisas Üniversitesi
dc.description.abstractPurpose: To compare the information quality, accuracy, and readability of patient-directed responses generated by large language models (LLMs), including ChatGPT-o3, ChatGPT-5.2, Gemini 3, and DeepSeek, regarding robotic-assisted total knee arthroplasty (RA-TKA). Methods: Thirty frequently asked patient questions were identified using LLM outputs and Google search queries. Responses were evaluated for information quality using the DISCERN and Quality Analysis of Medical Artificial Intelligence (QAMAI) instruments, for clinical accuracy using a 5-point ordinal rating scale, and for understandability and readability using the PEMAT Understandability and Flesch-Kincaid Reading Ease scores. Results: Median DISCERN scores were 46.0 (range, 35.0–50.0) for ChatGPT-o3, 45.75 (28.5–51.0) for ChatGPT-5.2, 43.75 (32.0–47.5) for Gemini 3, and 42.0 (32.0–50.0) for DeepSeek, with a significant overall difference among models (p < 0.001). The 5-point clinical accuracy scores were similar across models (median 4.0; p = 0.636). Median QAMAI scores were 23.0 for all four models, without a significant between-model difference (p = 0.462). PEMAT Understandability scores differed significantly among models (p < 0.001), with median scores of 90.0 for ChatGPT-o3 and ChatGPT-5.2, 88.0 for Gemini 3, and 85.0 for DeepSeek. Flesch-Kincaid Reading Ease scores also differed significantly (p < 0.001); Gemini 3 demonstrated higher readability than both ChatGPT models, whereas DeepSeek demonstrated higher readability than ChatGPT-o3. Conclusion: The evaluated LLMs demonstrated generally acceptable clinical accuracy but differed across measures of written information quality, understandability, and readability. No significant difference was detected using QAMAI. Although Gemini 3 and DeepSeek demonstrated greater readability in selected comparisons, median responses across all models remained above recommended patient-education reading levels. LLM-generated responses should therefore be regarded as supplementary rather than standalone sources of patient information regarding RA-TKA. © 2026 Elsevier B.V.
dc.identifier.doi10.1016/j.knee.2026.104635
dc.identifier.issn0968-0160
dc.identifier.pmid42727207
dc.identifier.scopus2-s2.0-105049982709
dc.identifier.scopusqualityQ2
dc.identifier.urihttps://doi.org10.1016/j.knee.2026.104635
dc.identifier.urihttps://hdl.handle.net/20.500.12794/3283
dc.identifier.volume63
dc.indekslendigikaynakScopus
dc.indekslendigikaynakPubMed
dc.language.isoen
dc.publisherElsevier B.V.
dc.relation.ispartofKnee
dc.relation.publicationcategoryMakale - Uluslararası Hakemli Dergi - Kurum Öğretim Elemanı
dc.rightsinfo:eu-repo/semantics/openAccess
dc.snmzKA_Scopus_20260922
dc.subjectArthroplasty
dc.subjectArtificial Intelligence
dc.subjectLarge Language Models
dc.subjectPatient Information
dc.subjectRobotic Knee Arthroplasty
dc.titleComparative quality, accuracy, and readability of large language model responses to patient questions about robotic-assisted total knee arthroplasty
dc.typeArticle

Files