Martini, Sebastian; Schluessel, Sabine; Aghamaliyev, Ughur; Rippl, Michaela; Deissler, Linda; Tausendfreund, Olivia; Nuebler, Desiree; Mueller, Katharina; Schmidmaier, Ralf; Drey, Michael (2026): Expert Evaluation of the Perceived Accuracy, Relevance, and Safety of Large Language Model–Generated Patient Information in Geriatrics: Cross-Condition Study. JMIR AI, 5: e91369. ISSN 2817-1705
Veröffentlichte Publikation
martini_chatgpt_jmir-ai05-2026_final.pdf
Abstract
Background:
Large language models (LLMs) are increasingly used to generate patient-oriented medical information. In geriatrics, such information must balance accuracy, relevance, and safety, as older adults may be particularly susceptible to misleading or harmful advice. However, systematic evaluations of expert perceptions across multiple geriatric conditions remain limited.
Objective:
This study aimed to explore geriatricians’ perceptions of the accuracy, relevance, and potential harm of LLM-generated patient information across common geriatric conditions and to examine variability and interrater agreement in expert ratings.
Methods:
In this cross-sectional expert rating study, 10 geriatricians evaluated 50 LLM-generated statements covering 5 geriatric conditions (sarcopenia, osteoporosis, urinary incontinence, depression, and dementia). Statements addressed diagnostic, etiological, prognostic, risk-related, and therapeutic aspects. Experts rated perceived accuracy, relevance, and potential harm using 5-point Likert scales. Rating distributions were summarized using medians and IQRs. The Kendall coefficient of concordance (W) was used exploratorily to assess agreement in the relative ordering of statements within predefined strata. Readability was assessed using Flesch-Kincaid Grade Level and Flesch Reading Ease.
Results:
Expert ratings indicated high perceived accuracy (median 4.32, IQR 4.01-4.59 and perceived relevance (median 4.51, IQR 4.06-4.66), while perceived potential harm remained low (median 1.59, IQR 1.17-1.92). IQR values ranged from 0.00 to 1.38 with most values clustering below 0.5, indicating limited dispersion in expert ratings. Agreement in the relative ordering of statements varied across domains, with W values ranging from 0.27 to 0.62 (median 0.53, IQR 0.46-0.58), indicating moderate concordance. No statements combined low perceived accuracy with high perceived potential harm. Readability analysis indicated generally accessible language, with a median Flesch-Kincaid Grade Level of 8.3 (IQR 7.4-9.6) and a median Flesch Reading Ease score of 60.8 (IQR 50.1-66.9).
Conclusions:
LLM-generated patient information for common geriatric conditions was rated as largely accurate and relevant, with low potential harm in typical scenarios. Variability in expert emphasis and the exploratory nature of agreement analyses highlight the limitations of perception-based evaluation. Future studies should incorporate guideline-based validation, readability optimization, and patient-centered outcomes to more comprehensively evaluate the safety and suitability of LLM-generated information for geriatric patient education.
| Dokumententyp: | Artikel (Klinikum der LMU) |
|---|---|
| Organisationseinheit (Fakultäten): | 07 Medizin > Klinikum der LMU München > Medizinische Klinik und Poliklinik IV (Endokrinologie, Nephrologie, weitere Sektionen) |
| DFG-Fachsystematik der Wissenschaftsbereiche: | Lebenswissenschaften |
| Veröffentlichungsdatum: | 17. Aug 2026 05:02 |
| Letzte Änderung: | 17. Aug 2026 05:02 |
| URI: | https://oa-fund.ub.uni-muenchen.de/id/eprint/2577 |
![Publikation bearbeiten['Plugin/Screen:render_action_img_suffix' not defined] Publikation bearbeiten](/style/images/action_view.png)