AIHealthcare Analytics

Knowledge Wiki

Generative AI for Patient Education — Quality, Readability, and Accuracy Assessment

OVERVIEW
Rev 1 Jul 6, 2026 18:44 UTC 17 sources

Content

Accuracy and Misinformation Risk

An adversarial audit of five popular chatbots across 50 health questions found 49.6% of responses were problematic (30% somewhat, 19.6% highly), with no significant differences among platforms except Grok generating disproportionately more highly problematic responses. Performance was strongest for vaccines and cancer, weakest for stem cells, athletic performance, and nutrition. Reference quality was poor, with a median completeness score of 40% and hallucinations precluding any chatbot from producing a fully accurate reference list. For alcohol and breast cancer risk, variability in GenAI output was identified as a public health concern, with only 6% of outputs noting alcohol's Group 1 carcinogen status.

Readability Findings

ChatGPT-4o can simplify surgical patient education materials to approximately a 10.5 SMOG score (down from 12.1), though reaching the AMA-recommended sixth-grade level remains challenging due to medical complexity. Across studies, AI-generated patient education materials consistently exceed recommended reading levels. ChatGPT-5 scored higher than DeepSeek V3 on DISCERN, PEMAT-P, GQS, and CLEAR quality measures for foot and ankle disorders, while DeepSeek V3 produced simpler, more readable content. For knee osteoarthritis, a fine-tuned Gonarthrosis Advisor outperformed ChatGPT-5 on DISCERN quality scores.

Cardiovascular and Specialty Applications

In cardiovascular health, ChatGPT provided correct diagnostic hypotheses in 43% of cases, 5% of supplementary exam recommendations, and 10% of laboratory test recommendations compared to physician records. For cancer rehabilitation queries, ChatGPT-4 demonstrated mean accuracy of 3.93/5 but frequently lacked exercise dosage specifics and safety precautions. For Alzheimer's disease information, GenAI outputs exceeded recommended reading levels and exhibited uncertain, inauthentic, and negative tones.

Recommendations

A hybrid approach combining AI-assisted simplification with clinician review is recommended. Patient and caregiver involvement in evaluating AI-generated clinical communication tools is essential. Evaluation frameworks should extend beyond readability to include comprehensibility, relevance, usability, emotional impact, and empowerment potential.


Sources & Provenance

Source Article Evidence Harvested
harvested Generative artificial intelligence-driven chatbots and medical misinformation: an accuracy, referencing and readability audit. Peer-Reviewed 2026-07-06
harvested PubMed 42157 Peer-Reviewed 2026-07-06
harvested PubMed 42127 Peer-Reviewed 2026-07-06
harvested PubMed 42157 Peer-Reviewed 2026-07-06
harvested Enhancing Readability of Surgical Patient Education Resources Using ChatGPT. Peer-Reviewed 2026-07-06
harvested PubMed 42142 Peer-Reviewed 2026-07-06
harvested PubMed 42151 Peer-Reviewed 2026-07-06
harvested PubMed 42176 Peer-Reviewed 2026-07-06
harvested PubMed 42107 Peer-Reviewed 2026-07-06
harvested PubMed 42172 Peer-Reviewed 2026-07-06
harvested Content analysis of output from generative artificial intelligence chatbots when prompted about breast cancer and alcohol consumption. Peer-Reviewed 2026-07-06
harvested PubMed 42238 Peer-Reviewed 2026-07-06
harvested PubMed 42157 Peer-Reviewed 2026-07-06
harvested Comparison of ChatGPT-5 and DeepSeek V3 for Artificial Intelligence-Assisted Patient Education in Foot and Ankle Disorders. Peer-Reviewed 2026-07-06
harvested PubMed 42142 Peer-Reviewed 2026-07-06
harvested Gonarthrosis Advisor vs ChatGPT-5: quality and readability of artificial intelligence-generated patient education for knee osteoarthritis. Peer-Reviewed 2026-07-06
harvested Protein Language Models in Virology: A Review of Advances and Applications. Peer-Reviewed 2026-07-06

Related Pages

Revision History (1 revisions)
Rev 1 Jul 6, 2026 18:44 UTC
← Back to Wiki Index