Generative AI for Patient Education — Quality, Readability, and Accuracy Assessment
OVERVIEWContent
Accuracy and Misinformation Risk
An adversarial audit of five popular chatbots across 50 health questions found 49.6% of responses were problematic (30% somewhat, 19.6% highly), with no significant differences among platforms except Grok generating disproportionately more highly problematic responses. Performance was strongest for vaccines and cancer, weakest for stem cells, athletic performance, and nutrition. Reference quality was poor, with a median completeness score of 40% and hallucinations precluding any chatbot from producing a fully accurate reference list. For alcohol and breast cancer risk, variability in GenAI output was identified as a public health concern, with only 6% of outputs noting alcohol's Group 1 carcinogen status.
Readability Findings
ChatGPT-4o can simplify surgical patient education materials to approximately a 10.5 SMOG score (down from 12.1), though reaching the AMA-recommended sixth-grade level remains challenging due to medical complexity. Across studies, AI-generated patient education materials consistently exceed recommended reading levels. ChatGPT-5 scored higher than DeepSeek V3 on DISCERN, PEMAT-P, GQS, and CLEAR quality measures for foot and ankle disorders, while DeepSeek V3 produced simpler, more readable content. For knee osteoarthritis, a fine-tuned Gonarthrosis Advisor outperformed ChatGPT-5 on DISCERN quality scores.
Cardiovascular and Specialty Applications
In cardiovascular health, ChatGPT provided correct diagnostic hypotheses in 43% of cases, 5% of supplementary exam recommendations, and 10% of laboratory test recommendations compared to physician records. For cancer rehabilitation queries, ChatGPT-4 demonstrated mean accuracy of 3.93/5 but frequently lacked exercise dosage specifics and safety precautions. For Alzheimer's disease information, GenAI outputs exceeded recommended reading levels and exhibited uncertain, inauthentic, and negative tones.
Recommendations
A hybrid approach combining AI-assisted simplification with clinician review is recommended. Patient and caregiver involvement in evaluating AI-generated clinical communication tools is essential. Evaluation frameworks should extend beyond readability to include comprehensibility, relevance, usability, emotional impact, and empowerment potential.