NOHARM — Stanford-Harvard Clinical AI Safety Benchmark
ENTITYContent
Overview
NOHARM (Numerous Options Harm Assessment for Risk in Medicine) is a clinical AI safety benchmark developed by researchers at Stanford University School of Medicine and Harvard Medical School. The benchmark was originally posted to arXiv in December 2025 and subsequently updated, with findings widely reported in mid-2026.
Methodology and Scope
NOHARM evaluated 24 AI systems — spanning general-purpose large language models and medical-specialized tools — on their clinical safety performance. The benchmark focuses specifically on the potential for AI systems to produce harmful medical recommendations, distinguishing it from accuracy-focused benchmarks that do not weight patient safety outcomes.
Key Findings
Four medical-specialized AI tools topped the rankings in a statistical tie, outperforming general-purpose models on safety metrics. The study provided empirical support for the argument that domain-specialized clinical AI carries meaningfully lower patient safety risk than general-purpose chatbots deployed in healthcare contexts. The results were cited in connection with OpenEvidence's reported physician preference advantage over competing systems.
Significance
NOHARM represents one of the most comprehensive independent safety evaluations of clinical AI to date, and its findings have been referenced in discussions around AI procurement standards, reimbursement policy, and regulatory frameworks for AI-enabled medical tools.