AIHealthcare Analytics

Knowledge Wiki

Generative AI as a Research Tool — Systematic Reviews, Literature Analysis, and Research Support

OVERVIEW
Rev 1 Jul 6, 2026 18:44 UTC 11 sources

Content

Systematic Review Screening

ChatGPT 5.0 outperformed ChatGPT 4.0 in full-text screening for systematic reviews (accuracy 0.77 vs. 0.71; κ=0.55 vs. 0.43), approaching but not matching human performance. In a comparative study of four evidence synthesis projects, ChatGPT Plus (GPT 4.1) showed high agreement with human reviewers for systematic reviews with well-defined criteria (κ=0.73–0.86, sensitivity ≥0.88) but only moderate agreement for scoping and narrative reviews (κ=0.56–0.59). AI-assisted screening is recommended as a triage aid rather than a replacement for manual screening, particularly in exploratory contexts.

Qualitative Thematic Analysis

Generative AI demonstrated meaningful alignment with human thematic analysis in surgical education research, with substantial conceptual overlap and no unique or contradictory themes identified by AI. AI-generated themes tended to integrate individual and group-level constructs into broader relational themes, while human analysts generated more discrete themes. A modular human-AI pipeline for thematic analysis across digital health interview studies found Anthropic Claude (Opus) produced the most consistently human-aligned themes, while Gemini and ChatGPT showed different trade-offs in speed, fidelity, and usability.

Citation Accuracy

Across studies, AI citation accuracy is substantially problematic. DeepSeek-R1 achieved the highest citation accuracy (78.6%) in anterior segment ophthalmology research, followed by ChatGPT and Copilot (51.4% each) and Gemini (12.9%). In a general health misinformation audit, chatbot hallucinations and fabricated citations precluded any chatbot from producing a fully accurate reference list (median reference completeness 40%). AI-assisted citation generation requires rigorous human verification.

AI-Assisted Teams vs. Human Teams

An experiment randomizing 288 researchers to 103 teams found AI-assisted teams (94% vs. 91% reproduction rates) performed comparably to human-only teams on most outcomes, but human-only teams identified significantly more major coding errors. AI-led teams achieved only a 37% reproduction rate, detected fewer errors, and proposed weaker robustness checks. Expert human judgment currently remains indispensable for reliable empirical verification.


Updates to Existing Pages

Sources & Provenance

Source Article Evidence Harvested
harvested PubMed 41956 Peer-Reviewed 2026-07-06
harvested PubMed 42020 Peer-Reviewed 2026-07-06
harvested PubMed 42023 Peer-Reviewed 2026-07-06
harvested PubMed 42033 Peer-Reviewed 2026-07-06
harvested PubMed 42100 Peer-Reviewed 2026-07-06
harvested PubMed 42056 Peer-Reviewed 2026-07-06
harvested PubMed 42133 Peer-Reviewed 2026-07-06
harvested PubMed 42148 Peer-Reviewed 2026-07-06
harvested PubMed 42117 Peer-Reviewed 2026-07-06
harvested PubMed 42181 Peer-Reviewed 2026-07-06
harvested PubMed 42228 Peer-Reviewed 2026-07-06

Related Pages

Revision History (1 revisions)
Rev 1 Jul 6, 2026 18:44 UTC
← Back to Wiki Index