Generative AI as a Research Tool — Systematic Reviews, Literature Analysis, and Research Support
OVERVIEWContent
Systematic Review Screening
ChatGPT 5.0 outperformed ChatGPT 4.0 in full-text screening for systematic reviews (accuracy 0.77 vs. 0.71; κ=0.55 vs. 0.43), approaching but not matching human performance. In a comparative study of four evidence synthesis projects, ChatGPT Plus (GPT 4.1) showed high agreement with human reviewers for systematic reviews with well-defined criteria (κ=0.73–0.86, sensitivity ≥0.88) but only moderate agreement for scoping and narrative reviews (κ=0.56–0.59). AI-assisted screening is recommended as a triage aid rather than a replacement for manual screening, particularly in exploratory contexts.
Qualitative Thematic Analysis
Generative AI demonstrated meaningful alignment with human thematic analysis in surgical education research, with substantial conceptual overlap and no unique or contradictory themes identified by AI. AI-generated themes tended to integrate individual and group-level constructs into broader relational themes, while human analysts generated more discrete themes. A modular human-AI pipeline for thematic analysis across digital health interview studies found Anthropic Claude (Opus) produced the most consistently human-aligned themes, while Gemini and ChatGPT showed different trade-offs in speed, fidelity, and usability.
Citation Accuracy
Across studies, AI citation accuracy is substantially problematic. DeepSeek-R1 achieved the highest citation accuracy (78.6%) in anterior segment ophthalmology research, followed by ChatGPT and Copilot (51.4% each) and Gemini (12.9%). In a general health misinformation audit, chatbot hallucinations and fabricated citations precluded any chatbot from producing a fully accurate reference list (median reference completeness 40%). AI-assisted citation generation requires rigorous human verification.
AI-Assisted Teams vs. Human Teams
An experiment randomizing 288 researchers to 103 teams found AI-assisted teams (94% vs. 91% reproduction rates) performed comparably to human-only teams on most outcomes, but human-only teams identified significantly more major coding errors. AI-led teams achieved only a 37% reproduction rate, detected fewer errors, and proposed weaker robustness checks. Expert human judgment currently remains indispensable for reliable empirical verification.
Updates to Existing Pages
Sources & Provenance
| Source | Article | Evidence | Harvested |
|---|---|---|---|
| harvested | PubMed 41956 | Peer-Reviewed | 2026-07-06 |
| harvested | PubMed 42020 | Peer-Reviewed | 2026-07-06 |
| harvested | PubMed 42023 | Peer-Reviewed | 2026-07-06 |
| harvested | PubMed 42033 | Peer-Reviewed | 2026-07-06 |
| harvested | PubMed 42100 | Peer-Reviewed | 2026-07-06 |
| harvested | PubMed 42056 | Peer-Reviewed | 2026-07-06 |
| harvested | PubMed 42133 | Peer-Reviewed | 2026-07-06 |
| harvested | PubMed 42148 | Peer-Reviewed | 2026-07-06 |
| harvested | PubMed 42117 | Peer-Reviewed | 2026-07-06 |
| harvested | PubMed 42181 | Peer-Reviewed | 2026-07-06 |
| harvested | PubMed 42228 | Peer-Reviewed | 2026-07-06 |