Blog
AI in Research #4: 5 Recent Studies, 5-Minute Journal Club
⏱️ Reading time: ~5 minutes
Our featured five studies today don’t settle whether AI can replace doctors. They show the conditions under which it helps, falls short, or becomes risky. The real question isn’t whether AI is “good”, but for which task, in which population, and under what oversight it becomes clinically useful.
🩺 Can an AI Find Diagnoses That Doctors Missed the First Time?
Jaech et al. | NEJM AI | 18 Jun 2026 | LLM-Assisted Reanalysis of Unsolved Rare Disease Genomes Increases Diagnostic Yield
Study design: Retrospective reanalysis of old, unsolved cases
Participants: 376 rare-disease genetic tests with no prior answer, across four disease groups (developmental disorders, neuromuscular disease, sudden unexpected death, early psychosis)
Task: An AI tool re-read each old case’s notes, symptoms, and gene variants and proposed candidate diagnoses; genetics experts then reviewed, tested, and confirmed which ones held up.
Of the AI’s suggestions, 18 of 376 cases (4.8% overall) were confirmed as new diagnoses after expert review, ranging from 1.0% to 13.3% depending on disease type. The AI proposed candidates; it did not diagnose on its own.
Why it matters AI-assisted re-review can recover diagnoses missed the first time, adding value to old tests rather than replacing the original testing or expert confirmation.
Key takeaway Yield varies a lot by disease type. Don’t treat the 4.8% as one fixed number.
🩺 Does AI Beat Doctors at Reading Real Skin Photos, Not Just Textbook Images?
Anriot et al. | JAMA Dermatology | 3 Jun 2026 | Limits of Artificial Intelligence Models for Skin Cancer Diagnosis in Realistic Settings
Study design: Real-world diagnostic accuracy study
Participants: 1,117 real skin lesion photos; 652 physicians; three AI systems, including a newer model called PanDerm
Task: Physicians and AI systems judged the same real-world skin photos for skin cancer.
Physicians averaged 65.9% correct. The best AI model reached 72.2%, above that overall physician average, but experienced dermatologists alone still outperformed AI (up to 74.2%).
Why it matters This AI model beat the average physician across all experience levels, but not the experienced dermatologists specifically. The finding depends entirely on which comparison group is used: “AI beats doctors” is only true against a mixed-experience average, not against specialists.
⚠️ Limitations
The AI missed some dangerous cancer types, including acral melanoma.
Most doctors and photos were French and of European ancestry, with darker skin tones underrepresented, limiting how well results generalize.
Key takeaway Foundation models beat mid-level clinicians but not experts, and the study’s own population skew limits how far that finding travels.
🩺 Are Doctors Already Using AI Tools Their Hospital Hasn’t Approved?
Petersson et al. | JMIR | 2026 | Shadow AI in Swedish Health Care: Qualitative Analysis of Physicians’ Free-Text Answers
Study design: Cross-sectional survey with qualitative content analysis of free-text answers
Participants: 357 physicians across Swedish healthcare organizations (~64% response rate), surveyed December 2023 to January 2024
Task: Physicians answered open-ended questions about how AI has changed their work and what shapes their decision to use it.
Analyzing the free-text responses, researchers found that much of this AI use involved tools that were never formally procured, certified, or approved for clinical use — mainly ChatGPT, accessed through personal accounts. Physicians described reaching for these unapproved tools on their own initiative, usually to save time or fill a gap their approved systems didn’t cover, not out of carelessness.
Why it matters Unofficial AI use isn’t a rare edge case; it’s already part of daily practice, often ahead of hospital policy.
Key takeaway This is a Swedish snapshot from December 2023–January 2024; The underlying gap: staff needs outpacing official AI policy, plausibly extends elsewhere including Switzerland.
🩺 How Often Could Following AI’s Medical Advice Actually Cause Harm?
Wu et al. | arXiv preprint | Dec 2025, updated Jul 2026 | First, do NOHARM: a medical safety benchmark and randomized study of physician and AI teaming on clinical consultations
Study design: A safety benchmark, plus a separate randomized study
Participants: 20 AI chat models and 4 dedicated medical AI tools, tested on 1,100 written, simulated case scenarios modeled on real specialist referrals, not live patients. Separately, 101 US physicians each worked through cases under three conditions — no AI, a provided AI assistant, and free choice of any resource — in randomized, counterbalanced order.
Task: Measured how often following the AI’s advice unedited, on these simulated cases, could have caused harm, and whether AI assistance improved doctors’ performance.
Following AI advice directly carried a chance of serious harm in up to 24.6% of the simulated cases for weaker systems, over 80% of it from things the AI left out rather than wrong information stated. Doctors using AI help scored better than those using usual resources, but often ignored good AI suggestions, so they still scored below the AI alone.
Why it matters Unedited AI advice carries a real, measurable risk. Doctors and AI did best together, but only when the doctor actually used the AI’s good suggestions.
Key takeaway This is a preprint using written case scenarios, not live patients.
🩺 Do Doctors and AI Even Agree on What Makes an AI’s Answer Good?
Shi et al. | npj Digital Medicine | 2 Jul 2026 | Physicians and artificial intelligence diverge in evaluating large language models on real clinical cases
Study design: Multicenter evaluation study
Participants: More than 400 physicians across seven specialties and several countries rated real AI-generated answers to real cases. Researchers also built AI “judge” agents matched to each physician’s specialty and experience level, to see if AI could take over the grading.
Task: Physicians and their matched AI judges rated the same answers to real, de-identified clinical cases.
Which AI model looked “best” changed depending on who scored it: senior versus junior physicians, or different practice settings, often ranked the same answers differently. The AI judges tracked the general direction of physician scores but missed finer calls, like whether an answer’s reasoning fit the clinical context.
Why it matters Grading AI with AI isn’t a reliable shortcut. “Physicians rated this AI highly” depends heavily on which physicians were asked.
Key takeaway Physician panels don’t fully agree with each other either, so no single panel is a fixed measuring stick for an AI tool’s quality.
📌 Bottom line
None of this argues for enthusiasm or rejection. AI can recover missed rare-disease diagnoses on reanalysis, and modern skin-cancer models now beat mid-career dermatologists — but not experts, and not on the cancer subtypes or skin tones underrepresented in training data. Physicians are already using unapproved AI daily to fill real gaps, often for good clinical reasons, not out of carelessness. Meanwhile, taking AI advice at face value carries a measurable risk of harm, mostly from what it leaves out rather than what it gets wrong — and even AI’s own judgment of “good” AI output doesn’t reliably match what physicians think. The pattern across all five: AI adds real value where research, reanalysis, triage, and second opinions are concerned, but safety, generalizability, and clinical accountability still depend on a physician staying in the loop.
📚 Sources
Jaech A, et al. LLM-Assisted Reanalysis of Unsolved Rare Disease Genomes Increases Diagnostic Yield. NEJM AI. 2026. DOI: 10.1056/AIcs2501343
Anriot J, Yan S, Coste C, et al. Limits of Artificial Intelligence Models for Skin Cancer Diagnosis in Realistic Settings. JAMA Dermatology. 2026. DOI: 10.1001/jamadermatol.2026.1492
Petersson L, Irgang L, Mauritzon I, Holmén M. Shadow AI in Swedish Health Care: Qualitative Analysis of Physicians’ Free-Text Answers. J Med Internet Res. 2026;28:e93484. DOI: 10.2196/93484
Wu D, Nateghi Haredasht F, et al. First, do NOHARM: a medical safety benchmark and randomized study of physician and AI teaming on clinical consultations. arXiv preprint. 2025 (updated Jul 2026). arXiv:2512.01241
Shi P, Li J, Yang Z, et al. Physicians and artificial intelligence diverge in evaluating large language models on real clinical cases. npj Digital Medicine. 2026. DOI: 10.1038/s41746-026-02942-6
Liked this? Get new articles in your inbox.
Subscribe