Synapse
⌘+K
Synapse
PulseExploreClubsResearchersJournals
Instagram
HomeClubsExplore
September 23, 2025Open Access

Large language models for automatable real-world performance monitoring of diagnostic decision support systems: a comparison to manual doctor panel review in a prospective clinical study

View Full Paper
Ask AI
Bookmark
Share

Authors

FCFabienne CotteMSMarcel SchmudePBPhilipp Bode

Discussion

Loading...

Member takes

Overview

Prospective clinical study compares large language models for diagnostic accuracy with manual review, suggesting significant potential for automated monitoring.

Key Points

  • Large language models can reproduce clinician review of diagnostic encounters with close agreement, enabling efficient monitoring.
  • GPT-5 achieved an accuracy of 84.7% in classifying encounters, indicating high sensitivity and moderate specificity compared to manual reviews.
  • The eligibility filtering step in diagnostic accuracy analysis was identified as the main source of divergence needing improvement.
  • Embedding large language model approaches could facilitate automated performance monitoring and regulatory compliance across healthcare systems.

Cite This Study

Cotte et al. (2025) studied this question.

synapsesocial.com/papers/68d440a6713b0b5dfea7feb9https://doi.org/10.1101/2025.09.20.25336227
View Full Paper
Ask AI
Bookmark
Share

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Rapidly Benchmarking Large Language Models for Diagnosing Comorbid Patients: Comparative Study Leveraging the LLM-as-a-Judge Method2025 · 1 citations
  2. 2Evaluating large language models in real-world hematologic clinical decision-making: Performance, limitations, and clinical implications2025
  3. 3Performance Comparison of Human Doctors and Large Language Models in Tuberculosis Triage, Diagnosis, and Management:An Experimental Study (Preprint)2025
  4. 4Large Language Model Symptom Identification From Clinical Text: Multicenter Study2025
  5. 5Clinical Assessment of Large Language Models: A Comprehensive Multi-domain Performance Study for Healthcare Applications2025 · 1 citations