Synapse
⌘+K
Synapse
PulseExploreClubsResearchersJournals
Instagram
HomeClubsExplore
August 24, 2025Open Access

Evaluating Large Language Model Diagnostic Performance on JAMA Clinical Challenges via a Multi-Agent Conversational Framework

View Full Paper
Ask AI
Bookmark
Share

Authors

KSKarl L. SangwonJZJeff ZhangRSRobert Steele

Discussion

Loading...

Member takes

Overview

Assessment of diagnostic accuracy in multi-agent setups suggests that benchmarks may overestimate clinical reasoning.

Key Points

  • Accuracy dropped significantly across all models when diagnostic formats shifted from vignettes to conversations, and from multiple-choice to free-response, highlighting limitations.
  • In multiple-choice vignettes, O1 achieved 79.8% accuracy, while in conversation free-response, it fell to 31.7%; similar trends were observed in other models tested.
  • The approach involved adapting 815 diagnostic cases into two formats, utilizing AI agents for subjective history and objective data during multi-agent conversations.
  • The findings suggest that current static benchmarks inadequately reflect clinical reasoning, supporting the need for open-source frameworks to improve evaluations.

Cite This Study

Sangwon et al. (2025) studied this question.

synapsesocial.com/papers/68af79ab7567bf4f94ff18a9https://doi.org/10.1101/2025.08.20.25334087
View Full Paper
Ask AI
Bookmark
Share