Synapse
⌘+K
Synapse
PulseExploreClubsResearchersJournals
Instagram
HomeClubsExplore
September 5, 2025JMIRx MedOpen Access

Rapidly Benchmarking Large Language Models for Diagnosing Comorbid Patients: Comparative Study Leveraging the LLM-as-a-Judge Method

View Full Paper
Ask AI
Bookmark
Share

Authors

PSPeter SarvariZAZaid Al-fagih

Discussion

Loading...

Member takes

Overview

Comparative study assesses diagnostic abilities of 21 large language models, highlighting LLM-as-a-judge method and retrieval-augmented generation.

Key Points

  • Gemini 2.5 achieved the highest diagnostic hit rate of 97.4%, outperforming other large language models.
  • The analysis involved 18 LLMs on 1000 MIMIC-IV hospital admissions, using prompts and temperature settings.
  • Retrieval-augmented generation improved the hit rate of GPT-4o 05-13 by an average of 0.8%, indicating better diagnostic accuracy.
  • Significant model performance variation was noted across different prompts, suggesting the need for diverse datasets.

Cite This Study

Sarvari et al. (2025) studied this question.

synapsesocial.com/papers/68c2384fb210217d64776a63https://doi.org/10.2196/67661
View Full Paper
Ask AI
Bookmark
Share