Synapse
⌘+K
Synapse
PulseExploreClubsResearchersJournals
Instagram
HomeClubsExplore
September 10, 2025International Journal of Medical InformaticsOpen Access

Diagnostic performance of newly developed large language models for critical illness cases: A comparative study

View Full Paper
Ask AI
Bookmark
Share

Authors

XWXing WuYHYu HuangQHQing He

Discussion

Loading...

Member takes

Overview

Comparative study assesses diagnostic accuracy and response quality of large language models in ICU, suggesting strong potential for clinical support.

Key Points

  • ChatGPT-o3 achieved the highest diagnostic accuracy of 72%, outperforming other models in critical illness cases.
  • The study compared four large language models, finding ChatGPT-o3, DeepSeek-R1, and ChatGPT-4o significantly better than DeepSeek-V3 in diagnostic performance.
  • In a cross-sectional study, 50 critical illness cases evaluated the models' diagnostic accuracy and response quality in intensive care unit settings.
  • Significant trends indicate that focusing on domain-specific fine-tuning could enhance the models' diagnostic capabilities in clinical settings.

Cite This Study

Wu et al. (2025) studied this question.

synapsesocial.com/papers/68c24317b210217d647a55c2https://doi.org/10.1016/j.ijmedinf.2025.106088
View Full Paper
Ask AI
Bookmark
Share