Comparative study assesses diagnostic accuracy and response quality of large language models in ICU, suggesting strong potential for clinical support.
Key Points
ChatGPT-o3 achieved the highest diagnostic accuracy of 72%, outperforming other models in critical illness cases.
The study compared four large language models, finding ChatGPT-o3, DeepSeek-R1, and ChatGPT-4o significantly better than DeepSeek-V3 in diagnostic performance.
In a cross-sectional study, 50 critical illness cases evaluated the models' diagnostic accuracy and response quality in intensive care unit settings.
Significant trends indicate that focusing on domain-specific fine-tuning could enhance the models' diagnostic capabilities in clinical settings.