Systematic review and meta-analysis reveal ChatGPT-4 outperforms ChatGPT-3.5 in medical licensing exams, indicating strong AI capabilities in medical education.
Key Points
ChatGPT-4 achieved an impressive accuracy of 81.8% in medical licensing exams, surpassing ChatGPT-3.5's 60.8% accuracy.
In in-training residency exams, ChatGPT-4's accuracy was 72.2%, compared to 57.7% for ChatGPT-3.5, highlighting a significant gap.
The study included 53 articles to assess ChatGPT's performance across various specialties, putting AI's capabilities under scrutiny.
Findings suggest ChatGPT-4 can enhance medical education but still cannot fully replace the necessary skills in medical practice.