Synapse
⌘+K
Synapse
PulseExploreClubsResearchersJournals
Instagram
HomeClubsExplore
August 22, 2025Frontiers in Artificial IntelligenceOpen Access

A multi-model longitudinal assessment of ChatGPT performance on medical residency examinations

View Full Paper
Ask AI
Bookmark
Share

Authors

MSMaria Eduarda Varela Cavalcanti SoutoAFAlexandre Chaves FernandesASAnne Emanuelle Cipriano da Silva

Discussion

Loading...

Member takes

Overview

Multi-model analysis reveals varying accuracy across medical areas in residency examinations, suggesting utility in education.

Key Points

  • ChatGPT-4 achieved 81.27% accuracy on medical residency exams, indicating strong performance for AI models in medical education.
  • GPT-4o outperformed GPT-4 with an accuracy of 85.88%, highlighting differences in performance among models.
  • Evaluation included 1,041 questions categorized by cognitive levels according to Bloom’s taxonomy, demonstrating the importance of assessing comprehension levels.
  • Careful integration of artificial intelligence in medical education is crucial, given ethical considerations and potential limitations in clinical practices.

Cite This Study

Souto et al. (2025) studied this question.

synapsesocial.com/papers/68af736e7567bf4f94fedbc0https://doi.org/10.3389/frai.2025.1614874
View Full Paper
Ask AI
Bookmark
Share

Also Consider

Synapse has enriched one closely related paper. Consider it for comparative context:

  1. 1Bloom’s taxonomy of cognitive learning objectives2015 · 1,039 citations