Synapse
⌘+K
Synapse
PulseExploreClubsResearchersJournals
Instagram
HomeClubsExplore
October 11, 2025Journal of Educational Evaluation for Health ProfessionsOpen Access

Performance of GPT-4o and o1-Pro on United Kingdom Medical Licensing Assessment-style items: a comparative study

View Full Paper
Ask AI
Bookmark
Share

Authors

BVBehrad VakiliAAAnizana AhmadMZMahsa Zolfaghari

Discussion

Loading...

Member takes

Overview

This analysis reveals accuracy differences in AI models for the UK Medical Licensing Assessment, suggesting development areas.

Key Points

  • GPT-4o achieved an accuracy of 88.8%, while o1-Pro outperformed with 93.0%, indicating clear model quality variance.
  • Statistical analysis using McNemar’s test confirmed the significant performance advantage for o1-Pro across the evaluated items.
  • Specialized medical fields revealed variable results, with both models excelling in surgery and psychiatry, but differences in dermatology and imaging.
  • Despite high overall scores, isolated weaknesses in general practice point to the need for traditional study methods alongside AI.

Cite This Study

Vakili et al. (2025) studied this question.

synapsesocial.com/papers/68e9b1d9ba7d64b6fc13306chttps://doi.org/10.3352/jeehp.2025.22.30
View Full Paper
Ask AI
Bookmark
Share