Performance of several large language models when answering common patient questions about type 1 diabetes in children: accuracy, comprehensibility and practicality
Cross-sectional comparative analysis reveals large language models can provide reliable answers in pediatric diabetes, suggesting their potential role in healthcare.
Key Points
Large language models showed similar performance when answering common questions about type 1 diabetes in children.
ChatGPT-4o achieved the highest mean score of 3.78 ± 1.09, while Gemini scored the lowest at 3.40 ± 1.24.
The evaluation used a standard prompt and was assessed by pediatric endocrinologists using the General Quality Scale.
Despite no significant differences, the study highlights the promise of advanced models in providing patient-friendly answers.