Observational analysis evaluates triage accuracy in the emergency department using large language models, indicating a promising future for AI support.
Key Points
Gemini 2.5 flash achieved the highest triage accuracy at 73.8%, demonstrating strong potential for AI in emergency care.
A total of 1,057 triage conversations were analyzed, revealing significant variations in model performance across different LLMs.
Using both zero-shot and few-shot prompting improved outcomes, highlighting the flexibility and adaptability of LLMs in clinical situations.
The findings support the integration of LLMs for non-critical triage, benefiting patient care in diverse clinical environments.