Assessment of diagnostic accuracy in multi-agent setups suggests that benchmarks may overestimate clinical reasoning.
Key Points
Accuracy dropped significantly across all models when diagnostic formats shifted from vignettes to conversations, and from multiple-choice to free-response, highlighting limitations.
In multiple-choice vignettes, O1 achieved 79.8% accuracy, while in conversation free-response, it fell to 31.7%; similar trends were observed in other models tested.
The approach involved adapting 815 diagnostic cases into two formats, utilizing AI agents for subjective history and objective data during multi-agent conversations.
The findings suggest that current static benchmarks inadequately reflect clinical reasoning, supporting the need for open-source frameworks to improve evaluations.