Authors
Loading...
Comparative study assesses diagnostic abilities of 21 large language models, highlighting LLM-as-a-judge method and retrieval-augmented generation.
Sarvari et al. (2025) studied this question.