Analysis shows that large language models yield good concordance with expert recommendations for hematologic malignancies, suggesting potential for decision support tools.
Key Points
Decision support tools exhibited good concordance with expert recommendations in hematologic malignancies, particularly non-Hodgkin lymphoma.
Aggregate competence scores ranged from 849 to 964 out of 1140 across 38 complex cases evaluated by large language models.
Assessment involved 3 large language models, with varying performance highlighted for specific cancer types such as multiple myeloma and non-Hodgkin lymphoma.
Findings indicate that language models may provide significant support in malignant hematology, emphasizing the need for careful human oversight.