Synapse
⌘+K
Synapse
PulseExploreClubsResearchersJournals
Instagram
HomeClubsExplore
December 8, 2025Blood

Evaluating large language models in real-world hematologic clinical decision-making: Performance, limitations, and clinical implications

View Full Paper
Ask AI
Bookmark
Share

Authors

ACAngela ConsagraJWJiasheng WangGRGustavo Rivero

Discussion

Loading...

Member takes

Overview

Observational analysis evaluated diagnostic accuracy in hematology, suggesting LLMs need more rigorous testing for clinical use.

Key Points

  • GPT-o3 achieved 58% agreement with expert assessments in hematology, indicating room for improvement in AI performance.
  • Expert ratings showed a consistent average score for GPT-o3 at 3.48 out of 5, underscoring its diagnostic relevance.
  • Analysis used 30 complex myelodysplastic syndrome cases to assess LLM capabilities in clinical decision-making.
  • Despite advancements, frequent hallucinations were observed, highlighting significant limitations in current AI models for specialized care.

Cite This Study

Consagra et al. (2025) studied this question.

synapsesocial.com/papers/69362f6e4fa91c937236e136https://doi.org/10.1182/blood-2025-4349
View Full Paper
Ask AI
Bookmark
Share

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Performance of different large language models (LLMs) as decision support tools across various hematologic malignancies2025
  2. 2Artificial intelligence‑driven virtual tumor board enhances precision care in myelodysplastic syndromes (MDS)2025
  3. 3Evaluating artificial intelligence (AI) as a clinical decision support tool for AML patients2025 · 1 citations
  4. 4Clinical Assessment of Large Language Models: A Comprehensive Multi-domain Performance Study for Healthcare Applications2025 · 1 citations
  5. 5Application of Large Language Models in Complex Clinical Cases: Cross-Sectional Evaluation Study2025 · 6 citations