Synapse
⌘+K
Synapse
PulseExploreClubsResearchersJournals
Instagram
HomeClubsExplore
September 18, 2025

Construction and Empirical Study of an Evaluation Dataset for Large Language Models in the Field of TCM Stroke (Preprint)

View Full Paper
Ask AI
Bookmark
Share

Authors

HLHongyan LongYDYang DengYGYaoguang Guo

Discussion

Loading...

Member takes

Overview

Empirical study highlights performance differences of large language models in TCM stroke evaluation.

Key Points

  • DeepSeek-R1 showed a 17 percentage point advantage in knowledge recall tasks compared to GPT-4o.
  • GPT-4o excelled in complex reasoning tasks, achieving a 90.5% scoring rate in interpreting classic texts.
  • The Traditional Chinese Medicine - Stroke Evaluation Dataset was systematically constructed for accurate assessment.
  • Large models trained on Chinese data have advantages in static knowledge, while general-purpose models excel in dynamic reasoning.

Cite This Study

Long et al. (2025) studied this question.

synapsesocial.com/papers/68d433a3713b0b5dfea72d2ehttps://doi.org/10.2196/preprints.81545
View Full Paper
Ask AI
Bookmark
Share

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Performance of Large Language Models versus Traditional Chinese Doctors in Migraine Diagnosis and Herbal Prescription: Approaching a Turing Point (Preprint)2025
  2. 2XuanHuGPT: parameter-efficient fine-tuning of large language model in the field of traditional Chinese medicine2025 · 2 citations
  3. 3Assessing the adherence of large language models to clinical practice guidelines in Chinese medicine: a content analysis2025 · 6 citations
  4. 4BianCang: A Traditional Chinese Medicine Large Language Model2025 · 25 citations
  5. 5GastroTCM: a large language model assistant for gastroenterology in traditional Chinese medicine2026 · 2 citations