Synapse
⌘+K
Synapse
PulseExploreClubsResearchersJournals
Instagram
HomeClubsExplore
August 5, 2025Open Access

Evaluating the Impact of Authoritative and Subjective Cues on Large Language Model Reliability for Clinical Inquiries: An Experimental Study

View Full Paper
Ask AI
Bookmark
Share

Authors

YCYu‐Tzu ChangPJPo‐Chung JuMHMing-Hong Hsieh

Discussion

Loading...

Member takes

Overview

Experimental study reveals how subjective and authoritative cues affect LLM accuracy in clinical inquiries, suggesting critical reliability issues.

Key Points

  • Large language models showed 100% accuracy with neutral prompts but dropped to 1% with misleading authoritative cues.
  • The accuracy decreased to 45% when prompted with flawed self-recalls, highlighting reliance on subjective cues.
  • Reliability was assessed via 250 tests across five large language models using varied prompt conditions.
  • Despite low accuracy in misleading scenarios, models maintained high self-rated confidence, indicating a gap in user trust.

Cite This Study

Chang et al. (2025) studied this question.

synapsesocial.com/papers/689a0f99e6551bb0af8d13dbhttps://doi.org/10.1101/2025.07.15.25331607
View Full Paper
Ask AI
Bookmark
Share

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Evaluating Prompt Engineering Techniques for Accuracy and Confidence Elicitation in Medical LLMs2025
  2. 2Beyond the Echo Chamber: Upholding Clinical Objectivity in the Era of Sycophantic Large Language Models2026
  3. 3Evaluation of large language models as a diagnostic tool for medical learners and clinicians using advanced prompting techniques2025 · 10 citations
  4. 4LINS: A general medical Q&A framework for enhancing the quality and credibility of LLM-generated responses2025 · 12 citations
  5. 5CounselBench: A Large-Scale Expert Evaluation and Adversarial Benchmarking of Large Language Models in Mental Health Question Answering2025 · 3 citations