Synapse
⌘+K
Synapse
PulseExploreClubsResearchersJournals
Instagram
HomeClubsExplore
December 4, 2025Journal of Korean Academy of Psychiatric and Mental Health NursingOpen Access

Verification of the Validity and Reliability of Therapeutic Communication Responses Generated by Large Language Models (LLMs): A Comparative Study of Prompt-Engineering Strategies

View Full Paper
Ask AI
Bookmark
Share

Authors

GKGeun Myun Kim

Discussion

Loading...

Member takes

Overview

Methodological design assesses validity and reliability of generated responses in psychiatric nursing education, indicating the effectiveness of prompt strategies.

Key Points

  • Role-prompt yielded the highest content validity at 3.54±0.39, while inter-rater reliability showed only slight agreement (κ=0.07).
  • Semantic consistency was highest for Chain-of-Thought at 0.76±0.12; significant differences were found among strategies (p=.008).
  • Analysis utilized Python for comparing four prompt-engineering strategies on therapeutic communication responses from GPT-5.
  • These findings highlight the importance of optimal prompt strategies to enhance the validity and reliability of therapeutic outputs.

Cite This Study

Geun Myun Kim (2025) studied this question.

synapsesocial.com/papers/6930e8bdea1aef094cca326ahttps://doi.org/10.12934/jkpmhn.2025.34.s1.23
View Full Paper
Ask AI
Bookmark
Share

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Evaluation of large language models as a diagnostic tool for medical learners and clinicians using advanced prompting techniques2025 · 10 citations
  2. 2Evaluating Prompt Engineering Techniques for Accuracy and Confidence Elicitation in Medical LLMs2025
  3. 3AI meets psychology: an exploratory study of large language models’ competence in psychotherapy contexts2025 · 6 citations
  4. 4Prompt design and comparing large language models for healthcare simulation case scenarios2025
  5. 5Impact of Prompt Engineering on the Performance of ChatGPT Variants Across Different Question Types in Medical Student Examinations (Preprint)2025