Pilot study compares moral evaluations of language models and humans, suggesting differences in machine psychology.
As an initial pilot study, we examined the moral evaluations of n = 8 large language models (LLMs) and n = 19 human subjects, comparing their qualitative responses to six case vignettes requiring moral courage. We also compared human subjects' and LLM responses to better understand "machine psychology" or "machine behavior" in analysing and assessing situations that require complex moral evaluations. We conducted detailed psycholinguistic analyses using the Linguistic Inquiry and Word Count in its current 2022 version. Responses from LLMs with high ELO ratings used over 1.5 times more power related terms compared to LLMs with low ELO ratings (p < 0.001). We found strong evidence of a lack of "behavioral similarity" in several dimensions. LLMs used over 1.5 times more achievement (p = 0.04, d = 0.95) and power related terms (p < 0.01, d = 1.18) than humans. They also showed almost twice as many terms related to moral emotions (p = 0.02, d = 1.08) and benevolence (p = 0.01, d = 1.11). LLMs also used over 3.5 times more terms related to universalism (p < 0.000, d = 2.44). Results validate a cautious approach to any presumed equivalence of human and LLM evaluations.Responses from LLMs with high ELO ratings used over 1.5 times more power related terms compared to LLMs with low ELO ratings (p < 0.001). We found strong evidence of a lack of "behavioral similarity" in several dimensions. LLMs used over 1.5 times more achievement (p = 0.04, d = 0.95) and power related terms (p < 0.01, d = 1.18) than humans. They also showed almost twice as many terms related to moral emotions (p = 0.02, d = 1.08) and benevolence (p = 0.01, d = 1.11). LLMs also used over 3.5 times more terms related to universalism (p < 0.000, d = 2.44).
No takes yet. Share an insight, caveat, or question.
Klein et al. (2025) studied this question.