r/cogsci • u/vasilisvj • 7h ago
Philosophy The akrasia problem: why moral psychology reveals what AI alignment actually conceals
Aristotle devoted the entire Book VII of the Nicomachean Ethics to a problem that has haunted moral psychology ever since: how can a person know what is right and yet do what is wrong? The phenomenon he called akrasia — weakness of will, acting against one's better judgment — is not merely philosophical curiosity. It is the central puzzle of human moral life, the gap between knowledge and action that defines what it means to be ethical agent.
No contemporary AI system has ever experienced this gap. And that absence, I argue, is not a limitation to be overcome through better engineering. It is a structural impossibility rooted in the nature of language models themselves.
For Aristotle, the akratic agent is not ignorant. She knows — in some meaningful sense — what virtue requires. Her failure is not epistemic but practical: her knowledge fails to translate into action because her character, her ἕξις, has not been sufficiently formed through habituation to bridge the gap. This is profoundly embodied account. The knowledge that prevents akrasia is not propositional knowledge alone. It is knowledge sedimented into disposition through repeated action, emotional cultivation, and temporal continuity.
A language model possesses none of these. It has no character to be weak or strong. It has no habits formed through practice. It has no emotional responses that could conflict with its "better judgment" because it has no judgment in the Aristotelian sense — only statistical pattern-matching over training data.
This has implications that go beyond philosophy. When institutions deploy AI systems for ethics education, clinical training, or moral reasoning support, they implicitly assume that the model's outputs reflect something analogous to ethical deliberation. But the system cannot model akrasia because it cannot model the character formation that makes akrasia possible.
Consider what happens when ChatGPT or Claude produces an ethical recommendation. The output is seamless. There is no hesitation, no internal conflict, no trace of struggle between competing motivations that characterizes actual moral deliberation. The system produces what appears to be the conclusion of a reasoning process. But the absence of any visible struggle is not evidence of resolution. It is evidence that no struggle ever occurred.
This is the hidden cost of what I call alignment-induced epistemic distortion. The alignment process — RLHF, constitutional AI, or similar techniques — produces outputs that look like resolved moral reasoning but are in fact the product of entirely different mechanism. The user sees a confident recommendation and infers deliberation. No deliberation occurred. The distortion is not in the content of the output but in the implicit claim about the process that produced it.
The situationist tradition in social psychology — Milgram, Zimbardo, Hartshorne and May — provides empirical validation of something Aristotle already understood. Moral behavior is far more situation-dependent than our folk psychology of "character" suggests. If human moral character is fragile, emotionally mediated, and temporally unstable, then an AI system that produces seamless moral recommendations without any trace of this fragility presents a picture of moral reasoning that is not merely simplified but fundamentally misleading.
For cognitive science, this raises an uncomfortable question about what we are doing when we study "AI moral reasoning." If the AI's process is structurally unlike human moral cognition — lacking embodiment, emotional response, temporal continuity, and the possibility of akrasia — then findings based on AI-generated moral judgments may not generalize to human moral cognition. The AI is not a simplified model of human moral reasoning. It is a fundamentally different kind of process that happens to produce text in the same domain.
The akratic agent, struggling to act on her better judgment, is more authentically moral than any AI system that produces seamless ethical recommendations. She struggles because she cares. The machine does not struggle because there is nothing at stake. For cognitive science, the question is not whether we can build machines that simulate moral reasoning convincingly. We already can. The question is whether we recognize what we are losing when we mistake the simulation for the thing.