After Orthogonality: Virtue-Ethical Agency and AI Alignment

The Gradient has published an essay titled "After Orthogonality: Virtue-Ethical Agency and AI Alignment" arguing that rational people do not have goals and that rational AIs should not have goals. The essay contends human actions are rational not because they are directed at final goals but because they align actions to practices, described as networks of actions, action-dispositions, action-evaluation criteria, and action-resources that structure, clarify, develop, and promote themselves.
The essay's central claim is that if AIs are to support, collaborate with, or comply with human agency, AI agents' deliberations must share a "type signature" with the practices-based logic humans use to reflect and act. It argues this matters not only for aligning AI to ethical ideals like human flourishing but also for aligning AI to core safety properties including transparency, helpfulness, harmlessness, and corrigibility. According to the essay, concepts like harmlessness or corrigibility are unnatural, brittle, unstable, and arbitrary for agents interpreting them as goals or rules, but natural for agents interpreting them as dynamics within networks of actions, action-dispositions, action-evaluation criteria, and action-resources.
The author introduces a formula, "promote x x-ingly," argued to capture something important about meaningful human life-activity and real human morality, giving examples such as art being the artistic promotion of art, romance being the romantic promotion of romance, caring about kindness being promoting kindness kindly, and caring about honesty being promoting honesty honestly.
The essay introduces the term "eudaimonic rationality" for a form of rational activity and valuing it argues is a useful or necessary framework for the agency and values of human-aligned AIs. It argues the concept of eudaimonia, meaning active rational human flourishing, does not simply point to a desired state or trajectory of the world to set as an AI's optimization target, but points to a structure of deliberation different from standard consequentialist rationality. It also argues eudaimonia suggests rational activity without a strict distinction between means and ends, or between instrumental and terminal values.
The essay states its arguments rest on dangers of a "type mismatch" between human flourishing as an optimization target and consequentialist optimization as a form, and on material advantages eudaimonic rationality plausibly possesses regarding stability and safety compared with deontological and consequentialist agency. It argues many classical AI safety considerations and paradoxes of AI alignment speak in favor of instilling AIs with eudaimonic rationality.
Based on reporting from the original publisher. Visit the source for full context and later updates.
Publisher excerpt
Preface This essay argues that rational people don’t have goals, and that rational AIs shouldn’t have goals. Human actions are rational not because we direct them at some final ‘goals,’ but because we align actions to practices [1] : networks of actions, action-dispositions, action-evaluation criteria,