AivexaNewsSearch
AI news for builders and product teamsChecked every hour

The Artificiality of Alignment

Collected Oct 1, 2026

An essay published by The Gradient argues that the current trajectory of alignment research is "misaligned" with the possibility that AI causes widespread, concrete and acute suffering, and is instead solving the problem of building a product people will pay for.

The essay, which first appeared in Reboot, says public discourse about AI risk conflates speculative future danger with present-day harms and confuses large models with algorithmic and statistical decision-making systems. It focuses on OpenAI and Anthropic, which it says own the most powerful models and take alignment and superintelligence most seriously in public communications.

Citing a New York Times interview, the essay says Nick Bostrom, author of Superintelligence, defines alignment as ensuring increasingly capable AI systems are aligned with what the people building them seek to achieve. It notes OpenAI names building superintelligence as a primary goal, and quotes the company's stated reasons: belief it will lead to a much better world, and that stopping it would be unintuitively risky and difficult.

On techniques, the essay describes reinforcement learning with human feedback (RLHF) and reinforcement learning with AI feedback (RLAIF, also known as Constitutional AI), used for ChatGPT and Claude respectively. It says both train a preference model aligned to "helpfulness, harmlessness, and honesty" (HHH), built through iterative pairwise comparisons by a human or AI.

The essay identifies two problems: determining which values and whose, and that feedback-based techniques address a chatbot surface layer rather than fundamental model capabilities. It says RLHF/RLAIF is tailored to building better products, and cites OpenAI lobbying for reduced regulation while publicly advocating government involvement, plus Anthropic raising hundreds of millions at a time. The author states they are not arguing alignment research is not worth pursuing.

Read at The Gradient

Based on reporting from the original publisher. Visit the source for full context and later updates.

Publisher excerpt

This essay first appeared in Reboot . Credulous, breathless coverage of “AI existential risk” (abbreviated “x-risk”) has reached the mainstream. Who could have foreseen that the smallcaps onomatopoeia “ꜰᴏᴏᴍ” — both evocative of and directly derived from children’s cartoons —