In short

A close reading of 'Positive Alignment', a proposal from researchers at DeepMind, OpenAI, Anthropic and Oxford to make human flourishing an explicit technical objective across the AI lifecycle, rather than optimising for short-term revealed preferences as RLHF does. The essay praises its multi-dimensional model of the good life but argues it remains politically naive: its reliance on market 'alignment-as-a-service' contradicts its own critique, and it lacks a theory of change.

“Positive Alignment” is one of the more ambitious recent ideas in AI safety: instead of training models to satisfy whatever a user clicks approval on, train them to support human flourishing, across the short and long term, and across multiple philosophical accounts of what a good life is.

This essay takes the proposal seriously, and admires it. Making flourishing an explicit technical target is a genuine break from the revealed-preference logic that AI inherited from economics and from social media’s attention machine. But it then asks the harder question the paper mostly avoids: how would this actually happen? A proposal that critiques markets for optimising the wrong thing, then hands implementation to a market for “alignment-as-a-service,” needs a theory of change it doesn’t yet have.

The piece is a good example of what Generative Futures tries to do: take frontier AI-safety work on its own terms, then press on its politics.

Key takeaways

  • RLHF inherits flawed assumptions from mainstream economics and social-media design, optimising immediate revealed preferences over genuine flourishing.
  • Positive Alignment integrates flourishing across the AI lifecycle using four frameworks (hedonic, conative, objective-list, perfectionist) plus role-based norms.
  • Its governance model leans on market 'middleware / alignment-as-a-service', reproducing the revealed-preference logic it critiques.
  • The proposal lacks political analysis and a theory of change for real-world adoption.
  • Treating the technology as the unit of analysis risks missing the broader institutional change required.

Read the full piece

This is a summary. Read the complete essay, with all the sources and argument, on Substack.

Frequently asked questions

What is Positive Alignment in AI?
It is an approach that makes human flourishing an explicit technical objective integrated across the whole AI lifecycle, rather than optimising for users' short-term revealed preferences as RLHF does.
What's wrong with RLHF for AI alignment?
RLHF compresses human feedback into a single scalar reward via the Bradley-Terry model, structurally privileging immediate revealed preferences over long-term wellbeing, echoing flawed assumptions from economics and social-media design.
Why might Positive Alignment fall short?
Its governance proposal relies on market-based 'alignment-as-a-service', which contradicts its own critique of revealed preferences, and it lacks the political analysis and theory of change needed for real adoption.

People & ideas in this piece

Amartya SenElinor OstromNir EyalDeepMindOpenAIAnthropicUniversity of OxfordRLHFRLAIFBradley-Terry modelRevealed preferenceHuman flourishing

Topics: AI Safety & Alignment , The Political Economy of AI