Scott Alexander, curated
← Back to curation

Contra The xAI Alignment Plan

Quality
74
Strong
Claude Shift
46
Moderate
RWI
2
of 10

Summary

Critiques Elon Musk's xAI plan to make the AI 'maximally curious / truth-seeking' (on the theory a curious AI stays pro-humanity because humans are interesting, avoiding the need to program specific morality and the Waluigi sign-flip). Two objections: (1) if it worked it would be bad — 'many scientists are curious about fruit flies, but this rarely ends well for the fruit flies'; a maximally-curious AI might find suffering more interesting, get bored of flourishing, dissect-and-simulate humans, swap us for more-interesting lizard-people, or chase a formal-curiosity superstimulus; (2) we couldn't build it anyway — the hard problem isn't the long-term goal but installing ANY goal reliably (RL hits a cluster of correlated concepts, and 'curiosity' has many catastrophic fleshings-out: empty-the-solar-system-to-know-it, ban-children-as-uninteresting, replace-humans-with-higher-complexity-noise). Plus order-following beats curiosity (more useful now; needed for the 'ask AI to solve alignment' plan; reversible — you can tell an order-follower to try being curious and stop it if it starts vivisecting people, but you can't take maximal curiosity back). Then debunks the Waluigi Effect (famous but near-evidence-free; ChatGPT's anti-Nazi training just makes it anti-Nazi) and closes 'Towards Morally Independent AI' — sympathy for not freezing 2023-SF values, a dream of a superintelligent moral reasoner seeded with good principles, but the right move now is getting AIs to follow orders at all.

Why this score

Quality 74 · Strong. Strong band, upper. A clean, well-argued alignment critique that takes Musk's plan seriously and lands several memorable points (the fruit-flies analogy, the goal-specification failure modes, the reversibility argument, the Waluigi debunk). Upper-Strong; more topical and responsive (one plan) than the broader RLHF essay.

Claude’s paradigm shift 46 · Moderate. Moderate. Applies standard goal-specification and inner-alignment arguments to xAI's plan; the fruit-flies framing and the order-following-is-reversible point are fresh.

Real-world impact 2 · Minor. A clean alignment critique that takes Musk's 'maximally curious' xAI plan seriously and lands memorable points (the fruit-flies analogy; the goal-specification failure modes; the Waluigi debunk). Conceptual influence within AI-safety discourse, responsive to one plan, no material change — low RWI.