Can This AI Save Teenage Spy Alex Rider From A Terrible Fate?
Read the original on Astral Codex Ten →
Summary
Vivid explainer of Redwood Research's robust-injury-classifier project (fine-tune GPT on 4,300 Alex Rider fanfics -> human-label violence -> train a classifier -> bounty-hunt adversarial examples -> retrain) as a concrete test of prosaic alignment. The adversarial failures (SEO-spam derails, poetic euphemism) motivate the payoff: prompting an agent for what it will do differs from putting it in the situation (the Generalissimo metaphor), so agentic alignment needs situation-prompts + interpretability + ELK, not just this classifier.
Why this score
Quality 75 · Excellent. Strong / low-Excellent edge — one of the better popular explanations of the adversarial-robustness problem; clear, memorable, does real conceptual work (the Generalissimo agentic-vs-prompted point, the three-things-needed synthesis) and aged well. 75.
Claude’s paradigm shift 52 · Moderate. Notable — a fresh, lucid framing/explainer building on Redwood's and ARC's work.
Real-world impact 3 · Moderate. Influential within AI-safety discourse. 3.