GPT-2 As Step Toward General Intelligence
Read the original on Slate Star Codex →
Follows up on
↳ Do Neural Nets Dream Of Electric Hobbits? — Essay · Feb 2019
Summary
Feb-2019 argument that GPT-2 — and humans — are both 'brute-force statistical pattern-matchers' that distill experience into a 'slurry' and rebuild from it; the only difference is how finely you blend. Catalogues GPT-2's untaught emergent faculties (counting to five, acronymizing AAPSLB, English<->French translation with no French in the corpus, TL;DR summarization, Pope-pastiche poetry, the curated GPT-2 'Moloch') and argues prediction is 'the golden key' that incidentally forces a model to learn arithmetic, history, biology and law. A wake-up call against 'AGI is impossible / unrelated to current work.'
Why this score
Quality 81 · Excellent. Excellent floor (81). A provocative, generative reframing — 'scaled next-token prediction -> general capability' — that aged remarkably well: emergence, unsupervised task-acquisition and the scaling intuition are exactly what GPT-3+ confirmed. Sits just above the Right-Is-The-New-Left 80 / Marijuana-MMTYWTK 78 Excellent floor and below the dedicated AI standouts; loose/polemical patches ('your mom...', the hand-wavy P!=NP bit) keep it out of the mid-80s, but the core insight is correct and was non-obvious in 2019.
Claude’s paradigm shift 62 · Notable shift. 62 — Major-shift territory. In Feb 2019 (pre-GPT-3) the claim that a next-token predictor was already doing the same thing brains do, and would scale to general intelligence, was a genuinely contrarian, ahead-of-consensus frame with only partial precedent, and it durably installed the 'prediction is general learning' view in rationalist AI discourse.
Real-world impact 3 · Moderate. 3 — influential within the rationalist/AI-discourse sphere; helped popularize the scaling/emergence intuition among an educated subculture, but no direct material or policy effect (the band-3 'niche professional sphere' tier).