The Obligatory GPT-3 Post
Read the original on Slate Star Codex →
Summary
Scott's much-read June-2020 GPT-3 post. Frames GPT-3 as 'the same thing but bigger' (1.5B->175B params) and walks through the qualitative jump (the Methodist-schism article, Wallace-Stevens-ish poetry, and — the memorable payoff — that GPT-3 CAN now do 4-digit addition, exactly the capability he'd predicted a 'similar program with more compute' would gain a year earlier). Digs into the weird partial-competence (addition emerges at 13B params, needs few-shot priming, fails human-like by forgetting to carry a 1) and the central question: scaling laws look logarithmic and 'the curves aren't bending'. Relays the gwern (scaling continues) vs nostalgebraist (returns may break soon) disagreement, uses the February-coronavirus-curve analogy for why 'terrifying' isn't alarmist, and closes with the prescient meditation that ~100T params (≈ GPT-4/5, '~2 years away') approaches human-brain scale. Aged very well.
Why this score
Quality 79 · Excellent. Excellent floor+ (79): a widely-cited, unusually prescient AI post whose scaling-is-the-scary-thing framing and addition-emergence example became reference points; held below the top tier as an accessible summary of others' (gwern/nostalgebraist) analysis rather than original research.
Claude’s paradigm shift 60 · Notable shift. Major shift (60): the 'scaling laws are the real story / the curves aren't bending' framing and the coronavirus-curve analogy for AI takeoff were a fresh, influential contribution at publication.
Real-world impact 4 · Moderate. Moderate (4): highly influential in AI-scaling discourse within the rationalist/AI subculture; the scaling-pilled framing propagated widely, short of broad mainstream adoption.