Scott Alexander, curated
← Back to curation

Testing AI's GeoGuessr Genius

Quality
71
Strong
Claude Shift
40
Moderate
RWI
2
of 10

Follow-up reading

Highlights From The Comments On AI Geoguessr — Highlights (companion) · May 2025

Summary

Scott stress-tests OpenAI's o3 at GeoGuessr with increasingly impossible personal photos (a featureless plain → Llano Estacado, ~110mi off; bare rocks atop Kala Pattar, Nepal → exactly right; a dorm room → right era/wrong place; zoomed grass → miss; a brown rectangle of the Mekong → 4th guess Mekong), with careful anti-cheating controls (metadata stripped, images flipped). The frame is the 'chimp and the helicopter' question about superintelligence — is there a bin of strategies as far beyond us as helicopters are beyond chimps, or just imaginable 'starships'? The GeoGuessr feats are his first visceral 'staring at the helicopter' moment, but the resolution is reassuring: o3 uses human-comprehensible cues (vegetation, sky, rock type), can't crack literally-impossible images, and sits at the top of the human range (Sam Patterson's head-to-head) — 'just very, very smart.' Ends on the frog-boiling worry that this calm might just be how things feel after they happen.

Why this score

Quality 71 · Strong. Strong band. An engaging, well-controlled experiment wired to a genuine conceptual hook (the chimp-helicopter superintelligence question) with a real, tentative update (AI capability here is human-comprehensible cue-reading, not magic). Strong rather than higher because the payload is one demonstration plus a borrowed analogy.

Claude’s paradigm shift 40 · Moderate. Moderate. The chimp-helicopter framing is a stock AI-risk analogy; the empirical GeoGuessr demonstration and the 'very smart not magic' read are the fresh part.

Real-world impact 2 · Minor. An engaging, well-controlled experiment (o3 at GeoGuessr with anti-cheating controls) wired to a genuine conceptual hook (the chimp-helicopter superintelligence question) with a tentative update — AI capability here is human-comprehensible cue-reading, not magic. Conceptual influence within AI discourse, one demonstration, no material change — low RWI.