Scott Alexander, curated
← Back to curation

Looking for information on scoring calibration

Quality
54
Solid
Claude Shift
25
Slight
RWI
2
of 10

Summary

A short technical question with one sharp observation: standard scoring rules (log score, squared error) conflate calibration with domain knowledge — someone who always says 70% and is right 70% of the time scores better than someone who always says 60% and is right 60%, though both are equally well-calibrated. Scott notes the only pure-calibration measures he's seen are calibration curves (sets of ordered pairs or single points), and asks whether there's an accepted way to reduce a calibration curve to a single number representing calibration alone, separate from domain ability.

Why this score

Quality 54 · Solid. Minor/Solid boundary. Largely a request for information, but it carries a genuinely sharp methodological observation (scoring rules entangle calibration and domain knowledge), so it scores just above the bare-admin floor per the prompt/question convention.

Claude’s paradigm shift 25 · Slight. Slight — a well-posed question with a useful observation, but no developed answer of its own.

Real-world impact 2 · Minor. Largely an information request, but carrying a sharp methodological observation (scoring rules entangle calibration with domain knowledge); a discourse-internal point with no material-world reach → RWI 2.