Scott Alexander, curated
← Back to curation

Antagonizing Opioid Receptors for (Prevention of) Fun and Profit

Quality
73
Strong
Claude Shift
40
Moderate
RWI
2
of 10

Summary

Connects temporal-difference reinforcement learning to a real, underused addiction treatment. Dopamine encodes reward-prediction error and back-propagates to predictive stimuli (the monkey/bell/juice experiment); alcohol's opioid surge in the VTA makes the TD system link 'wanting' to drinking, which repeated enough becomes addiction. The elegant payoff is the Sinclair Method: take naltrexone (an opioid antagonist) and then drink — the reward never registers, the dopamine system fires prediction-error and downgrades the drinking-reward link, and the craving fades (claimed 25-78% success, apparently beating AA/willpower). Characteristic honest puzzlement: why no side effects, if you're knocking out the brain's learning system? — 'you'd think somebody would have noticed.' Notes the same pathway may treat smoking, self-harm, kleptomania, overeating.

Why this score

Quality 73 · Strong. Strong-ish explainer — clear, informative, and grounded in neuroscience, with a genuinely interesting and somewhat counterintuitive real-world payoff (extinguish a craving by indulging it under an opioid blocker). Held at low-Strong because TD learning and the Sinclair Method are both established; the value is the lucid synthesis connecting them.

Claude’s paradigm shift 40 · Moderate. Moderate — TD learning and the Sinclair Method pre-exist; the contribution is the clean mechanistic story linking the neuroscience to the treatment.

Real-world impact 2 · Minor. A clear neuroscience explainer connecting temporal-difference reinforcement learning to a real, underused addiction treatment (the Sinclair Method — extinguish a craving by drinking under an opioid blocker). Conceptual/explainer influence within rationalist/neuro discourse; the science is established, so a lucid synthesis with no material change — low RWI.