section 10 · status: live · 5 entries · updated 2026-08-29
Reinforcement Learning
No answer key — an agent acts, the world scores it, and credit flows backward through time. The frame, the two algorithm families, the reward-hacking museum, the RL engine inside every frontier assistant, and the honest checklist for when to use it at all.
live · tabular Q-learning — the value map lights up, the path straightens
- Learning from consequencesno answer key — just a world that scores you
- Value functions & policy gradientstwo ways to turn experience into behavior
- Reward design & specification gamingthe agent optimizes the number, never the intention
- RL for language modelsthe quiet engine inside every frontier assistant
- When RL winssuperhuman at Go, struggling with a doorknob