section 09 · status: live · 5 entries · updated 2026-08-29
AI Safety & Alignment
Models do what the training signal rewarded, which is not always what anyone wanted. The gap has a name, an engineering practice, an adversary, an evaluation discipline, and a genuinely open long game — taken in that order.
live · preference training reshaping an output distribution — the tail thins, never to zero
- What alignment meansthe gap between what you optimized and what you meant
- How models are alignedpreferences in, constitution over, refusals shaped
- Jailbreaks & prompt injectionthe model does what the text says — whoever wrote it
- Evaluating safetya passed eval is a lower bound, not a verdict
- The long gameoverseeing systems smarter than the overseers