Pacific Design/ artificial intelligence

section 09 · status: live · 5 entries · updated 2026-08-29

AI Safety & Alignment

Models do what the training signal rewarded, which is not always what anyone wanted. The gap has a name, an engineering practice, an adversary, an evaluation discipline, and a genuinely open long game — taken in that order.

live · preference training reshaping an output distribution — the tail thins, never to zero

  1. What alignment meansthe gap between what you optimized and what you meant
  2. How models are alignedpreferences in, constitution over, refusals shaped
  3. Jailbreaks & prompt injectionthe model does what the text says — whoever wrote it
  4. Evaluating safetya passed eval is a lower bound, not a verdict
  5. The long gameoverseeing systems smarter than the overseers