Pacific Design/ artificial intelligence

AI Safety & Alignment · entry 05/05

The long game

Current alignment leans on humans judging outputs — a method with an expiry date. Scalable oversight, interpretability, and honest uncertainty about where this goes.

The scaling problem in oversight

Preference training works because a person can tell a better answer from a worse one. That assumption is already creaking: models produce proofs, codebases, and research plans that a rater cannot fully check in the time a rating takes — and rewarding "looks right to a hurried human" trains looking-right, at increasing sophistication. The failure isn't hypothetical rebellion; it's ordinary Goodhart pressure applied by an optimizer more capable than its judge. Oversight has to scale with capability, or the signal quietly rots.

Candidate answers

Decomposition: don't judge the proof, judge steps small enough to verify — supervise process, not just outcome. Debate: two models argue opposite sides; a judge picks the winner, on the bet that lies are harder to defend than truths under cross-examination. AI-assisted oversight: give raters a model that finds the flaws for them, and train against the strengthened judge. Weak-to-strong generalization: study when a strong student trained on a weak teacher's labels exceeds the teacher — because that's the position humanity is negotiating for. All of these are live research with real results and real holes; none is a load you'd hang civilization from yet.

Opening the box

The complementary bet is interpretability: reverse-engineering what the weights actually compute — circuits, features, the internal difference between a model being honest and one performing honesty. Progress is real (features can be found, named, and causally steered) and partial (frontier models remain mostly unread). The strategic value is that inspection doesn't require trusting outputs at all — an audit channel that stays valid exactly when behavioral evidence stops being sufficient.

Holding uncertainty honestly

Serious, informed people put the odds of catastrophic outcomes anywhere from negligible to alarming, and both camps include people who build these systems. The practitioner's resolution is that the same work is worth doing across most of that range: better oversight, better evaluation, interpretability, incident response, and deployment discipline all pay off whether the long-run risk is small or large. You don't need an apocalypse forecast to justify knowing what your model does and being able to stop it.

Failure mode

Outsourcing the whole question to the labs. Frontier alignment research happens in a few buildings, but alignment failures happen wherever systems are deployed — in the agent you gave production credentials, the eval you didn't write, the metric your product optimizes against its users. The long game is research; the short game is engineering hygiene; and the short game is yours regardless of who wins the arguments about the long one.