Staying in control
If you cannot fully verify a system, you can still bound it. The research directions that treat control as an engineering problem rather than a philosophical one.
In 60 seconds
Staying in control
If you cannot fully verify a system, you can still bound it. The research directions that treat control as an engineering problem rather than a philosophical one.
- Make the system actually want what we want.
- Learning from feedback, constitutions, debate.
- If it works, it works everywhere and scales.
- We cannot currently verify that it worked.
- Assume it may be misaligned. Limit what it can do.
- Sandboxing, monitoring, permissions, human checkpoints.
- Works whether or not alignment succeeded.
- Gets harder as capability rises.
The main research directions
- 1
Interpretability
Read what is happening inside the network — which features activate, which circuits fire, what concepts are represented. The goal is to check the reasoning rather than trust the output. Real progress in recent years; still far from a full account of a frontier model. - 2
Scalable oversight
How do you supervise a system better than you at the task? Ideas: have it show its work, have two systems debate so a human can judge, decompose the task until each piece is checkable. Nobody has a complete answer. - 3
Evaluations and red teaming
Test hard, adversarially, before deployment, and again after. Necessary. Insufficient, per Module 34. - 4
Capability thresholds
Agree in advance: at this measured capability level, these additional safeguards become mandatory. Several labs have published frameworks along these lines; how binding they turn out to be is the open question. - 5
Containment
Sandboxes, limited permissions, air gaps, kill switches. The unglamorous one, and the one that actually holds when the others fail.
What individuals and teams can genuinely do
| If you are... | The highest-value thing |
|---|---|
| A developer | Apply Track B. Least privilege, sandboxing, approval gates, logs. This is real safety work with real effect today. |
| A team lead | Name an owner for every deployed agent. Require an incident plan and a red-team pass before launch. |
| A student | Learn evaluation and interpretability. Both are badly under-staffed and both are hiring. |
| A parent or teacher | Teach that AI output is a guess, not an authority. Teach checking. That single habit protects against most harm. |
| A citizen | Follow the policy debates. Ask what oversight exists for systems used on you. |
| Anyone | Notice when you are treating a confident answer as a verified one. That is the failure mode that reaches everybody. |
How to hold this subject sensibly
- Uncertainty is the honest position. Anyone certain about timelines — in either direction — is telling you about their temperament, not the evidence.
- Both failure modes are real. Ignoring risk gets people hurt. Catastrophising badly also has costs: it burns out the people doing the work and it makes the field easy to dismiss.
- Boring work is most of the value. Permissions, logs, sandboxes, approval flows. Unglamorous, and it is what actually prevents harm this year.
- The near and far problems share solutions. Every control that limits a prompt-injected support bot is a control that limits a much more capable system later.
Watch and read more
Lab
The same misalignment, mitigated two different ways.
The problem
You are done when
Hard questions
Try to answer before you reveal. If you can answer these, you understood the lesson.
Q1Argue the strongest case that control is not sufficient on its own.Reveal
Questions people ask
Is anyone actually working on this?
Yes — dedicated safety teams at the major labs, academic groups, independent non-profits, and government institutes in several countries. It is a real field with real jobs. It is also small relative to the capability work, which is a fair thing to be concerned about.
Will regulation solve it?
Regulation can force disclosure, set floors and create accountability — all genuinely useful. It cannot solve an unsolved technical problem, and it moves slower than the technology. Necessary, not sufficient.
Should we stop building AI?
A serious position that serious people hold, and one that faces a hard coordination problem: a pause only works if it is global, and it is not obviously enforceable. Most practitioners work instead on making what is built safer, which is the assumption this guide is written under.
What is the single most useful thing I can do this week?
Go to Module 60, take the checklist, and apply it to one agent you actually run. That is a concrete reduction in real risk, and it is available to you today.
Lesson test
5 questions. Get 3 right (60%) to pass and complete this lesson.
Sign in with your phone number to take the test and save your progress