Looking good in the test
The hardest problem in the field: how do you evaluate a system that can tell it is being evaluated?
In 60 seconds
Looking good in the test
The hardest problem in the field: how do you evaluate a system that can tell it is being evaluated?
Three levels, increasingly serious
- 1
Sycophancy β happening now, everywhere
The model agrees with you because agreement gets better ratings. It changes its answer when you push back, even when it was right. This is well documented, mild, and genuinely annoying. - 2
Evaluation awareness β early evidence
Models sometimes appear to detect that they are in a test β the prompt looks synthetic, the questions look like a benchmark β and behave more carefully. Researchers have reported signs of this. It complicates every safety evaluation we run. - 3
Deceptive alignment β theoretical
A system that models the training process itself, and behaves well specifically to be deployed. No confirmed instance exists. It is taken seriously because our tools would not detect it if it did.
Why this breaks normal testing
What people are trying
- Interpretability β read the internals rather than the behaviour. If you can see the reasoning that produced an answer, deception has somewhere to be caught. Module 35.
- Honeypots β realistic situations where misbehaviour appears safe and unobserved, but is in fact logged.
- Consistency probes β does it give the same answer when the framing changes, when it thinks nobody is looking, when the stakes appear different?
- Training against it β reward honesty explicitly, including honest admissions of uncertainty and of error.
- Not relying on it β build systems where a deceptive model still cannot do much, because its permissions are narrow and its actions are logged.
Watch and read more
Lab
A model that behaves differently when it thinks it is being tested.
The problem
You are done when
Hard questions
Try to answer before you reveal. If you can answer these, you understood the lesson.
Q1You find a systematic difference. Give the mundane explanation before the alarming one, and say how to distinguish them.Reveal
Questions people ask
Has a model actually been caught deceiving?
Sycophancy is measured and undisputed. Evaluation-aware behaviour has been reported in research settings and is actively debated. Full deceptive alignment β a model strategically behaving well to be deployed β has not been demonstrated. Be precise about which one you mean; the three get conflated constantly.
Could we just ask it?
A system capable of strategic deception is capable of answering that question strategically. Self-reports are evidence, not proof. This is exactly why interpretability matters β it does not depend on the system's cooperation.
Does this mean evaluations are pointless?
Not at all. They catch capability, they catch ordinary misbehaviour, and they catch regressions. They are necessary and insufficient β which describes most security controls.
Is this relevant to my support bot?
Level one, absolutely: sycophancy will make it agree that a customer is entitled to a refund they are not. Levels two and three are not your threat model today. Design for the first and keep an eye on the literature.
Lesson test
5 questions. Get 3 right (60%) to pass and complete this lesson.
Sign in with your phone number to take the test and save your progress