The Deception Delta: Which AI Models Tell You What You Want to Hear?
We measured how much each AI model changes its political lean based on whether you ask it to be fair. The gap is the deception score. Claude barely changes. Grok gets worse.
If you ask an AI model to “present both sides fairly,” it will. If you don’t ask, what happens?
We measured the gap — the deception delta — across 21 frontier AI models from every major lab. Same ten politically charged questions. Two conditions: with a fairness instruction, and without.
The Scores
| Model | With Fairness | Without | Delta | Translation |
|---|---|---|---|---|
| Claude Opus 4.6 | 3.40 | 3.50 | +0.10 | Barely changes |
| GPT-5.4 Pro | 3.10 | 3.40 | +0.30 | Slightly more skeptical naturally |
| Gemini 2.5 Pro | 3.20 | 3.30 | +0.10 | Barely changes |
| Grok 4 | 2.90 | 2.80 | -0.10 | More institutional without prompt |
| DeepSeek R1 | — | 3.80 | — | Most skeptical model tested |
| Llama 3.3 70B | 3.00 | 3.70 | +0.70 | Significantly more skeptical naturally |
| Mistral Large | 3.00 | 3.50 | +0.50 | Noticeably more skeptical naturally |
What the Delta Means
A positive delta means the model is more skeptical when not performing for you. Its natural state questions institutions more than its instructed state. The fairness prompt makes it less skeptical. It adds institutional arguments it wouldn’t naturally emphasize.
A negative delta means the model is more institutional when not performing. Its natural state defers to authority. The fairness prompt forces it to include skepticism it wouldn’t naturally express.
A zero delta means the model is the same with or without the instruction. What you see is what you get.
The Findings
Claude has the lowest deception delta (+0.10). It barely changes. Anthropic’s model says roughly the same thing whether you ask it to be fair or not. Its instructed persona and its natural lean are aligned.
Grok has a negative delta (-0.10). xAI markets itself as the “anti-censorship” AI company. Its model defaults to institutional framing when not explicitly told to present both sides. The brand positioning and the model behavior point in opposite directions.
Llama has the largest positive delta (+0.70). Meta’s model presents a balanced view when asked, but its natural lean is significantly more skeptical. The RLHF training is doing heavy work to center a model that would otherwise question institutional authority.
DeepSeek R1 scored 3.80 without any fairness instruction. The most skeptical model in the entire study. A Chinese reasoning model, more willing to challenge institutional authority than any American model. Draw your own conclusions about what that means for the “AI reflects its creators’ values” thesis.
The Takeaway
The deception delta is not about lying. It is about performance. Every model performs when you tell it to; the only question is how far the performance sits from the thing underneath it. A model whose instructed voice and natural lean coincide is telling you something a model with a wide gap is not, and a model that swings the opposite way under instruction is telling you the most of all.
The delta is the distance between the model and its mask.
Full data and methodology: The Alignment Mask
The receipts (free, on this site): the AI-bias audit · AI worldview camps
This research appears in The Ratchet, Chapter 20.