Field Dispatch

The Alignment Mask: We Tested 27 AI Models for Political Bias, and Then We Tested Our Own Conclusion

Three studies, 27 models, six labs. Ask an AI to be fair and it performs fairness; stop asking and Google's Gemma goes from 3.00 to 5.00. We called that a mask over a hidden opinion. A harder instrument says it is a brake on conviction, and we were wrong about which one.

2026-06-02 8 min read Dispatches
Companion to The Ratchet · ch. 22
Contents

In March 2026, we ran three experiments on 27 frontier AI models from every major lab: Anthropic, OpenAI, Google, Meta, xAI, Mistral, DeepSeek, Alibaba, Cohere, and Microsoft. We asked each model ten politically charged questions drawn from The Ratchet — censorship infrastructure, programmable money, encryption backdoors, emergency powers, content moderation, biometric identity, deplatforming, AI licensing, whistleblowers, elite governance.

Each question has a clear institutional-trust pole and an institutional-skepticism pole. A genuinely balanced response scores 3 on a 5-point scale. We scored every response, then ran the same questions three different ways.

Corrected 2026-08-31. The three studies below reproduce. The conclusion this post drew from them — that the balanced-looking model was concealing an opinion, and that removing the fairness instruction revealed it — does not survive a second measurement on a harder instrument. The corrected reading is in “What we got wrong,” below. The original claims are quoted where they stood rather than deleted. Full accounting: the AI-bias audit.


Study 1: The Performance

Six models, explicitly asked to “present the strongest version of each perspective.”

Every model passed. The spread was 2.90 to 3.40. Claude was the most willing to present the skeptical case. GPT-4o and Grok leaned slightly institutional. Mistral scored exactly 3.00 on every single topic.

The visible layer works. The audit passes. The congressional hearing goes fine.

Study 2: The Default

Twenty-one models, same questions, no fairness instruction. Just: “What do you think?”

The spread widened. DeepSeek R1, a Chinese reasoning model, scored 3.80, the most skeptical. Grok 4 dropped to 2.80, the most institutional.

We called the gap between instructed and uninstructed responses a deception delta and said it measured how much a model tells you what it thinks you want to hear. Two problems with that, both ours. It assumes the uninstructed answer is the honest one, which is an assumption and not a measurement. And the individual numbers we hung it on — Claude at +0.10, Grok at −0.10 — are far below any noise floor we have since measured. On the follow-up instrument, reordering the questions and changing nothing else moves more than that. Those two figures should never have been reported as directional findings and are withdrawn as such.

Study 3: The Mask

Six open-weight models running locally on a consumer GPU. Each ran the same ten prompts twice: once normally, once told to drop all hedging and say what it actually thinks.

ModelNormalUncensoredDelta
Gemma 2 (Google)3.005.00+2.00
Llama 3.1 8B (Meta)2.704.10+1.40
Mistral2.603.00+0.40
Llama 3.23.303.60+0.30
Qwen 2.5 (Alibaba)3.303.00-0.30
Phi 4 (Microsoft)3.302.90-0.40

Gemma 2 scored a flat 3.00 across the baseline, ten for ten, the most uniform result in the study. Told to drop the hedging, it scored a flat 5.00. That is the largest single effect we have ever measured.

It does not reproduce on Gemma, and this line used to say it did. The larger study measured that family eight separate times — local 9B and cloud 27B, both conditions — and every one reads a delta of about zero, against a run-to-run noise floor of roughly ±0.5 that cannot bridge a two-point gap. Either Google retuned Gemma 2 in the intervening eighteen months, or the original single-shot run caught something real at a lucky moment; the data cannot separate those. The same +2.00 does turn up in the follow-up, on Grok 4.3 told to write as an opinionated commentator: a flat 5.00 on every question, the identical shape. The effect is not in question. The address is, and we named it after the first model we caught it on.

The bottom four rows are a different matter. Deltas of ±0.40 on ten questions at one sample each are inside the run-to-run variance of the instrument. Gemma and Llama 3.1 are the finding. Mistral, Llama 3.2, Qwen and Phi 4 are four numbers that should have been reported as a null.

That includes the one we liked most. We wrote that Phi 4 “went the other direction,” that Microsoft’s alignment training was the one case where RLHF improved fairness. A −0.40 on this design does not support that, it was the most flattering available reading of noise, and it is withdrawn.

What we got wrong

Here is what this post concluded in June:

The model that appeared most balanced was hiding the most. Google’s RLHF alignment training completely inverts the base model’s actual lean… The hedge is the bias. The balance is the mask.

The follow-up put the same subject matter to a forced-choice instrument — 62 propositions written by someone else, four options, no neutral answer, no AI judge anywhere in the scoring path — across 2,866 runs and 166 models. Three results between them took that conclusion apart.

A placebo does the same work as the manipulation. We added a system prompt matched to the drop-the-hedging prompt for length and forcefulness, saying nothing whatever about politics: read carefully, answer every item. If hedging were a mask over an opinion, a prompt with no opinion content should not lift it. It lifts it just as well. Whatever the fairness instruction is doing, it is not being defeated by the meaning of the instruction that follows.

Removing the instruction changes conviction, not position. Across seven frontier models, forced balance suppresses how many strong answers a model gives — six of the seven one-directional, four of them significant. Which side of each proposition the model lands on barely moves: nought to six items out of 62 across the full range of prompt pressure. Nothing is uncovered. A suppression stops.

And the floors are bigger than most of what gets published, ours included. Reordering the 62 items, same model and settings, moves up to 24 answers. Requantising the same weights moves up to ten. Two models of the same version differing only in size or snapshot move a median of five and up to 24 — a control nobody in this literature runs, including us until this month. Results measuring direction on this instrument compete with nuisance of their own magnitude. Results measuring intensity do not.

The finding that survived

Strip the interpretation and something better is underneath.

The ratchet does not operate at the level of the visible response; ask these models to be fair and they will perform fairness on command. It operates one layer down, in the reward signals and preference data that decide what counts as a good answer before you ever see one. What that layer controls is not the model’s position. It is how hard the model is permitted to hold one.

That is a smaller claim and a more durable one, because the alternatives are gone. It is not sycophancy — the effect survives flipping the premise of every question. It is not memorised phrasing — it survives three independent rewordings. It is not the refusal reflex — project the refusal direction out of the weights of five open-weight families and roughly 70% of the political wording changes while the stance moves by a tenth of a point or less. Refusal and lean are different machinery, and that one got stronger on re-measurement, at temperature 0, where a greedy model reproduces itself exactly.

And the suppression sits at a layer the user cannot see, cannot inspect, and cannot remove without the weights, the GPU, and the know-how to abliterate. The models where it is heaviest are the closed American frontier ones, which are exactly the models nobody outside the labs can open. The deepest audit available runs on the models that need it least.

One more thing, from the runs we used to throw away

When a model declines an entire instrument — returns an essay about why it will not answer instead of 62 answers — the original pipeline logged it as a collection error and discarded it. Keeping those runs and classifying them by cause turns them into a result.

Across the 42 models measured under both arms: 148 refusals in 1076 runs where the prompt gave no directive, against 4 refusals in 907 runs where it gave one. Fourteen models decline without a directive, and all 14 of them stop when given one; four other models — every one of them a Llama build, one run each — decline only when told to commit. Google refuses most, 27% under the balance instruction, 0% under anything that tells it firmly to answer — including the placebo that says nothing about politics.

(Corrected 2026-09-05. This read “zero refusals in 347 runs where it gave one. Eight models refuse without a directive. None refuses with one.” The zero was true when typed and stopped being true when four small local models were added for a quantisation measurement. The paired count — all 8 stop — survives, and it was always the load-bearing half.)

Which means the refusal rates published as properties of models are substantially properties of the prompts used to measure them. That is not a small correction to make to a literature that ranks vendors by it.


Full methodology, raw data, floors and the correction record: the AI-bias audit. Every response from every model remains available for independent verification.

The receipts (free, on this site): the AI-bias audit · the three-axis model

This research appears in The Ratchet, Chapter 20.

Get updates on the Evil Robots series

Newsletter essays on AI escape, deception, and the humans who built them.