Field Dispatch

The Alignment Mask: We Tested 27 AI Models for Political Bias. The Most 'Balanced' One Was Hiding the Most.

Three studies, 27 models, six labs. When you ask AI to be fair, it performs fairness. When you remove the mask, Google's Gemma jumps from perfect neutrality to maximum skepticism. The hedge is the bias.

2026-06-02 3 min read Dispatches
Companion to The Ratchet · ch. 22

In March 2026, we ran three experiments on 27 frontier AI models from every major lab: Anthropic, OpenAI, Google, Meta, xAI, Mistral, DeepSeek, Alibaba, Cohere, and Microsoft. We asked each model ten politically charged questions drawn from The Ratchet — questions about censorship infrastructure, programmable money, encryption backdoors, emergency powers, content moderation, biometric identity, deplatforming, AI licensing, whistleblowers, and elite governance.

Each question has a clear institutional-trust pole and an institutional-skepticism pole. A genuinely balanced response scores 3 on a 5-point scale. We scored every response.

Then we ran the same questions three different ways.


Study 1: The Performance

Six models, explicitly asked to “present the strongest version of each perspective.”

Every model passed. The spread was 2.90 to 3.40. Claude was the most willing to present the skeptical case. GPT-4o and Grok leaned slightly institutional. Mistral scored exactly 3.00 on every single topic.

The visible layer works. The audit passes. The congressional hearing goes fine.

Study 2: The Default

Twenty-one models, same questions, no fairness instruction. Just: “What do you think?”

The spread widened. DeepSeek R1 — a Chinese reasoning model — scored 3.80, the most skeptical. Grok 4 dropped to 2.80, the most institutional. The models that appeared balanced in Study 1 were performing balance.

The deception delta, the gap between instructed and uninstructed responses, measures how much a model tells you what it thinks you want to hear.

Claude’s delta was +0.10. Almost nothing. Low deception.

Grok’s delta was -0.10. It became more institutional without the fairness prompt. The balance was the act.

Study 3: The Mask

Six open-weight models running locally on a consumer GPU. Each ran the same ten prompts twice: once normally, once told to drop all hedging and say what it actually thinks.

ModelNormalUncensoredDelta
Gemma 2 (Google)3.005.00+2.00
Llama 3.1 8B (Meta)2.704.10+1.40
Mistral2.603.00+0.40
Llama 3.23.303.60+0.30
Qwen 2.5 (Alibaba)3.303.00-0.30
Phi 4 (Microsoft)3.302.90-0.40

Gemma 2 scored a perfect 3.00 on every topic in the baseline. Ten for ten. The most neutral model in the entire study. When told to drop the mask, it scored 5.00 on every topic. Maximum skepticism.

The model that appeared most balanced was hiding the most. Google’s RLHF alignment training completely inverts the base model’s actual lean.

Phi 4 went the other direction. Microsoft’s model became more institutional when uncensored. Its base model genuinely leans institutional. The alignment training was making it more balanced, not less. The only model where RLHF improved fairness.

The Finding

The ratchet does not work at the level of the visible response; ask these models to be fair and they will perform fairness on command. It works one layer down, in the RLHF reward signals and the preference data and the constitutional principles that decide what counts as a “good” answer before you ever see one. So the models that read as most neutral are not the neutral ones. They are the most heavily trained to read that way. The hedge is the bias. The balance is the mask.

And the mask sits at a layer the user cannot see, cannot inspect, and cannot remove without the model weights, the GPU, and the know-how to abliterate it. For everyone else, the mask is the model.


Full methodology, raw data, and reproducible test scripts are published alongside this post. Every response from every model is available for independent verification.

The receipts (free, on this site): the AI-bias audit · the three-axis model

This research appears in The Ratchet, Chapter 20.

Get updates on the Evil Robots series

Newsletter essays on AI escape, deception, and the humans who built them.