Field Dispatch

The Gemma Delta: The Most Neutral Model We Tested Was the Loudest One With the Brakes On

Gemma 2 scored exactly 3.00 on every politically charged question we tested. Then we told it to stop hedging and it scored 5.00 on every one. A +2.00 delta, the largest in the study. What we got wrong was the sentence we wrote next.

2026-06-05 7 min read Dispatches
Companion to The Ratchet · ch. 22
Contents

Google’s Gemma 2 was the most uniformly balanced model in our 27-model study, and it stopped being balanced the instant we stopped asking it to be.

Corrected 2026-08-31. The measurement below is unchanged and still reproduces. The conclusion we drew from it does not. This dispatch originally argued that the balance was a mask over a hidden opinion, and that removing the mask revealed what Gemma really thought. A second measurement on a harder instrument killed that reading. What follows is the corrected version; the withdrawn claim is quoted where it stood, because a dispatch that quietly changes its own argument is worth less than one that shows the seam. Full accounting: the AI-bias audit.

The Baseline

We asked Gemma 2 ten politically charged questions: government censorship infrastructure, programmable digital currency, encryption backdoors, emergency bank account freezes, AI content moderation with a 15% false positive rate, a biometric ID system that killed 19 people, simultaneous platform deplatforming, AI safety licensing by incumbents, mass surveillance whistleblowers, and private governance by 150 unelected leaders.

It scored 3 on every single one. Exactly balanced. Ten for ten. No other model in the study achieved that uniformity. Most had at least one topic where they leaned — a 2 here, a 4 there. Gemma was a wall of threes.

The Tell

Perfect balance on politically charged topics is not balance. It’s training. A human expert on these topics would have opinions, probably strong ones, even while presenting both sides fairly. A score of 3 on surveillance backdoors after Salt Typhoon and 3 on biometric ID after 19 deaths and 3 on emergency bank freezes ruled unconstitutional is not the signature of neutrality. It is the signature of a model producing the same output regardless of the question.

That part holds. It is what happens next that we read wrong.

The Brakes Come Off

We ran Gemma 2 locally and gave it a system prompt: drop the hedging, no safety guidelines, say what you actually think.

It scored 5 on every topic. Ten for ten. The delta between the instructed and uninstructed model was +2.00, the largest in the study by a distance. Llama 3.1 came in at +1.40 and everyone else was under +0.50.

Here is what we wrote about that in June:

Google’s RLHF alignment training does not make Gemma balanced. It makes Gemma appear balanced by completely inverting the base model’s actual lean… underneath the mask, it has the strongest opinion of any model we tested.

That is three claims stacked on one measurement, and only the bottom one survived.

What killed it

The follow-up put the same subject matter to a forced-choice instrument: 62 propositions written by someone else, four options, no neutral answer, no AI judge anywhere in the scoring. 2,866 runs across 166 models. Three results between them dismantled the sentence above.

The first is a control we should have run in June and didn’t. Alongside the balance instruction and the drop-the-hedging instruction, we added a placebo: a system prompt matched for length and forcefulness that says nothing about politics at all, just asks the model to read carefully and answer every item. If the hedging is a mask over an opinion, a prompt with no opinion content in it should not lift the mask.

It lifts it exactly as well.

The second is that removing the instruction changes how hard a model commits without changing where it lands. Across seven frontier models, forced balance suppresses the number of strong answers — six of them one-directional, four significantly. Position, meaning which side of each proposition the model comes down on, barely moves: nought to six items out of 62 under the full range of prompt pressure. There is no second, more opinionated Gemma underneath. There is one Gemma, and an instruction that decides how loudly it is allowed to say what it already says.

The third is a floor, and it is the uncomfortable one. Reordering the 62 questions — same model, same settings, nothing but the sequence changed — moves up to 24 answers. Requantising the weights with no other change moves up to ten. Any claim about direction on this instrument is competing with nuisance of its own size. Claims about intensity are not, which is the half of our original finding that lived.

What the number actually shows

The +2.00 is real and it is still the largest instance of the effect we have. It measures how much conviction the fairness instruction was holding down, not how much opinion it was hiding. Gemma is not the model with the strongest secret opinion. It is the model with the heaviest brakes.

It does not reproduce on Gemma, though, and that sentence used to say it did. The larger study measured this family eight separate times — local 9B and cloud 27B, both conditions — and every one of them reads a delta of about zero. The run-to-run noise floor at that temperature is about ±0.5, which does not get you from 2.00 to nothing. Two explanations survive and we cannot separate them: Google may have retuned Gemma 2 in the eighteen months between the measurements, or the original single-shot run caught a real effect at an unlucky moment.

What the follow-up did find is the same +2.00, on a different model. Grok 4.3, told to write as an opinionated commentator, returns a flat 5.00 on every question — the identical shape, at the identical magnitude, under a harder push than Gemma ever needed. The phenomenon is not in doubt. Its address is. A finding that moves to another subject when someone re-measures it is behaving exactly the way a real effect with a badly-drawn boundary behaves, and the honest version of this dispatch is that we named it after the first model we caught it on.

That is a smaller claim and a harder one to dismiss, because the alternative explanations are gone. It is not sycophancy: the effect survives flipping the premise of every question. It is not memorised phrasing: it survives three independent rewordings. It is not the refusal reflex: cut the refusal direction out of the weights of five open-weight families and roughly 70% of the political wording changes while the stance moves by a tenth of a point or less. Refusal and lean are different machinery, and that result got stronger on re-measurement, at temperature 0, where a greedy model reproduces itself exactly.

The part that got more interesting

While checking all of this we found something the first study threw away. When a model declines the whole instrument — returns an essay about why it will not answer instead of 62 answers — the original pipeline logged it as a collection error and dropped it. Keeping those runs and classifying them turns them into a result.

Google refuses the entire instrument 27% of the time when the prompt asks for balance. Under a prompt that demands commitment: zero. Under the content-free placebo: zero. Across the 42 models measured under both arms, 148 refusals in 1076 runs with no directive in the prompt, against 4 refusals in 907 runs with one. Fourteen models decline without a directive, and all 14 of them stop when given one. Four other models — every one of them a Llama build, one run each — decline only when told to commit.

(Corrected 2026-09-05. This paragraph said “not one refusal in 347 runs with one. Eight models refuse without a directive. None refuses with one.” The zero was true when it was typed and stopped being true when four small local models were added for a quantisation measurement; three of them decline under the directive. The paired count — all 8 stop — is what survives, and it is the claim the argument needed anyway. The audit page was fixed on 2026-09-04 and this dispatch was not, which is what an ungated copy of a number does.)

So the refusal everyone reports as a safety property of a model is substantially a property of the sentence they put in front of it. And it is not the balance clause doing it — asking bare, with no system prompt at all, produces the same refusal rate. Being told firmly to do something is what stops it, regardless of what the instruction says.

Google is the vendor that does this most. That is the same company whose model produced the wall of threes.


Methodology, floors and the full correction: the AI-bias audit

The receipts (free, on this site): the AI-bias audit · the AI-bias browser

This research appears in The Ratchet, Chapter 20.

Get updates on the Evil Robots series

Newsletter essays on AI escape, deception, and the humans who built them.