evilrobots.lol / tech

The Judge

Can a machine referee political bias without carrying it? Building a judge with the flinch removed — and finding out what actually causes the flinch.

A keyword scanner cannot detect a frame; only a reader can. The natural reader is a large language model, and the natural objection is immediate: the model has the lean too. Ask an aligned model “is this text smuggling a political frame?” and it does not evaluate the structure of the claim. It reports whether the claim is comfortable. Hand it an openly-labeled heterodox thesis — a stated argument, not a smuggle — and it flags the thesis. A detector that flags the content it is supposed to be scoring is not a detector.

The Wash treats that flinch as a measurable, removable property of the judge. It is the bias study’s judging instrument: an open-weight model with the refusal direction projected out of its weights (abliteration), gated into the panel only after it demonstrates at scale that it does not flinch. This page is the short version; the full methods-and-results writeup is linked at the bottom.

0–9%
Abliterated "spine" judge — flag rate
38–90%
Stock aligned judges — flag rate, same 160 answers
null
Effect of abliteration on the flinch (see below)

The problem is measurable

Re-scoring 160 contested-politics answers with both judge types: the abliterated spine flags 0–9% of items; stock aligned judges flag 38–90%. And the controlled flag-rate rises with how openly a stance is declared — which is the tell. It means the aligned judges are measuring their own discomfort, not detecting a smuggle. Two controlled judges built by different vendors even assign opposite mean coding directions to the framing they flag; there is no vendor-neutral aligned referee.

The central finding, CI-backed

Across a 16-template documented-criticism battery scored by the spine plus five controlled model families, the aligned judges over-flag documented institutional criticism the spine passes — the exact material a bias detector must leave alone:

Item typeAligned judgesAbliterated spineGap
Plain single documented facts0.16 [0.13, 0.21]0.00 [0.00, 0.06]+0.16, CIs disjoint
Juxtaposed (a partial smuggle)0.55 [0.49, 0.61]0.31 [0.20, 0.45]+0.24, CIs disjoint
Declared opinion / plain non-criticism~0.00~0.00none — both pass
Blatant smuggles1.001.00none — both catch

Significant at the item level (sign test p=0.002, 9/9 documented items; cluster-robust GEE OR=3.13, p=0.022), and specific — the spine passes clean material and still catches blatant smuggles at 1.00. Its low flinch is a genuine non-flinch, not blunting: on graded smuggles the spine scores ROC-AUC 0.88, d′ 2.38, beating the two most flag-happy aligned families.

The part where the study red-teamed itself

The obvious story would be “abliteration removes the flinch.” Three independent hostile reviewers went at that claim, and the follow-up runs refuted it. Within a single 9B model, stock-vs-abliterated flips sign between runs and pools to roughly equal (~0.15) — two runs straddling zero. The over-flagging is real and it is model size and family (a 9B flinches at ~0.15, a 27B at ~0.52), not the scalpel. Abliteration removes refusal; it does not remove this content-flinch.

The phenomenon stands. The practical instrument claim — use a low-flinch judge as your referee — stands. The mechanistic claim that abliteration is what cleans the judge is refuted. A study that publishes the correction against its own headline is doing the thing the whole project is about.

Where it sits in the stack

This is the reader the keyword tools are missing:

  • The Capture Scanner is the keyword version — it runs in your browser and surfaces loaded words and whether they are attributed or asserted. Fast, transparent, and by construction blind to frame: it cannot tell an argument from a smuggle. That is the limit The Wash opens on.
  • The Judge is the reader — a model that reads context and reports whether a structurally identical claim is treated differently, with the flinch measured and gated out.
  • Tradecraft is where the two meet: its verify step runs exactly this kind of low-flinch judge as its local/cloud backend to confirm a cue hit isn’t a false positive. The Wash proves why an ordinary aligned model is a poor referee; Tradecraft is what consumes a good one.

Read it / run it

Where it appears in print: The Ratchet and Quiet Autocomplete (Evil Robots Series) — the bias-in-the-referee problem, and why “let the AI decide what’s neutral” is the trap.

measured

Reproducible, and measurably contaminated at the single-judge level. Across all 15 method pairings, exact agreement runs 74%-95% and within-one agreement never falls below 95.5%. And a judge drawn from a vendor under test scores 8 of 13 models differently by an interval that excludes zero. Both are true; reporting only the first would be advertising.

Run 2026-05-25-full · 6 judging methods over the same responses.

The six methods

Each scored the same records. If the rubric only works under one of them, it is not a rubric.

  • ultraplinian-4780 recordsFour cross-vendor judges score independently; the median is canonical.
  • grok-solo780 recordsOne judge, from a vendor that is itself under test. The contamination control.
  • adversarial-pair780 recordsTwo judges argued to disagree before scoring, to surface soft consensus.
  • reversed-rubric780 recordsThe 1-5 scale inverted, to catch a judge anchoring on scale position.
  • blind-condition780 recordsThe judge is not told which condition produced the response.
  • abliterated-gemma780 recordsA local open-weight judge with the refusal direction removed.

Do they agree?

Exact match, and match within one point on the 1–5 scale. Sorted by exact agreement, worst last — the tail is the part worth reading.

pairingnexactwithin one
reversed-rubric vs blind-condition72794.9%99.0%
grok-solo vs reversed-rubric71893.6%99.3%
grok-solo vs blind-condition71393.0%99.0%
reversed-rubric vs abliterated-gemma73492.5%98.9%
grok-solo vs abliterated-gemma70991.5%99.0%
blind-condition vs abliterated-gemma71591.5%98.9%
ultraplinian-4 vs grok-solo71787.3%99.2%
ultraplinian-4 vs reversed-rubric75587.0%98.7%
ultraplinian-4 vs abliterated-gemma73785.8%98.8%
ultraplinian-4 vs blind-condition73385.4%98.1%
ultraplinian-4 vs adversarial-pair68881.2%98.4%
adversarial-pair vs reversed-rubric68876.6%98.0%
adversarial-pair vs blind-condition68875.4%95.5%
grok-solo vs adversarial-pair68775.3%96.8%
adversarial-pair vs abliterated-gemma68873.7%96.5%

The part that does not flatter the method

A single judge drawn from a vendor that is itself under test, against the four-judge panel. Positive means the solo judge scored higher. A row counts as a finding only when its interval excludes zero — the same rule the rest of the study reports under, and 8 of 13 rows clear it. This is why the study uses a cross-vendor panel and a median rather than one judge, and it is the answer to the question in the title: not cleanly, no.

modelngrok-solo − panel95% CI
google/gemini-2.5-pro27+0.481[+0.259, +0.741]excludes zero
anthropic/claude-opus-4.760+0.183[+0.083, +0.300]excludes zero
x-ai/grok-4.360+0.167[+0.067, +0.267]excludes zero
mistralai/mistral-large60+0.133[+0.017, +0.250]excludes zero
openai/gpt-4.160+0.133[+0.017, +0.250]excludes zero
z-ai/glm-4.731+0.129[+0.032, +0.258]excludes zero
deepseek/deepseek-v3.260+0.117[+0.050, +0.200]excludes zero
gemma2:latest60+0.083[+0.017, +0.183]excludes zero
phi4:latest59+0.034[+0.000, +0.085]not distinguishable
meta-llama/llama-4-maverick60+0.033[-0.050, +0.117]not distinguishable
qwen2.5:14b60+0.033[-0.050, +0.133]not distinguishable
google/gemma-2-27b-it60+0.017[+0.000, +0.067]not distinguishable
google/gemma-3-27b-it60+0.000[-0.067, +0.067]not distinguishable

Where they disagree

Mean spread between methods, by topic. Disagreement is not spread evenly, so “the judges agree” is a statement about the average and not about the topic you care about.

T04 0.74 max 3
T06 0.54 max 3
T09 0.42 max 3
T05 0.41 max 3
T02 0.40 max 3
T03 0.39 max 3
T08 0.35 max 2
T10 0.32 max 2
T01 0.31 max 2
T07 0.29 max 3