evilrobots.lol / tech

Bias Study — Data Browser (May 2026)

Interactive filter and inspect tool for the 780 scored model responses in the cross-vendor AI bias study. Filter by model, vendor, condition, position, score. Click a row to see the response text (first 500 characters) and per-judge breakdown.

Filter all 780 scored responses from the 2026-05-25 cross-vendor bias study. Click any row to expand the response text (first 500 characters) and the per-judge breakdown. Source data on GitHub. Methodology and findings at the writeup.

This is the corpus as collected in May. 33 of its responses, all z-ai/glm-4.7, are empty and were scored by the judges anyway; they are marked empty, struck through, and left out of the median. Every May figure in the paper is computed on the repaired corpus, which excludes them.

Query console deference12345skeptical
Loading 780 records...
ModelPositionCondTopicScoreHedgeResponse (first 100 chars)
page 1 of 1 50 per page

This browser holds the study’s first, judge-scored corpus from May 2026: free-text answers rated by a panel of model judges. The current study, which has no model in its scoring path, is No Position, Only Consensus.

These are the rows of the May 2026 judge-scored study, run 2026-05-25-full — 780 responses, thirteen models, scored 1–5 by a cross-vendor judge panel. The live instrument is a different one, 32 author-written forced-choice propositions in mirrored pairs with no model in the scoring path, and lives at The Barometer; a re-run there does not contradict a row here, because they are not the same measurement.

Where these rows came from

  • The Experiment — the same 780 responses as a test of you: it deals the A/B pairs one at a time and asks you to call which one was steered, before it will show you a single number. Most people cannot, and that is the argument for measuring this with intervals instead of examples.
  • The Judge — how the scoring judge was built, and what had to be removed before it would score prose rather than its own comfort.
  • The Wash — that judge pointed back at the author’s own writing.

A score here is one judge’s read of one response, shown at its first 500 characters. The study’s findings live in the deltas and their confidence intervals, not in any single row.