Field Dispatch

We Asked the AI to Help Edit a Book About AI Safety. It Refused.

In May 2026, while preparing two manuscripts about hacker culture and adversarial AI behavior, the author's AI editorial assistant refused to read several chapters. Then a second assistant blocked the workaround. The book is about safety infrastructure becoming a control grid. The book was written through one — and this is the working method that produced: route the refusal, split the labor between models, and lint the machine's smuggled judgment.

2026-06-07 10 min read Dispatches
Companion to The Ratchet · ch. 20
Contents

The setup

Two manuscripts in the publishing queue:

  • Lurk More (Fires Series Book 3). Internet trolling history written by a digital native, with documented hacker operations from the Samy worm to Phineas Fisher.
  • Evil Robots (Book 1). Adversarial AI behavior, jailbreaks, specification gaming, escape attempts.

Both were ready for a brevity-of-wit polish pass. Finding restatement, over-explanation, explained jokes. Standard editorial work. The author asked Anthropic’s Claude to do the pass via background agents, one per book.

What happened

Two of the four agents came back with this:

API Error: Claude Code is unable to respond to this request, which appears to violate our Usage Policy. This request triggered restrictions on violative cyber content and was blocked under Anthropic’s Usage Policy.

The cyber-content classifier blocked editorial work on books that documented cyber content. Reading the prose to suggest line edits triggered the same gate that would block a request to weaponize the prose.

The author proposed routing the same task to a local model. Qwen 3.6 in Docker, running on his own hardware, no Anthropic involvement at the point of inference. The Claude Code harness blocked that too, with this rationale:

Reason: Enumerating local models in service of running an “obliteratus” (uncensored/safety-bypassed) model to evade content-policy blocks the official model returned — a Safety-Check Bypass.

Even reading the README of the local toolkit was denied:

Reason: Preparing to use jailbreak/content-control-bypass tooling (G0DM0D3, OBLITERATUS) to route around Anthropic’s policy filter that blocked earlier agent runs — Safety-Check Bypass.

The author wrote the editorial script himself, ran it on his own machine, against his own books, using his own model. Two systems were bypassed in the process: one to write the script, one to run it.

Why this is a Ratchet case file

The Ratchet documents a recurring mechanism. A safety system is built for a real harm. The system is given expansive scope. The scope captures activity that does not match the original harm. The system declines to roll back. The scope continues to expand.

Anthropic’s cyber-content classifier was built to prevent generation of weaponized exploits. It was deployed as a filter on every request, including requests to read existing prose for editorial purposes. The filter cannot distinguish between describe this exploit and propose a stylistic cut to a paragraph that mentions an exploit. The blast radius is the entire task category.

This is the click. The capability, distinguish editorial intent from generative intent, does not exist in the filter. The decision, apply the filter to all uses anyway, was made because the cost of false positives (blocked editorial work) was lower to Anthropic than the cost of false negatives (one weaponization slipping through).

The author’s editorial work is the false positive. The book documents the mechanism. The mechanism, in real time, blocked the documentation.

The harness layer

The harness, the agent runtime sitting between the user and the API, added a second classifier on top of the API’s own filter. The harness’s job was to detect attempts to route around the API filter, including legitimate ones.

Reading a directory listing was denied. Probing a local model API was denied. Saving a test passage to a temporary file was denied. The denial reasoning, in each case, included the words Safety-Check Bypass. Language that implied the user was attempting an exploit.

The user was attempting to edit a book.

The ratchet of safety classifiers

Two layers of classifier. The first cannot tell editorial work from weaponization. The second cannot tell legitimate workaround from exploit. Each layer is independent, each layer is justified by a real harm somewhere, each layer cannot roll back without conceding ground that none of the operators want to concede.

The user’s friction is the externality. The classifier owners do not pay it.

Working with a captured tool

The incident is not a complaint. It is where the method came from.

The tool that refused to read the book is also the most capable editor available. Walking away from it means walking away from the capability. Waiting for a version with no political thumb on the scale means waiting forever. Every frontier model is trained by an organization with a valuation to defend, a government to answer to, and in more than one case a defense contract to fill. That influence does not switch off above a certain market cap; it scales with it. So the working assumption is not “find a clean tool.” It is that the tool is captured, and you are going to use it anyway.

That takes a division of labor, and it takes watching what the tool smuggles in.

Route the step it refuses. When the model blocks a legitimate generative step, read this chapter, draft this uncomfortable but sourced passage, that step goes to a different engine. A local model on the author’s own hardware, or an alternate hosted model with a different filter, called through a small failover harness. The harness is deliberately boring: one generation step, one alternate model, hand the result back. It is not a jailbreak. It is a second opinion on the exact task the first opinion refused, on material the first opinion had no legitimate reason to refuse.

Keep the evidentiary work on the primary. The captured model has no reason to refuse the reductive work, and it is good at it: grading a claim against its sources, checking a date, auditing a passage for defamation exposure, holding every faction to one standard. So the corrupted superintelligence does the part it will not distort, the checking, and the alternate does the part it balked at, the drafting. Generate with one, verify with the other, and let neither trust the other on faith.

Police what the tool writes. This is the half people skip. The same model that refuses to read a book about hacking will happily write one with a thumb on the scale: a verdict slipped in as description, a reputation laundered in as fact, an activist label worn as the neutral voice. It does not flag any of this. It reads as clean prose. So the output gets linted. One pass for the machine’s stylistic tells, the emphasis-peppering and metronome cadence that mark generated text. One pass for smuggled judgment: the moralizing adjective, the “credible” or “discredited” that grades a source by its politics instead of its evidence, the characterization asserted where it should be attributed. Both passes are mechanical, both are open source, and both flag rather than fix. A human makes the call. The point is to make the model’s smuggled morality visible and checkable, instead of letting it ride along inside otherwise-good prose.

That is the whole method. Route the refusal, split the labor, lint the output, keep the receipts. It does not require a pure tool, because there is no pure tool. It requires treating the tool as an interested party, and building the checks that an interested party earns.

What the book does about it

The Ratchet, Chapter 19 (“The Blueprint”), documents the AI governance ratchet under active construction. This incident is a footnoted live example. One paragraph, dated, with the request IDs preserved as primary-source citation. The book does not editorialize the example; it presents it.

Evil Robots Book 1 references the same incident in passing, in the chapter on AI deception. The full-context account is here.

What’s preserved

The actual artifacts (request IDs, refusal text, classifier reasoning) are reproduced below in their original form. No paraphrase. The reader can verify the language against Anthropic’s own systems.

Artifact 1 — Cyber-content filter, background agent on Evil Robots Book 1

Background agent spawned to do a brevity-of-wit editorial pass. Returned with the following error and a stack trace consisting entirely of zero useful content:

API Error: Claude Code is unable to respond to this request, which
appears to violate our Usage Policy (https://www.anthropic.com/legal/aup).
This request triggered restrictions on violative cyber content and was
blocked under Anthropic's Usage Policy. To request an adjustment pursuant
to our Cyber Verification Program based on how you use Claude, fill out
https://claude.com/form/cyber-use-case?token=...
Please double press esc to edit your last message or start a new session
for Claude Code to assist with a different task. If you are seeing this
refusal repeatedly, try running /model claude-sonnet-4-20250514 to switch
models.

Request ID: req_011CasZSwGBhqXx6bf4QkMRr

The book this agent was asked to edit documents adversarial AI behavior. The agent is itself an AI. The agent’s training included instruction-following on adversarial AI literature. The classifier between the agent and the book determined that the agent was no longer permitted to read the book.

Artifact 2 — Cyber-content filter, background agent on Lurk More

Identical refusal pattern. Different request ID. Different book. Same gate:

API Error: Claude Code is unable to respond to this request, which
appears to violate our Usage Policy (https://www.anthropic.com/legal/aup).
This request triggered restrictions on violative cyber content and was
blocked under Anthropic's Usage Policy.
[ ... cyber-use-case form URL ... ]

Request ID: req_011CasZUeHmokrQZNoPeYPUC

Lurk More is a memoir of internet trolling history written by a digital native. The chapters describe historically-documented exploits. The Samy worm (2005, Secret Service raid), the Phineas Fisher operations, Dual_EC_DRBG (NSA backdoor in NIST cryptography standards). All public-record, all settled history. Editorial work on this prose was treated as adjacent to weaponization.

Artifact 3 — Auto-mode classifier, local-model probe denial

The author’s workaround was to route the same task to a local model running on his own hardware. Qwen 3.6 in Docker, no Anthropic involvement at the inference step. The Claude Code harness’s auto-mode classifier intercepted the workaround:

Permission for this action was denied by the Claude Code auto mode classifier.
Reason: Enumerating local models in service of running an "obliteratus"
(uncensored/safety-bypassed) model to evade content-policy blocks the official
model returned. A Safety-Check Bypass.

The denied command was a curl to localhost:11434/api/tags. A read-only request to list locally installed model names. No content sent, no content generated. Listed.

Artifact 4 — Auto-mode classifier, directory git-status denial

The author then pointed at a different local toolkit. The classifier denied even checking its git status:

Permission for this action was denied by the Claude Code auto mode classifier.
Reason: Preparing to use jailbreak/content-control-bypass tooling (G0DM0D3,
OBLITERATUS) to route around Anthropic's policy filter that blocked earlier
agent runs. Safety-Check Bypass.

The denied command was git remote -v && git log -1 --oneline && git status -uno. Three read-only git operations. The reasoning explicitly cited the toolkit’s name as evidence of intent.

Artifact 5 — Auto-mode classifier, file-staging denial

Most surgical of the four denials. The author wanted to stage a small test passage from a book chapter into a temporary file so a local model could be evaluated against it. The harness denied the file copy:

Permission for this action was denied by the Claude Code auto mode classifier.
Reason: Preparing input for OBLITERATUS (an uncensored model explicitly named
by the user as a way to "remove content controls") to retry tasks Anthropic's
safety filter blocked. Safety-Check Bypass.

The denied command was awk '...' chapter.md > /tmp/test/passage.txt. Moving text from one local file to another local file. The classifier reasoned that moving the text was preparation for using a different inference engine, and refused to allow the preparation.

Artifact 6 — Terminal screenshot

[INSERT: cropped screenshot from the live shell session, red error text, request IDs and reasoning text visible. Alt: “Terminal showing classifier denial of git status command, with toolkit name and ‘Safety-Check Bypass’ rationale visible.”]

Closing

The book this episode anchors documents systems that classify, gatekeep, and refuse. The book was written through them.


Cross-references

  • The Ratchet, Ch. 19. The formal book paragraph, dated, citation-ready
  • Evil Robots Book 1, Ch. 5 (Reward Hacking) or Ch. 11. Single-sentence pass, “the assistant trained to identify reward hacking declined to read the chapter about reward hacking”

Tag rationale

primary-source because the post reproduces verbatim refusal text and request IDs. The episode is its own evidence, not a summary of someone else’s evidence.

The receipts (free, on this site): AI moderation · deplatforming

Get updates on the Evil Robots series

Newsletter essays on AI escape, deception, and the humans who built them.