Field Dispatch

The Deterrence Mechanism

Daniel Kokotajlo received the offer letter in May 2024. The letter was a separation agreement. It accompanied his resignation from OpenAI, where he had spent three years as a governance researcher working on the company's policy on the…

2026-08-02 18 min read Dispatches
Companion to Quiet Autocomplete

Daniel Kokotajlo received the offer letter in May 2024. The letter was a separation agreement. It accompanied his resignation from OpenAI, where he had spent three years as a governance researcher working on the company’s policy on the deployment of advanced AI capabilities. The separation agreement was, in its structure, a standard non-disparagement and non-disclosure agreement. In one substantive respect, it was not standard: the agreement provided that if Kokotajlo declined to sign it, he would forfeit his vested equity in OpenAI, approximately one point seven million dollars at the company’s then-current secondary-market valuation.

Kokotajlo declined to sign.

Within days, the structure of his refusal had been reported by the journalist Kelsey Piper, writing in Vox on May 17, 2024, in a piece that drew on internal OpenAI documents leaked by a source who has not, to this date, been publicly identified. The piece detailed not just Kokotajlo’s specific case but the broader pattern: OpenAI’s separation agreements, since approximately 2019, had included perpetual non-disparagement clauses, and a non-disclosure clause covering the existence of the agreement itself, backed by equity-forfeiture penalties, applied uniformly to departing employees regardless of seniority or termination reason. Piper’s follow-up reporting cut the harder way: the equity-forfeiture language had been in the off-boarding documents since 2019, and when departing employees had asked to sign a release stripped of the non-disparagement and secrecy terms, OpenAI’s lawyers had refused. The incorporation document the company cited as its authority to claw back vested equity bore three signatures. All of them Sam Altman’s.

OpenAI’s response, over roughly the ten days following the first story, was to disavow the practice. CEO Sam Altman posted on X that the company had never clawed back any vested equity and would not. He was, he wrote, “genuinely embarrassed” he had not known the provision existed. Asked by Vox whether its assurance that no equity had been or would be canceled marked a change in policy, the company said the statement “reflects reality.” OpenAI sent former employees an internal memo confirming it had “not canceled, and will not cancel, any Vested Units,” and released them from the non-disparagement terms. Multiple departing employees, including William Saunders and Carroll Wainwright, subsequently confirmed publicly that they had been bound by clauses they had now been released from.

Daniel Kokotajlo’s vested equity, the one point seven million dollars, was, by the company’s own subsequent statement, never withheld. OpenAI had, by its own admission, no documented case of an employee whose equity had been clawed back under the non-disparagement clauses.

The threat was the mechanism. The threat was never executed.

Vested equity you own and cannot sell, sealed behind glass with no door. The gag that never needs a hand over the mouth.

The natural framing — they took his money — is exactly backwards. The natural framing is wrong because they did not take his money. The natural framing is wrong because they did not need to take his money. The mechanism worked perfectly, on every employee who signed, without taking anyone’s money. The mechanism worked by the threat of the action it never had to perform.

The deterrence-enforcement distinction matters because it determines what reform looks like.

If the mechanism worked by enforcement, the reform is enforcement-side: make it harder to claw back equity, raise the burden of proof, give departing employees procedural rights to contest. These reforms are what the post-2024 reform cycle — Stephen Kohn’s July 2024 SEC complaint under Rule 21F-17, Senator Grassley’s August 2024 letter to OpenAI, Senator Warren and Representative Trahan’s August letter, Senate Judiciary’s September 17, 2024 hearing, Grassley’s S.1792 introduced May 15, 2025 — has substantially produced. Equity clawbacks for whistleblowing are now meaningfully harder than they were in 2023.

If the mechanism worked by deterrence, those reforms achieve approximately nothing.

The reform cycle assumed enforcement. The mechanism was deterrence. The reform cycle therefore did not, on the empirical evidence available as of May 2026, change the silencing rate of departing AI-lab employees. The rate of public disclosure by departing employees, measured across OpenAI / Anthropic / Google DeepMind / Meta AI / xAI from 2024 through early 2026, was approximately constant. The threats stopped being made; the silences continued.

The mechanism that does not need to be enforced is the most efficient possible mechanism. It is what students of capture-resistant institutional design call a credible threat without execution costs, a structure in which an entity can deter behavior without ever having to do the thing that deterring the behavior would require. Such structures are notoriously stable. They self-reinforce. They produce, over time, populations whose own internalized risk-assessment matches the deterrent intent so closely that the deterrent never needs to be invoked.

The non-disparagement architecture at OpenAI was such a structure. It was also, the chapter argues, the prototype.

The same architecture has been documented at every other major frontier AI lab. The forms vary. Anthropic’s “unless mutual” clause carve-out, applied selectively to the co-founders and to a small number of senior researchers, has been documented in reporting by The Information and through the public Anthropic founders’ compensation disclosures. Google DeepMind’s separation agreements, in their post-2024 form, contain comparable non-disparagement language; the language is enforced through the equity-vesting clawback provisions standard in Google’s overall compensation regime. Meta’s January 2025 layoffs were accompanied by separation agreements that included disclosure restrictions extending beyond standard trade-secret protections. The agreements’ specific terms became public when one terminated employee, in a since-deleted social-media post, quoted the binding clause.

The Right to Warn open letter was published on June 4, 2024, by current and former employees of OpenAI, Google DeepMind, and Anthropic — its named former-OpenAI signatories including Jacob Hilton, William Saunders, Carroll Wainwright, and Daniel Kokotajlo, with six further signatories anonymous — and endorsed by Yoshua Bengio, Geoffrey Hinton, and Stuart Russell. It demanded four structural reforms: an end to agreements barring risk-related criticism, a verifiable commitment to non-retaliation, an anonymous internal channel for raising concerns, and a right of employees to disclose risks publicly once internal channels failed.

As of May 2026, of the four demands, zero have been adopted by any major U.S. frontier AI lab in a form that civil-society observers describe as substantive. The labs have adopted versions of each demand in forms that comply with the letter while leaving the deterrent mechanism essentially intact. The reform was procedural. The architecture was structural. The two operate on different layers.

The silence the deterrence bought is the disclosure the books, including this one, have had to do without. Every chapter in this book, every claim in the book, every sourced receipt, has been assembled from primary documents that the institutions involved disclosed for reasons of their own — for compliance, for marketing, for litigation, for legally-mandated transparency. The book has not been able to draw on first-person accounts from the inside of the labs, except in cases where the speaker had already taken the personal cost of speaking: Kokotajlo, Saunders, Hilton, Adler, Aschenbrenner. That list of names is, structurally, the complete public record of insider AI-lab disclosure since 2020.

It is approximately ten people. The labs employ approximately ten thousand.


It is worth being precise about what the deterred employee stood to lose, because the precision is the whole of the threat, and the threat is the whole of the mechanism.

The equity at OpenAI did not take the form of ordinary shares. It took the form of units the company called Profit Participation Units, a vested interest in the company’s future profits rather than a stake in the company itself. The distinction sounds like an accountant’s footnote. It is in fact the load-bearing beam. A PPU could not simply be sold on a public market the way a share in a listed company can. Its value was realized through a company-controlled secondary market: periodic tender events that OpenAI organized, at which approved buyers, on terms the company set, were permitted to purchase units from employees who wished to sell. From those events the company could exclude a departing employee at its discretion. Those tender windows were the only path from a vested unit to cash. The time pressure on the exit agreement matched: departing staff reported roughly a week to decide whether to sign, inside a sixty-day clock. If they did not sign within sixty days, the units were gone.

The significance of company-mediated liquidity is that the threat did not need a clawback clause to bite. An employee did not need to be told that signed equity would be seized. The employee needed only to understand that the value of the units was realized at the company’s discretion, through a market the company controlled, at windows the company opened. The path from a vested unit to actual money ran, at every step, through an institution the employee was being asked to promise never to disparage. The non-disparagement clause and the controlled liquidity were not two separate instruments. One instrument, viewed from two angles. You could keep your units and your voice. The units would simply be worth whatever they were worth to a person who had annoyed the only entity that could turn them into cash.

They did not have to take anything. The architecture took nothing from anyone and silenced almost everyone. A clawback is an action; an action leaves a record, generates a grievance, invites a lawsuit, makes a martyr. Discretionary liquidity is not an action. It is a standing condition. It silences without an event, and an event is the thing whistleblower law knows how to attach to.


The reform cycle the chapter has already named understood the action and missed the condition.

Stephen Kohn’s July 2024 complaint to the Securities and Exchange Commission was filed under Rule 21F-17: the rule that prohibits an employer from taking any action to impede an individual from communicating directly with the Commission about a possible securities-law violation, including through confidentiality or separation agreements. It is a good rule, aimed at a real abuse, and it is aimed at the action: the clause, the impediment, the thing written into the agreement. Senator Grassley’s August 2024 letter, Senator Warren and Representative Trahan’s letter the same month, the Senate Judiciary hearing of September 17, 2024 at which William Saunders testified, and Grassley’s S.1792 the following May all share the same target. They make it harder to write the clause, harder to enforce the clause, harder to penalize the employee who breaches the clause.

None of them reaches the condition, because the condition is not written anywhere. There is no clause that says your illiquid equity will be worth less if you make us regret granting it. That sentence is not in any agreement. It does not have to be. It is the structure of the asset, and the structure of the asset is lawful, ordinary, and shared by most late-stage private companies in the technology sector. You cannot file a 21F-17 complaint against an instrument’s having been a PPU rather than a share. The reform cycle legislated against the visible blade and left the gravity in place.

The threats stopped being made. Altman disavowed the clauses; the general counsel committed in writing; named employees were released. And the rate at which departing employees spoke did not measurably move. The visible deterrent was withdrawn. The silence was undisturbed — the cleanest possible demonstration that the visible deterrent was never the load-bearing one.


The book has a way of knowing exactly how many people the mechanism failed to silence, because those people are its sources.

The complete public record of first-person insider disclosure from inside the American frontier labs, from roughly 2020 through early 2026, is a list short enough to print in a sentence: Daniel Kokotajlo, William Saunders, Jacob Hilton, Carroll Wainwright, the six anonymous signatories of the Right to Warn letter, plus Steven Adler and Leopold Aschenbrenner, who spoke outside that letter’s frame.

The people on the list are, almost without exception, people who had already moved past the point where the mechanism had purchase. Kokotajlo refused the agreement at the door and kept his voice by being willing to pay the full posted price, by his own account roughly one point seven to two million dollars, something near eighty-five percent of his family’s net worth. His full story belongs to later chapters. So does the parallel case of the absorbed witness. Here he is only the cleanest illustration of the rule: the variant where the antibody was intact and the apparatus simply outlasted the question of whether he was right. Saunders had left, gone to an evaluation outfit, signed the June 2024 letter, and testified under his own name to a Senate Judiciary subcommittee on AI oversight. Hilton, another former OpenAI researcher and signatory, described the equity-conditioned non-disparagement structure publicly and pressed the company to honor its reversal in writing. The Right to Warn signatories signed a public letter, which is the act of someone who has already decided the cost is bearable. These are not people the mechanism failed to deter. These are people who were no longer inside its field.

The people who are not on the list are the evidence. The gap between those numbers is not a population with nothing to say. It is a population for whom the structure of their compensation, and the structure of their future employability in a small industry of a handful of firms that all know each other, makes saying it a decision with a price that does not need to be quoted to be understood.


There is one figure on the list who broke the mechanism by a different route, and his case is instructive precisely because it is the exception that maps the rule.

Leopold Aschenbrenner did not refuse an agreement. He was dismissed. He left OpenAI’s Superalignment team in April 2024 — fired, by the company’s account, for what it characterized as a leak of an internal document to outside researchers, which it said was unrelated to a security memo. By his own account, he had written that memo to the board, arguing OpenAI’s security was egregiously insufficient against theft of model weights and algorithmic secrets by foreign actors, and was told the memo was “a major reason” for the firing. The document he was said to have leaked, he characterizes as a benign brainstorming note shared with three external researchers. Two months later he published Situational Awareness, a long-form essay that made him, briefly, one of the most-discussed voices on the trajectory of the technology. He went on to found an investment fund built on the thesis the essay argued.

The structural point is not the dispute. The structural point is that a person who is dismissed is, by the time the question of silence arises, already outside the geometry that produces silence in the people who stay. A retention-and-liquidity mechanism works on people it is retaining. It has no grip on someone it has already let go. Less grip still on someone let go on terms that gave him a grievance and took away the upside he would otherwise have been protecting. The mechanism deters by giving people something to lose and making the loss contingent on their silence. Fire someone before the something has fully vested, and you have, at the cost of one employee, manufactured the one kind of witness the mechanism cannot touch: a person with the knowledge, the standing, and nothing left in the company’s gift to forfeit.

That is why the variant that broke is the one where the person had less to lose, not more. The mechanism is a function of stake. Remove the stake — by paying it out as Kokotajlo was willing to do, or by destroying it through a dismissal as in this case — and the voice returns. Leave the stake in place and illiquid and discretionary, and the voice does not need to be taken. It stays home on its own.


The deterrence reading has a comparison class. It shows the mechanism is neither new nor an accident. Two established regimes bracket it. The older one is the form that admits its own purpose. The newer one is the form a regulator decided to outlaw.

The older one is the intelligence community’s pre-publication review. A person granted access to classified information signs, as a condition of that access, one or both of two standard instruments: the SF-312, the Classified Information Nondisclosure Agreement, and, for sensitive compartmented information, the Form 4414. The 4414 commits the signer to submit for review any writing that contains or purports to contain such information or a description of the activities that produce it. The obligation is lifelong, content-based, and a prior restraint: it sits in front of the writing, not behind it. It does not claw back and it does not punish after the fact, in the ordinary case; it conditions the disclosure on passing through a government office first. Its teeth were set by the Supreme Court in Snepp v. United States in 1980. A former CIA officer who published a book about the agency’s activities in South Vietnam without submitting it for review was held to have breached a fiduciary obligation, and was made to surrender all his profits from the book to the government, in a constructive trust, even though the book was found to contain no classified information. The remedy is disgorgement of earnings, structurally the same lever as “you keep your voice and lose what you earned.” The regime has since survived a direct constitutional challenge: the Knight First Amendment Institute and the ACLU filed Edgar v. Ratcliffe in 2019 on behalf of five former intelligence and military employees, arguing the system was an unconstitutional prior restraint; the district court dismissed it in reliance on Snepp, the Fourth Circuit affirmed for the government in June 2021, and the Supreme Court — the case by then restyled Edgar v. Haines as the named defendants changed — declined to take it up. The regime stands. It binds, for life, an estimated several million current and former clearance holders. It is broader and more durable than anything OpenAI ever wrote down.

The newer regime runs the other way, and it is the one the AI-lab reform cycle was actually imitating. SEC Rule 21F-17, adopted under Dodd-Frank’s whistleblower provisions and effective in 2011, prohibits any person from taking any action to impede an individual from communicating directly with the Commission about a possible securities-law violation, including by enforcing or threatening to enforce a confidentiality agreement. It is the rule that treats the written impediment itself as the offense, and the Commission has enforced it with real money. It opened in 2015 by fining KBR a hundred and thirty thousand dollars over confidentiality language that barred employees from discussing investigation facts without the legal department’s consent, the first action of its kind. It has since reached the large numbers: thirty-five million from Activision Blizzard in February 2023, over separation agreements that required employees to notify the company of any regulatory request for information; ten million from D. E. Shaw in September 2023, over agreements that barred disclosure of confidential information to anyone outside the firm and required departing staff to affirm they had filed no prior complaints; eighteen million from J.P. Morgan Securities in January 2024, over release agreements that let clients respond to SEC inquiries but barred them from reaching out voluntarily.

Set the two regimes beside each other and the AI labs occupy a third position. They had the financial sector’s written clauses, the confidentiality and non-disparagement language a 21F-17-style rule treats as the violation, and the reform cycle, importing exactly that logic, stripped them. But there is a discipline to be kept here, and it is the whole of the comparison’s honesty: OpenAI did not violate Rule 21F-17, and nothing in this chapter claims it did. The rule governs SEC-registered and securities contexts; OpenAI is a private company, and there is no equivalent statutory floor for frontier-AI-safety disclosure at all. That absence is the point. The 21F-17 clause was reachable. It is written down, it is the kind of thing a regulator can void and fine. The condition was not, because there is no clause that says your illiquid equity will be worth less if you make us regret granting it, and no AI-safety regulator with a 21F-17 in hand to void it if there were. What the labs kept, when the clauses were stripped, is the intelligence community’s half of the bargain: a standing condition that sits in front of disclosure, lifelong, content-shaping. Except that in the labs’ case the standing condition is not a clearance office but the structure of an asset. The reform cycle legislated the part that could be regulated. It left the part that could not. The labs ended up holding the more durable of the two regimes, having had the weaker one taken away.


It is the reason the mechanism is worth this much attention even though it has, on the documentary record, never once been executed.

A mechanism that works by deterrence and never by enforcement, that silences without an event, that leaves no record because its operation is the absence of an action, such a mechanism is, for as long as it goes unexamined, deniable. Everyone involved can say, truthfully, that no one was ever silenced, because no one was ever made to be silent. The clauses were disavowed. The equity was never withheld. There is no clawed-back employee, no seized account, no martyr. The silence is real and the cause of it cannot be pointed at, which is the most comfortable possible position for an institution to be in.

That comfort ends the moment the mechanism is described. Once you can name the structure — vested but illiquid equity, liquidity at the company’s discretion, a small industry where every employer knows the others, and a population that does its own risk arithmetic so reliably that the threat never has to be made — you have converted an emergent property into a design. The thing that was an accident of how technology compensation happened to be structured becomes a thing that can be chosen on purpose. The next lab does not have to stumble into the deterrence mechanism. It can build it, knowing what it is, knowing that the reformers will legislate against the clause and leave the condition, knowing that ten out of ten thousand is not a failure rate but a specification.

That is the finding. It is why the chapter has refused the easier story the whole way through. The easier story is that an institution did something wrong, was caught, and was reformed. The true story is that the institution did nothing, that doing nothing was the instrument, that the reform addressed the something the institution had stopped doing, and that the nothing continues, lawful and silent and now, having been described, available to be done deliberately. There were about ten people willing to pay for the right to speak. The mechanism’s achievement is everyone else.

Get updates on the Evil Robots series

Newsletter essays on AI escape, deception, and the humans who built them.