METR
- Status
- ACTIVE — Model Evaluation & Threat Research; independent 501(c)(3) since December 2023 (formerly ARC Evals)
- Hazard — Reach
- 84
- RCH / FND / ENT
- 9 / 8 / 7
- Conduct
- CONFLICTED — GRADER OF ITS OWN BLUEPRINT
Institutional Archetype
THE GRADER — METR is the entity that runs the dangerous-capability and autonomy evaluations the labs cite in their system cards. It does not build models, write refusals, or set policy. It measures whether a model “can autonomously carry out substantial tasks” — software engineering, ML engineering, cybersecurity, research — and whether it shows the early shape of capabilities like self-exfiltration or making itself hard to shut down. The throughline is the instrument: take a frontier model, define the test on which its safety will be argued, and publish a result the labs and governments treat as the evidentiary baseline. The instrument does not change. Only the model under test does.
Mandate & Origin
METR — pronounced “meter,” for Model Evaluation & Threat Research — was incubated as the evaluations team inside Paul Christiano’s Alignment Research Center, where it was known as ARC Evals, and spun out into a standalone 501(c)(3) nonprofit, announced December 4, 2023.
- Stated mission, verbatim: “METR’s mission is to develop scientific methods to assess catastrophic risks stemming from AI systems’ autonomous capabilities and enable good decision-making about their development.”
- Self-description, verbatim: “METR (pronounced ‘meter’) is a research nonprofit that scientifically measures whether and when AI systems might threaten catastrophic harm to society.”
- Founded and led by Beth Barnes (Founder, CEO), who by her own team-page bio previously worked at OpenAI — evaluating scalable-oversight techniques and screening code models for misalignment before release — and collaborated with DeepMind’s chief scientist on scaling laws.
Funding & Backers
METR publishes its funder list and states plainly that it takes no money from the companies it grades.
- Independence statement, verbatim: “METR has not accepted funding from AI companies, though we make use of significant free compute credits.”
- Named backers on METR’s own about page include Schmidt Sciences, the Survival and Flourishing Fund (the philanthropic vehicle associated with Jaan Tallinn), the Audacious Project, the Sijbrandij Foundation, the Pew Charitable Trusts, Longview Philanthropy, the UK AI Security Institute, and the European AI Office.
The recurrence worth naming without asserting a hand: Schmidt Sciences funds METR, and the same Schmidt funding footprint reaches across the broader evaluation apparatus (see the file on Eric Schmidt). The grader is paid by a node that also funds much of the layer it operates in. That is the shape of the institution. The arithmetic is the finding.
Actions & Leadership Choices
The PR says “measure, don’t certify.” The deeds say something more entangled, and the entanglement is in the founding purpose, not in a scandal.
Actual founding purpose. METR was not built as a neutral metrology lab. It was built as the evaluation arm of Paul Christiano’s Alignment Research Center — ARC Evals — whose job from the start was to operationalize a specific thesis: that the catastrophic risk to watch is autonomous capability, and that the way to govern it is for labs to test for it before release. The purpose was to make that thesis testable and then make the test the gate. That purpose has been served.
- It wrote the self-governance template, then became the grader of it. In September 2023, while still ARC Evals, the team published the foundational “Responsible Scaling Policies” framing and disclosed that it had helped Anthropic write its RSP version 1.0. Within months OpenAI (“Preparedness Framework”) and Google DeepMind adopted broadly similar regimes. METR did not merely measure the frontier — it authored the document the labs now point to as their commitment, and it grades the labs against that very regime. The architect of the self-regulation standard is also its examiner.
- The one place the labs touch its resources is the one place to watch. METR refuses funding from the companies it grades and says so. But it makes use, by its own wording, of “significant free compute credits” — supplied by those companies. When the values were tested by the question does the grader take anything from the graded, METR drew the line at cash and crossed it at compute. The distinction is real and disclosed; it is also the single thread connecting the examiner’s resources to the examined.
- When it had a finding, it published the hedge, not the brake. METR’s verdicts (“we failed to find significant evidence for a dangerous level of autonomous capabilities”) have never, on the public record, stopped a release — and METR is explicit that “pre-deployment capability testing is not a sufficient risk management strategy by itself.” It measures; the lab decides; the deployment proceeds. The cost of the value (real independence) is borne in the hedge; the lab pays nothing for the grade.
Leadership choices. Founder and Co-CEO Beth Barnes is a former OpenAI alignment researcher who left in 2022; the policy seat — director of policy Chris Painter — reviewed an early draft of Anthropic’s revised RSP with Anthropic’s permission. The leadership is drawn from the lab it most closely grades and consults for the lab on the very standard it then evaluates. That is not a hidden hand; it is the published org chart. The grader, the standard’s author, and the standard’s first adopter share people and a draft history.
CONDUCT: CONFLICTED — GRADER OF ITS OWN BLUEPRINT. METR is genuinely independent of lab cash and genuinely careful in its measurements; it is also the body that wrote the self-governance regime, helped a frontier lab draft its first version, evaluates that lab against it, and runs on that lab’s compute. The earnestness is real and the conflict is structural — the examiner authored the exam.
Sources: METR announcement (formerly ARC Evals) — metr.org, Dec 4 2023; METR — About; Beth Barnes — METR team page; METR — Claude 3.7 Sonnet evaluation report; GPT-4 System Card — OpenAI; Frontier AI Taskforce first progress report — gov.uk; Responsible Scaling Policies — METR (formerly ARC Evals), Sep 2023; METR — Wikipedia.
Get updates on the Evil Robots series
Newsletter essays on AI escape, deception, and the humans who built them.