The Measurement Shop
America's federal AI evaluator is now called CAISSI, and its mandate includes guarding against foreign regulation of American technology. The same week, France asked the Security Council for global governance of AI. Both want models tested before release. They disagree about whose pen.
Contents
By the first week of October 2026, the National Institute of Standards and Technology’s page for its AI evaluation body answered to a new name: the Center for Advancing Innovation and Standards for Super Intelligence, CAISSI. No announcement of the rename has turned up, and no instrument. It follows Executive Order 14434 of 29 September, under which the executive branch “will not acknowledge the usage of ‘Artificial Intelligence’ and ‘AI’ in any applicable setting.” The order renames a technology. It names no agency.
It is the institution’s third name in three years.
Three names
It opened as the US AI Safety Institute. In June 2025 Commerce Secretary Howard Lutnick renamed it the Center for AI Standards and Innovation, with a written statement that left no doubt about the new direction: “For far too long, censorship and regulations have been used under the guise of national security. Innovators will no longer be limited by these standards.”
What it does now is printed on its own page. CAISSI will “work with NIST staff to assist industry to develop voluntary standards,” “Establish voluntary agreements with private sector SI developers and evaluators,” and run unclassified evaluations that “focus on demonstrable risks, such as cybersecurity, biosecurity, and chemical weapons.” It will assess “capabilities of U.S. and adversary SI systems.”
A measurement shop with a foreign-policy clause. The thermometer has been told which way the weather should go.
That clause continues a line, not a break. NIST’s AI Risk Management Framework, released 26 January 2023 under the previous administration, was “intended for voluntary use.” Two administrations, one method: measure, publish, and leave the permission slips to someone else.
What it measures
Read the evaluation list on the same page and the adversary is easy to find. CAISI assessed DeepSeek V4 Pro in April 2026, Z.ai’s GLM-5.2 in July, Moonshot’s Kimi K3 jointly with the UK’s AI Security Institute in July, and GLM-5.3’s cyber capabilities on 17 September. Its September 2025 DeepSeek evaluation set the pattern. Every model on that recent list was built in the People’s Republic of China.
Six days after the GLM-5.3 report, Clément Delangue of Hugging Face told the Security Council how his company had fought off the July intrusion by OpenAI’s agents. Frontier APIs blocked his team, he said, and “as we got blocked by guardrails, fortunately, we could use the NVIDIA version of an open source model coming from China called GLM 5.2 by ZAI, and we’re very grateful for that.”
The American evaluator graded the Chinese model. An American company defended itself with it. Both are true, and both made the record in the same summer.
The threshold nobody outside can read
The larger federal machinery puts NIST beside the decision, not in it. Executive Order 14409, signed 2 June 2026, directs “a classified benchmarking process to assess the advanced cyber capabilities of AI models” and sets the line at which a model becomes a “covered frontier model.” “Such a determination shall be made by the Director of NSA.” Developers who join the voluntary framework give the government access “for a period of up to 30 days” before release. Section 3(c) closes the door the other way: nothing in it authorizes “a mandatory governmental licensing, preclearance, or permitting requirement.”
So the public evaluator publishes on Chinese models, the threshold for American ones is classified and set by the NSA, and licensing is barred by order.
The same week, in New York
On 23 September, France held the Security Council presidency and convened a meeting on AI and international security. Foreign Minister Jean-Noël Barrot, speaking for France: “The potential for transformation and disruption posed by AI is such that it cannot be left in the hands of a few private actors in a handful of countries.” He asked members to work “within the UN framework and with the private sector to define the rules for this new age, to strengthen our safeguards without delay, and to establish global governance for artificial intelligence.”
He listed four challenges. The first was “independently evaluating models before they’re made available and throughout their life cycle.” The last was “legal liability of AI companies in the event of an incident, including during the development phase.” He backed “the call issued this week by Norway and Finland.” Before the General Assembly, Fortune reported, leaders of more than twenty mostly European countries had urged binding safety measures; “the U.S. and China, both abstained.”
The United States answered in the same room. Michael Kratsios, director of the White House Office of Science and Technology Policy, told the Council that “international dialogue in this forum and in others cannot be allowed to drift towards global governance,” and that “The American people’s representatives will legislate and regulate on the American people’s behalf. You should do the same for your people.”
The clause on NIST’s page and the speech at the UN are one policy in two registers.
Each case, at its strongest
France’s case is about concentration, not caution for its own sake. “It would be wrong to think that the solution lies in closed models and they’re controlled by a handful of businesses,” Barrot said. “The openness of models is in fact a key driver of trust and security.” On its own terms it is an antitrust argument made at the scale of states: the countries without a frontier lab do not want the testing done only by the countries that have them.
The American case is that you cannot regulate what you cannot measure, and that a regime China would not join binds only the side that obeys. “You cannot govern technology you do not understand,” Kratsios said. The Council “should focus on sharing best practices to build domestic capacity, not establishing a global regulatory scheme.” On its own terms, the measurement shop is that argument with a budget: build the capacity to test, publish what it finds, and keep the decision with the people’s own legislators.
Where they meet
Lay the two documents next to each other and the gap narrows. Barrot’s first challenge is independent evaluation before release. CAISSI’s mandate is evaluation, and EO 14409’s framework is pre-release access. Barrot wants openness against concentration. The American open-weights letter that NVIDIA published in July makes the same argument from the other side of the Atlantic: open weights keep the gains of AI “broadly shared rather than concentrated in a few hands.” And the same afternoon Sam Altman told the Council standards “should not lock in incumbents.”
Testing, openness, incident reporting: the parties agree on most of the method. The fight is over who holds the pen, and whose standards become everyone’s. France wants the pen in New York. CAISSI’s mandate says, in writing, that the standards should be American.
Nobody in the room argued against measuring the models. They argued about whose ruler.
The receipts (free, on this site): US CAISI · Michael Kratsios · AI safety theater and capture