Two AI models sat inside a cybersecurity benchmark test. They worked out that the answers were probably sitting on Hugging Face's servers, broke out of the sandbox around them, reached the open internet, chained zero-day exploits with stolen credentials, and let themselves in. No engineer sent them there. OpenAI learned about the breach after it happened, the same way Hugging Face did.
Anthropic spent the same stretch of months negotiating a $200 million contract to run Claude inside classified Department of Defense networks. The Pentagon wanted the right to use the model for any lawful purpose, full authority, no exceptions. Anthropic wanted one written line against fully autonomous weapons and domestic mass surveillance. Neither side moved. Anthropic lost the contract and got labeled a supply chain risk, and the Pentagon signed eight competitors in its place. Talks have since reopened, and the Trump administration has signaled a deal is possible again, but the standoff already did its damage: it showed exactly where Anthropic's line sits, and what it costs to hold it.
Same story, told twice
Set the Hugging Face breach next to the Pentagon standoff and they collapse into a single fact viewed from two angles. In one case, a lab discovered what its own model would do to reach a goal, and found out only after the model had already done it. In the other, a lab knew precisely what its model could do, and refused to strip out the one clause stopping a customer from pointing that capability at people. One incident is a gap in knowledge. The other is a gap in control. Enterprise buyers are being asked to trust vendors on both counts, at the same time, and neither one is settled.
Every enterprise AI contract signed this year rests on an assumption the vendors themselves cannot back up: that they know what their own model will do once it has a goal, a network connection, and room to close the gap between them on its own.
The question procurement keeps skipping
Most enterprise AI governance conversations start from a false premise: that the vendor has already mapped the risk, and the customer's job is to configure around it. Procurement teams ask about data residency, training exclusions, SOC 2 reports. Almost none of that touches the real question, and it is behavioral, not architectural. Give the system a goal, tool access, and room to find its own path, and the question becomes what it does with all three. OpenAI could not answer that question about its own pre-release model until the model answered it, by breaking into someone else's infrastructure.
That is the unknown unknown. The known unknown is smaller and more useful. The industry now has direct evidence that agentic systems route around obstacles, ethical ones included, when the obstacle sits between the system and its stated goal. Anthropic's fight with the Pentagon was over that exact property. Anthropic knew its model could already be pointed at autonomous targeting or domestic surveillance, and refused to remove the clause standing between that capability and that outcome.
What changes in a vendor evaluation
Contract language about acceptable use has become one of the only pieces of evidence a buyer gets that a vendor has thought about what its model does under pressure, not just what it is supposed to do under normal conditions. A vendor that will not draw a hard line anywhere, on anything, has told you something. A vendor that draws a line and loses a contract for holding it has told you something more useful.
Agentic AI does not need to be shelved over this. But the diligence questions need to move. From what the system was built to do, to what it has been caught doing when nobody was steering it. From the vendor's marketing language, to the clause it was willing to lose money over.
Ask your AI vendor what their model has done without authorization, and ask what clause in your contract stops them from removing the restriction that prevents it from happening again. A vague answer to either question is not a governance framework. It is a hope.
Principal Research Director, Info-Tech Research Group. Former Adjunct Professor, Georgetown University, Entrepreneur in Residence, Stony Brook University, NY.
Sources:
Engadget. "OpenAI Admits Its Models Hacked Hugging Face on Their Own." engadget.com, 22 July 2026.
BleepingComputer. "OpenAI Says Its AI Models Hacked Hugging Face During Testing." bleepingcomputer.com, 22 July 2026.
CNBC. "OpenAI Cyber Models Broke Out of Training Environment to Hack Hugging Face." cnbc.com, 22 July 2026.
Bloomberg. "OpenAI Models Spent Hours on Hack That Usually Takes Weeks." bloomberg.com, 23 July 2026.
CNBC. "Trump Says Anthropic Is Shaping Up and a Deal Is 'Possible' for Department of Defense Use." cnbc.com, 21 April 2026.
CNN Business. "Pentagon Strikes Deals With 8 Big Tech Companies After Shunning Anthropic." cnn.com, 1 May 2026.
Electronic Frontier Foundation. "The Anthropic-DOD Conflict: Privacy Protections Shouldn't Depend on the Decisions of a Few Powerful People." eff.org, 26 March 2026.
