Fifteen thousand calls to paid cloud services in under an hour ran up about $50,000 in charges, and no attacker was involved.
An accounting agent running on artificial intelligence entered a runaway execution loop, made more than 15,000 high-cost application programming interface calls, and disrupted live business transactions. Mandiant published the case in its AI Risk and Resilience report on September 16, the second day of the Cyber Defense Summit at the Ronald Reagan Building in Washington, D.C. Mandiant calls it a consequence of weak controls. My read is simpler. A spend ceiling would have capped the bill. An owner with a kill switch would have ended it in minutes.
An accounting agent made 15,000 high-cost calls in under an hour
Every security vendor wants you to picture the autonomous attacker. The largest dollar figure in this story came from your own software, running under weak controls.
A spend ceiling would have capped the bill. An owner with a kill switch would have ended it in minutes.
Google's threat tracker clocks a credential-theft campaign at six hours
Google Threat Intelligence Group published its third-quarter 2026 AI Threat Tracker, which Help Net Security covered on September 8. It describes a Mandiant investigation from the second quarter. A financially motivated actor with access to an organization's cloud infrastructure used an AI coding chatbot, a prompt, and a set of agent instructions to plan and run a mass credential-harvesting campaign in under six hours. Thousands of third-party credentials were taken. The agent handled vulnerability scanning, troubleshooting, and Internet Protocol address rotation without manual intervention.
The same report says Google has not yet observed a fully autonomous pipeline run against targets in the wild. Defend against what exists today: stolen credentials moving faster than your response process.
Mandiant's report adds a second case. An attacker hijacked an active AI coding-assistant session at an unnamed software-as-a-service provider and used it to install an infostealer through a poisoned PyPI package. The attacker took GitHub OAuth tokens, then spread the Shai-Hulud worm across about 100 internal repositories. Mandiant does not say how the session was taken over.
Google answers a sub-7% true-positive rate with Mantis and Fairwind
Google puts the true-positive rate of naive AI code scanning below 7%, according to IT Brief's coverage of the announcement. Its answer is Mantis, a set of agent skills that finds a suspected flaw, reproduces it in a sandbox, writes a patch, and runs regression tests against it. Google open-sourced the core skills in late June under an Apache 2.0 license, and the repository states it is not an officially supported product. Its README warns that the suite generates and runs code and belongs only in isolated environments.
You do not need the repository to take the idea. Send developers a reproduced bug, and stop sending them a model's confidence score.
Google's cyber-tuned model follows the same logic. Gemini 3.8 Flash Cyber reaches 47.2% pass@1 on CWE-Bench, a patching benchmark run by Collinear, against 47.8% for a leading frontier model, and clears 70% on an internal 20-language discovery set. Access runs through the Fairwind Program for government authorities, critical infrastructure operators, and software maintainers. Gemini 4 Argon, announced September 30, reaches Fairwind partners first.
Google and Mandiant supplied both the numbers and the products
Google and Mandiant produced most of the figures above, and both sell the products that answer them. This account draws on the published agenda, news coverage, and the companies' own posts. Having run marketing at Network Solutions and VeriSign, I know a keynote is built to carry a roadmap. Read the benchmarks as vendor claims until someone outside Google reruns them.
The agenda itself is useful. Royal Hansen's session, "AI is here. What's our next move?", argued for automated defenses, hardware-based controls, safe coding practices, and authorization tied to intent. I have read the abstract and have not seen the talk.
Nick Andersen, acting director of the Cybersecurity and Infrastructure Security Agency, told reporters after his fireside with Sandra Joyce that he needs people who can "pivot from day to day." He wants hires who understand operational technology and control systems well enough to work across electricity, oil and natural gas, and water and wastewater. A faster console still needs that person when an incident moves from a cloud login to a plant-floor system.
The day-one session "From Breach to Lessons Learned" put GitHub's Alexis Wales, Mandiant's Dan Wire, LayerZero's Erez Maharshak, and Trellix's John Fokker on how internal and supply-chain breaches unfolded and how teams briefed leadership, employees, the press, and the board. Two keynotes were not recorded: the Brian Krebs fireside and the panel with security chiefs from Merck, RTX, and Mandiant.
Three Monday checks drawn from Mandiant's cases
Pick one production agent and write down who owns it, what it can call, what it can spend, and who can stop it.
Put a reproduction step in front of the next AI-generated vulnerability ticket, whether that step is Mantis or a sandbox your team builds.
Then ask your incident lead to draft the board note for your last serious event in the time an agent needs to finish a credential run. The missing sentence will show you where your communication plan fails.
