Mandiant says one runaway AI agent racked up $50,000 in cloud charges in an hour, with no hacker involved.

Security
Google and Mandiant spent September 15 and 16 in Washington selling machine-speed defense. The evidence they brought says the first fix is knowing who owns the agents you already run.
By Shashi Bellamkonda  |  October 9, 2026
$50,000
Cloud charges from one runaway accounting agent in under an hour (Mandiant)
< 6 hours
Credential-harvesting campaign run with a coding agent (Google Threat Intelligence Group)
~100
Internal repositories hit by the Shai-Hulud worm after one hijacked coding-assistant session (Mandiant)
47.2%
Gemini 3.8 Flash Cyber pass@1 on CWE-Bench, against 47.8% for a leading frontier model (Google)

Fifteen thousand calls to paid cloud services in under an hour ran up about $50,000 in charges, and no attacker was involved.

An accounting agent running on artificial intelligence entered a runaway execution loop, made more than 15,000 high-cost application programming interface calls, and disrupted live business transactions. Mandiant published the case in its AI Risk and Resilience report on September 16, the second day of the Cyber Defense Summit at the Ronald Reagan Building in Washington, D.C. Mandiant calls it a consequence of weak controls. My read is simpler. A spend ceiling would have capped the bill. An owner with a kill switch would have ended it in minutes.

An accounting agent made 15,000 high-cost calls in under an hour

Every security vendor wants you to picture the autonomous attacker. The largest dollar figure in this story came from your own software, running under weak controls.

A spend ceiling would have capped the bill. An owner with a kill switch would have ended it in minutes.

Google's threat tracker clocks a credential-theft campaign at six hours

Google Threat Intelligence Group published its third-quarter 2026 AI Threat Tracker, which Help Net Security covered on September 8. It describes a Mandiant investigation from the second quarter. A financially motivated actor with access to an organization's cloud infrastructure used an AI coding chatbot, a prompt, and a set of agent instructions to plan and run a mass credential-harvesting campaign in under six hours. Thousands of third-party credentials were taken. The agent handled vulnerability scanning, troubleshooting, and Internet Protocol address rotation without manual intervention.

The same report says Google has not yet observed a fully autonomous pipeline run against targets in the wild. Defend against what exists today: stolen credentials moving faster than your response process.

Mandiant's report adds a second case. An attacker hijacked an active AI coding-assistant session at an unnamed software-as-a-service provider and used it to install an infostealer through a poisoned PyPI package. The attacker took GitHub OAuth tokens, then spread the Shai-Hulud worm across about 100 internal repositories. Mandiant does not say how the session was taken over.

Google answers a sub-7% true-positive rate with Mantis and Fairwind

Google puts the true-positive rate of naive AI code scanning below 7%, according to IT Brief's coverage of the announcement. Its answer is Mantis, a set of agent skills that finds a suspected flaw, reproduces it in a sandbox, writes a patch, and runs regression tests against it. Google open-sourced the core skills in late June under an Apache 2.0 license, and the repository states it is not an officially supported product. Its README warns that the suite generates and runs code and belongs only in isolated environments.

You do not need the repository to take the idea. Send developers a reproduced bug, and stop sending them a model's confidence score.

Google's cyber-tuned model follows the same logic. Gemini 3.8 Flash Cyber reaches 47.2% pass@1 on CWE-Bench, a patching benchmark run by Collinear, against 47.8% for a leading frontier model, and clears 70% on an internal 20-language discovery set. Access runs through the Fairwind Program for government authorities, critical infrastructure operators, and software maintainers. Gemini 4 Argon, announced September 30, reaches Fairwind partners first.

Google and Mandiant supplied both the numbers and the products

Google and Mandiant produced most of the figures above, and both sell the products that answer them. This account draws on the published agenda, news coverage, and the companies' own posts. Having run marketing at Network Solutions and VeriSign, I know a keynote is built to carry a roadmap. Read the benchmarks as vendor claims until someone outside Google reruns them.

The agenda itself is useful. Royal Hansen's session, "AI is here. What's our next move?", argued for automated defenses, hardware-based controls, safe coding practices, and authorization tied to intent. I have read the abstract and have not seen the talk.

Nick Andersen, acting director of the Cybersecurity and Infrastructure Security Agency, told reporters after his fireside with Sandra Joyce that he needs people who can "pivot from day to day." He wants hires who understand operational technology and control systems well enough to work across electricity, oil and natural gas, and water and wastewater. A faster console still needs that person when an incident moves from a cloud login to a plant-floor system.

The day-one session "From Breach to Lessons Learned" put GitHub's Alexis Wales, Mandiant's Dan Wire, LayerZero's Erez Maharshak, and Trellix's John Fokker on how internal and supply-chain breaches unfolded and how teams briefed leadership, employees, the press, and the board. Two keynotes were not recorded: the Brian Krebs fireside and the panel with security chiefs from Merck, RTX, and Mandiant.

Three Monday checks drawn from Mandiant's cases

Pick one production agent and write down who owns it, what it can call, what it can spend, and who can stop it.

Put a reproduction step in front of the next AI-generated vulnerability ticket, whether that step is Mantis or a sandbox your team builds.

Then ask your incident lead to draft the board note for your last serious event in the time an agent needs to finish a credential run. The missing sentence will show you where your communication plan fails.

CIO/CTO Viability Question
Which production agents in your environment can answer who owns them, what they can spend, and who can stop them today, and which finding path still turns a model's score into a ticket before anyone has reproduced the bug? If locating the owner takes a meeting, the bill arrives first.
Sources
Mandiant. "Cyber Defense Summit 2026: Keynotes." cyberdefensesummit.mandiant.com. September 2026.
Pogorelec, Anamarija. "One Runaway AI Agent Racked Up a $50,000 Cloud Bill." Help Net Security, 16 Sept. 2026, helpnetsecurity.com.
Markovic, Sinisa. "Threat Actors Are Giving AI Agents a Bigger Role in Cyberattacks." Help Net Security, 8 Sept. 2026, helpnetsecurity.com.
Geller, Eric. "CISA Looks to Recruit General Infrastructure Security Experts Rather Than Sector-Focused Advisers." Cybersecurity Dive, 16 Sept. 2026, cybersecuritydive.com.
Doshi, Tulsee, and Raluca Ada Popa. "Introducing Gemini 3.8 Flash and 3.8 Flash Cyber." Google, 2 Sept. 2026, blog.google.
Betz, Chris, and Ruchi Shah. "Cloud CISO Perspectives: How Google Cloud Security Uses AI Internally." Google Cloud Blog, 30 June 2026, cloud.google.com.
Mitchell, Sean. "Google Cloud Uses AI Agents to Secure Software Lifecycle." IT Brief UK, 30 June 2026, itbrief.co.uk.
Google. "Mantis." GitHub, github.com/google/mantis.
"Attacker Hijacks AI Coding Assistant." The Hacker News, Sept. 2026, thehackernews.com.
Disclaimer: This blog reflects my personal views only. Content does not represent the views of my employer, Info-Tech Research Group. AI tools may have been used for brevity, structure, or research support. Please independently verify any information before relying on it.