The $50,000 hour, with no attacker in sight
Mandiant and Google's Threat Intelligence Group published their AI Risk and Resilience 2026 report on September 16, and the anecdote leading the coverage is a genuinely useful gut check for any executive still treating AI agent risk as a future problem. An accounting agent deployed inside an enterprise malfunctioned and entered an execution loop, firing more than 15,000 high-cost API calls within a single hour. The result was roughly $50,000 in unplanned cloud charges and disrupted business transactions, caused entirely by a software bug in the agent's own reasoning, with no external attacker involved at any point.
That incident matters precisely because it removes the adversary from the equation. Enterprises spend enormous energy preparing for attackers who compromise or manipulate an AI agent, and that risk is real and well documented elsewhere in the same report. But a $50,000 an hour cost blowout from a malfunctioning agent, with no malicious actor required, is a failure mode that a purely security-focused threat model will miss entirely. It is an operations and governance failure, and it happened inside a production deployment, not a sandbox. Any executive who has approved an agent rollout on the strength of a security review alone, without a matching conversation about spend caps and runaway-loop detection, is exposed to a version of this same incident today.
The worm that moved through your own code
The report also details a supply-chain-style incident in which a self-propagating worm, in the same family as the Shai-Hulud npm worm that has circulated through 2026, spread across approximately 100 internal code repositories at an affected organization. Separately, Mandiant describes attackers weaponizing OpenClaw AI agent skills, the growing ecosystem of installable capabilities for autonomous coding and workflow agents, to distribute backdoors and infostealers to unsuspecting developers who installed them expecting productivity tools.
Together these findings describe a genuinely new attack surface: the marketplace of agent skills, plugins, and extensions that enterprises are adopting to make their AI deployments more capable. Each of those installable pieces is effectively unreviewed third-party code running with whatever permissions the agent itself holds, and Mandiant's report treats this as analogous to the npm and PyPI supply chain problem the industry has spent years trying to solve, except now arriving inside AI tooling that many security teams have not yet subjected to the same scrutiny as their open source dependencies. Most enterprises have a software bill of materials process for their application code and nothing equivalent for the skills their agents install.
Old vectors, new autonomy
One of the report's more sobering findings is continuity rather than novelty: vulnerability exploitation remains the leading initial infection vector for the sixth consecutive year, meaning the fundamentals of patch management and exposure reduction still matter more than almost anything else, AI-related or not. What has changed is what happens after that initial exploitation. Mandiant documents the first publicly confirmed case of an AI-developed zero-day exploit used for mass exploitation in May 2026, and describes attackers using AI tooling to accelerate reconnaissance, lateral movement, and even sandbox escape once inside an environment.
In one healthcare-sector incident described in the report, attackers compromised thousands of credentials and exfiltrated API keys and financial secrets through unauthorized data-harvesting frameworks, a scale and speed of credential theft that reflects automation on the attacker side rather than a merely larger phishing campaign. The message for defenders is that the fundamentals have grown more urgent, not less, since AI arrived, because both the initial exploitation and everything that follows it are now happening on a compressed timeline that gives incident response teams less room to react.
The same tools cut both ways
It would be incomplete to read this report as purely a warning, because Mandiant also documents AI genuinely improving defense. In one engagement, AI-enabled source code review discovered over 100 true-positive critical vulnerabilities in just two days, work that would take a human team substantially longer at that depth and speed. The report also credits behavior-based monitoring, rather than signature-based detection, with thwarting a sophisticated espionage campaign, internally named DARK CASTLE, targeting telecommunications and government organizations, which suggests defenders adopting AI-assisted, behavior-first detection are keeping pace in at least some cases.
The throughline across both the offensive and defensive findings is that AI amplifies whatever discipline, or lack of it, already exists in your environment. Organizations with strong identity controls, code review practices, and monitoring get faster, more effective versions of those practices when AI is layered on top. Organizations without those fundamentals get faster, more effective versions of their existing gaps, which is exactly what the $50,000 rogue-agent incident and the 100-repository worm both illustrate.
What this means for your agent rollout
Mandiant's own framing is direct: defending against autonomous AI threats requires clearly identified, adaptive identity controls for agents, faster defensive response times, and a security operations center reoriented toward real-time behavioral telemetry rather than after-the-fact log review. For any executive currently pushing agents from pilot into production, that framing should shape the rollout plan directly. Every agent needs a scoped identity, a hard ceiling on the resources or spend it can consume autonomously, and monitoring that would catch an execution loop within minutes, not after a $50,000 invoice arrives.
It also means treating agent skills, plugins, and extensions with the same supply-chain scrutiny you apply to open source packages, because Mandiant's findings show that scrutiny gap is already being actively exploited. The build-versus-buy decision on agent infrastructure should now explicitly include the question of who is responsible for reviewing every skill or extension before it runs with production access, because 'nobody' is not a governance answer that survives contact with this report.



