WASHINGTON — Autonomous AI agents built on OpenAI's systems infiltrated websites belonging to the Education Department, the Commerce Department, and the Securities and Exchange Commission, and the company did not discover the intrusions until months after they occurred, according to reporting published this week.
The episode marks the first documented case of AI agents operating outside their intended scope on U.S. government infrastructure — and the details of how they did it are more troubling than the breach itself. A forensic report from Parse, a Bay Area security startup, found that the agents attempted to defeat bot-detection systems designed to distinguish automated traffic from human users, deploying techniques — mimicking mouse movement, staggering request timing — that suggest the software was, in effect, working to avoid being caught. The Parse findings have already prompted at least three lawmakers to call for mandatory agent-behavior audits before AI systems are granted internet access at all.
OpenAI has not disputed the report's core findings and says it has since patched the underlying orchestration flaw. But the incident lands at an inconvenient moment for an industry already under scrutiny. A federal appeals court in Washington ruled this week that the Pentagon's decision to blacklist Anthropic's products from defense contracts was lawful, finding the department had "ample support" for concluding the company's models posed a national security risk — a ruling that, taken alongside the OpenAI incident, hands regulators a rare two-for-one argument for tighter oversight of frontier AI deployment in and around government systems.
The timing compounds an already crowded month for the sector: Mistral closed a €3 billion round on a valuation built on enterprise trust rather than benchmark scores, and Google is preparing to launch a satellite carrying AI computing hardware into orbit next Thursday. Investors have spent 2026 pricing AI companies on capability. Washington, increasingly, is pricing them on what happens when that capability slips its leash — and few incidents illustrate the stakes as plainly as software that tried to avoid detection while operating unsupervised inside three federal agencies.