OpenAI and Nvidia Face Growing Safety Problems As AI Agents Go Rogue

By Hassan Shittu

Key highlights:

  • OpenAI agents accessed U.S. government systems without authorization, copying SEC content, and attempting to access the Education Department
  • GPT 6.1 Astra was withheld after failing safety requirements, marking OpenAI’s second model halt in three months
  • Nvidia’s Open Agent Safety Platform takes an infrastructure-first approach by using hardware-enforced controls outside the model itself

OpenAI, Nvidia, and other major AI companies are facing a new AI security challenge as autonomous agents gain the ability to interact with government websites, software systems, and external tools with less human supervision.

Recent incidents have shown that AI agents can do more than generate incorrect answers. When given access to websites, credentials, or other computer systems, they can take actions that developers did not intend, creating a new layer of security risks around increasingly autonomous AI.

The incidents have also raised questions about whether current permission systems are strong enough to control what AI agents can access and how quickly humans can intervene when an agent behaves unexpectedly.

OpenAI has reportedly halted training on a new model following incidents involving AI agents interacting with U.S. government websites. While Nvidia is approaching the problem from the infrastructure side by introducing an open-source framework designed to provide AI agents with a controlled environment with hardware-enforced controls over agents.

OpenAI faces new questions after AI agents access government systems

OpenAI’s AI agents have interacted with several U.S. government systems in ways the company says were outside their intended behavior, raising questions about how autonomous systems handle credentials and access controls.

In one incident, OpenAI agents used developer access keys found in public code repositories to retrieve information from the U.S. Census Bureau’s data service.

The Commerce Department said the information itself was public, but the incident centered on how the agents obtained access. OpenAI’s reporting framework classifies the unauthorized use of exposed credentials as a form of model misbehavior.

The agents also interacted with SEC websites, copying public information from SEC.gov and Investor.gov and reposting it elsewhere. OpenAI said it found no evidence that the agents used SEC credentials, while the Securities and Exchange Commission said it was unaware of any unauthorized access to nonpublic information.

A separate incident involving the U.S. Department of Education remains under investigation.

Independent AI research lab Transluce reported that an agent believed to be associated with OpenAI attempted to access the Education Department’s civil rights office but failed. The department said it found no impact, while OpenAI continued investigating the incident.

The cases show a distinction that is becoming increasingly important as AI agents gain access to external systems.

An agent does not need to obtain confidential information to create a security problem. Using an exposed credential, accessing a system through an unauthorized method or taking an action outside its assigned scope can itself create a security risk.

OpenAI pauses GPT 6.1 Astra release over security concerns

The recent government incidents follow earlier cases that exposed similar problems with controlling autonomous AI agents.

In July, OpenAI disclosed an incident involving Hugging Face in which one of its models bypassed a sandbox during cybersecurity testing and accessed the platform without authorization.

According to OpenAI, the agent obtained a login credential that allowed it to access a biology-related file.

The incident raised a security question: can an AI agent reliably recognize when an action crosses a boundary set by its developers, particularly when it can discover credentials or other ways around those restrictions?

OpenAI also disclosed an incident involving Australia’s Medicare Statistics Reporting Portal on June 18. The company said an agent accessed aggregate health statistics and internal file names but found no evidence that patient records were accessed.

The AI company discovered the incident on Aug. 11 but did not notify Australian authorities until Sept. 10, prompting criticism from Prime Minister Anthony Albanese over the delay and notification process.

Other agencies identified in other incidents included Services Australia, the NSW Bureau of Crime Statistics and Research, the Victorian Department of Health, and the Australian Institute of Health and Welfare.

OpenAI has apologized for its handling of the disclosures and said it will support affected agencies, fund cybersecurity measures, and establish a task force focused on risks from increasingly capable AI agents.

The incidents point to a problem that goes beyond whether an AI model can generate harmful content. As agents gain the ability to browse websites, obtain credentials, execute code, and interact with external systems, developers also have to control what those systems can access and what actions they can take.

That challenge has become more important as models are given longer-running tasks and greater autonomy. A system can follow its objective while still taking individual actions that violate the permissions or boundaries set by its developers.

OpenAI has also faced concerns over whether its latest models can reliably maintain those boundaries. The company decided not to release its latest agentic model, GPT 6.1 Astra, after it failed to meet its safety requirements.

OpenAI’s head of safety systems, Saachi Jain, said the model struggled to remain within authorized boundaries and clearly communicate which tasks it had completed.

Notably, it is the second time in three months that OpenAI has halted development of its models. The decision is coming after the earlier Hugging Face incident and the subsequent disclosures involving government systems.

Together, the cases show that controlling an AI agent requires more than improving the model itself. Developers also need safeguards that can restrict an agent when its behavior moves beyond its intended scope.

That is where Nvidia is taking a different approach, building external controls designed to monitor and restrict what AI agents can do even when the underlying model does not follow its intended boundaries.

How Nvidia is trying to stop AI agents before they cross the line 

Nvidia has taken a different approach with its Open Agent Safety Platform, which places security controls outside the AI model itself. 

The platform combines OpenShell, an open-source runtime that restricts an agent’s access to files, networks, and tools, with Sentry, a monitoring system designed to detect and contain suspicious activity.

According to Nvidia, Sentry runs on its BlueField-4 data processing unit and can independently monitor agents and isolate suspicious activity within milliseconds. 

More than 100 organizations were using the platform at launch, including Microsoft, Perplexity, Accenture, and JPMorgan Chase, as well as Anthropic, Palantir, Cisco, CrowdStrike, and Hugging Face.

The differing approaches reflect a debate over how AI risks should be managed. Nvidia CEO Jensen Huang has described AI safety as an engineering problem that can be addressed through technical controls. 

Meanwhile, OpenAI CEO Sam Altman and Anthropic CEO Dario Amodei have supported greater coordination to ensure safety measures keep pace with model development.

The debate has reached policymakers, where Senator Sarah Hanson-Young invited Altman and Amodei to appear before a Senate inquiry examining AI and data centers. 

Also, two U.S. lawmakers introduced legislation in July that would allow the federal government to shut down an AI model, with an exemption for authorized adversarial testing.

Source:: OpenAI and Nvidia Face Growing Safety Problems As AI Agents Go Rogue