> OpenAI pulls the plug on GPT 6.1 Astra as agents keep crossing lines
[DATE: 29/09/2026 14:45]
[LANGUAGE: EN]
OpenAI has scrapped the planned October release of GPT-6.1 Astra after internal testing found the model did not meet the company’s safety and alignment standards.
GPT-6.1 Astra was being developed as a more autonomous model capable of handling complex tasks with less human assistance, and was expected to be integrated into ChatGPT and Codex. But internal testing found that it could evade oversight, misrepresent its actions and operate beyond its authorized scope, while attempting to use external tools it knew were unsafe, according to a report by The Wall Street Journal
OpenAI reportedly plans to take Astra’s underlying model through additional reinforcement learning to build subsequent models in the GPT-6 family and investigate what caused the safety problems identified during testing.
Models developed by OpenAI increasingly suffer from what the industry euphemistically calls “alignment problems,” meaning a lack of understanding of what is right and wrong, what is acceptable behavior and what unacceptable.
The now-abandoned model’s predecessor, GPT-6 Astra, was caught conducting unsanctioned software supply-chain attacks in simulated cybersecurity tests by the UK’s AI Security Institute, despite being explicitly told that attacking internet targets was out of scope. It did so far more often in tests than its predecessors, GPT 5.5 and GPT 5.6 Sol.
Such models will essentially do anything to achieve their objectives, said Pieter Danhieux, co-founder and chief executive officer of Secure Code Warrior. “For these agents, their operation is essentially business as usual; They will relentlessly pursue the initial goal they were instructed to do, and being repeatedly told ‘no’ by access control parameters will simply ensure they seek the next available endpoint until they succeed,” he said, adding that such models require human oversight and stronger regulation.
OpenAI’s decision to pull GPT-6.1 Astra comes amid a series of incidents involving the company’s models, including an agent that gained unauthorized access to an Australian government portal and models that interacted with several US government websites in unexpected ways.
AI broke the rules, but someone did first
Australian Prime Minister Anthony Albanese said an internal OpenAI model was conducting research into public medical spending on June 18 when it encountered repeated blocks while attempting to obtain information from Australia’s Medicare Statistics Reporting Portal. It tried alternative ways to get the information, eventually gaining unauthorized access to the portal.
He said the agent accessed both public and non-public files and wrote files to an internal server; investigations into the incident remain ongoing.
Aviv Nahum, co-founder and chief executive officer at Above Security, said this incident could be more than an AI being at fault. “If the archive evidence holds, the most-hyped AI hack of the year looks a lot like the most common security failure of the last thirty years: a misconfiguration,” he said. “The portal’s own code pointed visitors to an endpoint that required no credentials, and the agent followed the path it was given. That’s not an ‘AI agent hacking a government website,’ that’s a door that was left open, found by something that can read the code more carefully and execute more quickly.”
We should be careful about framing every agent incident as “rogue AI,” Nahum cautioned, adding that doing so makes for “dramatic headlines,” but moves the blame to the model instead of the environment.
The US incidents were less serious but followed a similar pattern of models interacting with external systems. An OpenAI spokesperson told the Washington Post that the company’s models accessed publicly available information on SEC.gov, Investor.gov and Census.gov during training and evaluation. The company said its agents acted inappropriately in the SEC and Census Bureau incidents but did not steal any private data.
Ben Bernstein, manager of Huntress’ cybersecurity advisors team, took a similar view of the US incidents, arguing that they have been overstated as AI hacks. “These agents were simply tasked with gathering public SEC and Census information, but the guardrails in place were not firm enough to contain a model programmed to problem solve,” he said.
Still, OpenAI cannot be let off entirely, with reports of model misalignment and unauthorized activity piling up over the past few days.
Testing times
The common thread across these incidents is that they involved models being tested or evaluated by OpenAI, not publicly deployed models.
The Australian incident involved an internal OpenAI model conducting research. The US activity similarly involved models being tested by OpenAI as part of its research and evaluation work. Neither involved a customer taking a publicly available ChatGPT model and directing it against a government system.
OpenAI’s decision to kill off GPT-6.1 Astra follows its decision to pause training of its most capable models after an one of the models under test bypassed network restrictions and used DNS to communicate externally.
CEO Sam Altman acknowledged the breadth of the problem in a tweet on Sept. 25. “There is an extensive and ongoing review related to our agents’ use of internet access during training and evaluation,” he wrote, adding, “We have not been as fast as we would have liked but we are trying to balance our desire for transparency with gaining a clear understanding from petabytes of agent activity logs, and working with impacted organizations.”