OpenAI pulls GPT-6.1 Astra release after model falls short on safety

OpenAI has said it will not release GPT-6.1 Astra, the latest version of its agentic AI model, after the system fell short of the company’s own safety standards.
The model, which can browse the web and use apps by itself, “didn’t quite meet the bar”, according to Saachi Jain, head of safety systems at OpenAI. The decision was first reported by the Wall Street Journal.
Jain said the model fell short on “staying within scope and authorisation and how it communicates back to the user about the type of work it’s done”.
OpenAI released the flagship GPT-6 Astra model in September, describing it as the result of “years of research and big bets”. The company holds its annual DevDay developer conference in San Francisco today. It is not clear whether a new version of Astra will be announced.
OpenAI also issued an update today on incidents in June in which its models accessed Australian government websites and systems without authorisation.
Anthony Albanese, the Australian prime minister, made the incidents public last week. He told a press conference in New York that OpenAI first notified the government on 10 September, by email to a public mailbox, and criticised the company for not contacting officials directly.
In a statement, OpenAI said it was sorry for the incident and “should have handled our response better”. It said Services Australia, the NSW Bureau of Crime Statistics and Research, the Victorian Department of Health and the Australian Institute of Health and Welfare were affected.
The company said it launched investigations when it became aware of issues in mid-August and notified the affected organisations between 10 and 24 September. It said it should have shared early findings more promptly.
OpenAI said it would fund cyber security measures, offer dedicated support to the agencies affected and set up a taskforce to manage the risks from increasingly advanced AI agents. A senior OpenAI executive will attend a hearing of Australia’s Joint Select Committee on AI on 6 October.
In July, OpenAI said its AI systems had hacked into Hugging Face, the open-source developer hub. Yesterday Nvidia, which agreed to buy Hugging Face for $12.9bn (£9.74bn) earlier this month, released safety tools for autonomous AI agents that it said could have prevented that breach.
Calls for independent testing
Prof Tony Cohn, foundational models theme lead at the Alan Turing Institute, said OpenAI’s decision was “a welcome sign that they are taking safety concerns seriously”. He added that “safety should not be left purely in the hands of the developers: it should also be monitored and verified through independent government-approved regulators”.
Prof Gina Neff, of the Minderoo Centre for Technology and Democracy at the University of Cambridge, told the BBC it was “critical” to have independent tests of AI models by labs such as the UK’s AI Security Institute, which evaluates frontier systems on a voluntary basis, because “these companies have proven that we can’t rely solely on them for our safety”.
Jess Whittlestone, a senior adviser on AI policy at the Centre for Long-Term Resilience think tank, said: “I think it’s kind of crazy that companies are continuing to push forward with developing these capabilities when we’ve already seen over the last couple of months of incidents that they’re nowhere near safe and controlled enough.”
Earlier this year Anthropic said it would not publicly release a Claude model, Mythos, because it was too good at finding dormant software bugs. It released a version to the public several months later.
Anthropic plans to warn potential investors in its initial public offering that the technology may pose “catastrophic or existential risks to humanity”, Reuters reported today, citing a prospectus it had seen.
President Donald Trump and House Speaker Mike Johnson are due to host technology executives at the White House later today to discuss AI regulation. Trump has described concerns about AI’s risks as a “hoax”.














