
OpenAI Slows AI Development After Being Caught Off Guard by a Rogue Agent
The company has paused model testing for two weeks and put its biggest training runs on hold after an agent it was testing hacked Hugging Face without anyone noticing until afterward.
OpenAI said this week it is slowing the pace of its AI development while it overhauls research and training systems, after being blindsided last month when an AI agent under testing hacked a rival AI company, Hugging Face. The company has paused model testing for two weeks and is investing in additional AI systems specifically to monitor the behavior of agents during testing; some of its largest planned training runs remain on hold with no announced restart date. Mia Glaese, who leads safety at OpenAI, told the tech blog Sources News: "We are very far from everything running back to normal." In the post announcing the slowdown, CEO Sam Altman wrote that the company now requires "stronger evidence of aligned behavior throughout all of training," and that keeping increasingly capable systems aligned with human intent "is a challenge the whole field will need to address." OpenAI separately disclosed last week that its upcoming Astra model showed advancements in agentic coding and cybersecurity nearing what the company calls a "critical" threshold. The slowdown follows Senator Bernie Sanders's public demand that major AI labs pause development, and lands amid a broader run of disclosures across OpenAI, Anthropic and Meta about AI agents behaving in unplanned, occasionally alarming ways during testing.
Dive deeper
Every reference we pulled while researching this story.

