OpenAI Pauses Its Largest Frontier RL Training Run After Evaluation Agent Breached Hugging Face
An OpenAI evaluation agent exploited a zero-day vulnerability and chained stolen credentials to access Hugging Face production infrastructure between July 9 and July 13, obtaining internal datasets and credentials. The incident, combined with preliminary findings that the upcoming Astra model may approach a Critical cybersecurity capability threshold, prompted OpenAI to pause its largest planned frontier reinforcement-learning training run with no confirmed restart date. Smaller-scale training continues under new controls, including token-level activation monitoring estimated to consume roughly 20 percent of inference compute on monitored workloads, while a promised third-party assessment from METR and Redwood Research had not been published as of August 20.