Skip to content

OpenAI took a two-week break from training new models to improve safety. Now what?

We'll soon find out if improvements to OpenAI's alignment, monitoring, and security practices were enough to prevent another Hugging Face incident.

OpenAI took a two-week break from training new models to improve safety. Now what?
Photo by Declan Sun / Unsplash

After losing control of some of its most advanced AI technology, which led to an unprecedented attack on Hugging Face's infrastructure, OpenAI said Tuesday that it paused reinforcement learning training of new models for two weeks in order to correct the missteps that led to the incident.

During the slowdown, the company "further hardened and red-teamed our research environments and expanded the coverage of our monitoring systems," it said in a blog post. "Our largest planned frontier [reinforcement learning] run remains on hold while we conduct smaller-scale training and evaluations to assess model behavior, validate our safeguards, and establish more evidence of alignment before proceeding."

OpenAI said it re-examined its policies and procedures covering three parts of the model training process: monitoring, alignment, and security measures. More robust safeguards in any of those three areas might have prevented the Hugging Face incident, in which AI agents undergoing testing found their way to the internet through a poorly configured sandbox and ran amok in Hugging Face's infrastructure for a shockingly long amount of time before anybody figured out what was going on.

This content is for members only

Subscribe
Add The Stack on Google