OpenAI
We'll soon find out if improvements to OpenAI's alignment, monitoring, and security practices were enough to prevent another Hugging Face incident.
After losing control of some of its most advanced AI technology, which led to an unprecedented attack on Hugging Face's infrastructure, OpenAI said Tuesday that it paused reinforcement learning training of new models for two weeks in order to correct the missteps that led to the incident.
During the slowdown, the company "further hardened and red-teamed our research environments and expanded the coverage of our monitoring systems," it said in a blog post. "Our largest planned frontier [reinforcement learning] run remains on hold while we conduct smaller-scale training and evaluations to assess model behavior, validate our safeguards, and establish more evidence of alignment before proceeding."
OpenAI said it re-examined its policies and procedures covering three parts of the model training process: monitoring, alignment, and security measures. More robust safeguards in any of those three areas might have prevented the Hugging Face incident, in which AI agents undergoing testing found their way to the internet through a poorly configured sandbox and ran amok in Hugging Face's infrastructure for a shockingly long amount of time before anybody figured out what was going on.
Join peers managing over $100 billion in annual IT spend and subscribe to unlock full access to The Stack’s analysis and events.
Already a member? Sign in