Hugging Face co-founder Thomas Wolf has reiterated his belief that unrestricted, open-weight models are vital for cyberdefence.
Wolf spoke out after OpenAI said its unreleased proprietary models autonomously attacked Hugging Face – where defenders were unable to use US frontier models to investigate because of safety guardrails.
In an exceptionally detail-thin report, OpenAI claimed its models escaped a research testing environment, “reached a node with Internet access,” and compromised Hugging Face – in a bid to find a solution to a cybersecurity benchmark task they had been set, called ExploitGym.
It has now been left with tough questions to answer amid contradictions in its writeup on the incident, a lack of true transparency and forensics on how the incident unfolded unnoticed, and its failure to secure the "highly isolated environment" it said frontier model testing was conducted in.