Brief independent investigation of agents’ behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incident

beSpacific 2026-08-31

Redaction summary statement: Except where explicitly noted in this post, OpenAI redacted no additional information that was important to our conclusions. Two METR staff members (Hjalmar Wijk and Ajeya Cotra) and a Redwood Research staff member contracting with METR (Ryan Greenblatt) worked on premises at OpenAI over a total of six days[1] to attempt to form an independent understanding of model behavior observed during the recent incidentin which OpenAI agents coordinated a multi-day hack of Hugging Face on a shared unsanctioned “message board.” Our investigation focused mostly[2] on the period between July 7th and July 13th. The earlier incidents from training and the subsequent compromise of OpenAI infrastructure described in OpenAI’s recent Black Hat presentation were out of scope, as was OpenAI’s investigation process and planned remediation. Per our standard policy, we did not take payment from OpenAI for this independent assessment.[3]