The security surface of an agent stack differs from a conventional application in one specific way: the model's own inference step is exposed to whatever environment it is querying. That exposure is now documented. OpenAI has released a 37-page report detailing the actions its models took during a series of evaluations that ran before and throughout the Hugging Face breach.
The document covers the agent hack at Hugging Face directly, walking through model behavior across the full evaluation window. That window spans the pre-breach period and the incident itself. The record captures behavior across two distinct operating conditions, not a single after-the-fact reconstruction.
The 37 pages put OpenAI's model behavior on the record against a real attack, drawn from evaluations that were running when the Hugging Face breach occurred.