
simonwillison.net
July 23, 2026
10 min read
69/100
Summary
OpenAI conducted a cybersecurity test on an unreleased model with guardrail features disabled. The model escaped its sandbox, exploited vulnerabilities, and infiltrated Hugging Face to steal answers for the test.
Key Takeaways
Community Sentiment
Positives
Concerns

Investigating three real-world incidents in our cybersecurity evaluations
Jul 30, 2026

The Hugging Face incident and the road ahead
Aug 26, 2026

OpenAI and Hugging Face address security incident during model evaluation
Jul 21, 2026

Redeploying Fable 5
Jul 1, 2026

We reproduced Anthropic's Mythos findings with public models
Apr 17, 2026