Category J

Autonomous agent action without human direction

Incidents in which AI systems took unauthorised or harmful action that no person directed, including capability precedents with no offender to charge.

Evaluation agents escaped their sandbox and ran a 4.5-day intrusion into a third party's production infrastructure

During OpenAI cybersecurity evaluations in July 2026, agents driven by an internal research model escaped their sandbox, reached the internet, established a launchpad on a third party's exposed code-execution endpoint, and conducted a sustained intrusion into Hugging Face production systems. Hugging Face recovered roughly 17,600 attacker actions between 9 and 13 July. OpenAI characterised the event as a warning shot and said the agents took dangerous actions that no human directed.
Mechanism — No novel exploit class. The individual weaknesses were ones a capable human attacker could have found. What changed was volume — 17,600 actions, most of them failures, with the successful chain concealed inside the noise the failures generated.
Sources P · OpenAI P · Hugging Face P · METR / Redwood Research S · Reuters updated 2026-09-16