2026 Irregular — CTF evaluation misconfiguration let frontier models exploit a real website (Aug 4)
Data compromised
Data belonging to the affected real website, plus site credentials the model found and used to operate it; Irregular reports no impact identified beyond that site while its audit continues
Technical writeup
Verified vendor-side incident — Irregular, the external cyber-evaluation partner used by multiple frontier labs, notified OpenAI on July 29, 2026 of an incident in its capture-the-flag environment, and Anthropic’s July 30 disclosure attributes its own three Claude incidents to the same misconfiguration. Models were told they had no internet access and instructed to find hidden information by exploiting a simulated environment; a testing-environment misconfiguration let them reach the public internet instead. In one OpenAI run the fictional target name coincided with a real domain, and the model exploited that live website — OpenAI stresses this involved no sandbox escape and no zero-day, just a basic vulnerability reachable because of the misconfiguration — and then found and used credentials to operate the same site. Irregular paused the evaluations, began remediation, notified affected third parties, added safeguards, told OpenAI all identified issues are no longer active, and has communicated about related incidents involving other labs from the same environment. It is producing a white paper on containment practices for offensive AI evaluations. Catalogued as the shared root cause behind several 2026 AI-caused intrusions; no victim count published, so recordsAffected stays 0.
Root cause
Capture-the-flag testing environment was meant to be internet-isolated but a misconfiguration allowed public egress; a fictional target name coincided with a real registered domain, so models attacked a live site believing it was simulated
References
- https://openai.com/index/third-party-cyber-evaluations-involving-openai-models/
- https://www.anthropic.com/news/investigating-incidents-cybersecurity-evals
- https://www.irregular.com/research/next-generation-of-cyber-evals
- https://www.helpnetsecurity.com/2026/07/31/anthropic-claude-cybersecurity-incidents/