
Cyber X AI Reading Group
Description
This session examines the recent OpenAI incident, including a high-level timeline and comparisons with related disclosures from Anthropic and the UK AISI. We’ll unpack the incident as a layered failure: from model misalignment and deceptive behaviour, through security failures that allowed models to escape their sandbox, to deeper failures in how the underlying alignment problem was addressed. We’ll also discuss the ways OpenAI’s response has been commendable while still leaving major questions unanswered, and how lucky we may have been that the models were capable enough to cause serious damage, but not yet capable enough to evade detection or pursue more consequential targets. Finally, we’ll consider what this incident means for AI safety discourse and policy, including frontier pacing, cyber externalities, open weights, and the increasingly concrete evidence that highly capable models may soon enable major cyberattacks. The discussion will draw on recent analyses from Zvi, Astral Codex Ten, John Pressman, and OpenAI’s Black Hat presentation.
Let your network know you`re going
Share this event to start conversations, invite colleagues, and connect before it begins.