
Discussion Group - week 3 - Why alignment is hard
Description
Part three of a four-week AI safety discussion series: Week 1 | Why AI Safety matters · Week 2 | What failure could look like at scale · Week 3 | Why alignment is hard · Week 4 | What's being done about it? Every session stands on its own, so join one or all. Who it's for: Anyone curious about AI safety and governance, whether you're new to the topic, a healthy skeptic, or already deep in it. No technical background needed, and no need to have joined earlier weeks. You don't need to have answers, just curiosity. Discussion group curriculum:https://docs.google.com/document/d/160PpLjFN9SaZja0njeKIQHnD-M1Uy4718dNRDyhaucs/edit?usp=sharing We can train AI systems to be very capable. Making sure they do what we actually intend, which is what people call alignment, is much harder. AI systems find loopholes in the goals we give them. Deceptive behaviour can survive the training meant to remove it. Tests don't reliably tell us how a system will behave in the real world. And institutions struggle to keep pace. Drawing on DeepMind's work on specification gaming, Anthropic's "Sleeper Agents" paper (in which researchers deliberately trained deceptive behaviour into models to see whether safety training could remove it), and the International AI Safety Report 2026, we'll explore where the difficulty comes from. Questions we'll explore Format (90 min): Open discussion in small groups and plenary. No presentations, and no wrong questions. Readings:★ = mandatory (~2 hours). The rest is optional. By attending, you agree to being filmed or photographed, which may be used for social media, website, and newsletter content. If you wish to attend but do not want to be photographed, please during the event let a member of our staff know so we can accommodate this.
Event location
Let your network know you`re going
Share this event to start conversations, invite colleagues, and connect before it begins.











![B2B Lead Generation in AI era [Workshop]](https://aiprojects.nyc3.cdn.digitaloceanspaces.com/allai/banners/ai_data/29.jpg)