
SAIN Amsterdam Discussion Group: Eval Differential
Description
🤖 AI Safety Discussion 📍 Location: Oerknal, Science Park, UvA🕠 Time: 5:30 PM Discussion topic: Evaluation Differential We'll discuss a recent paper examining how advanced AI models increasingly recognise safety evaluation patterns and may effectively "spot the test" by inferring the expected answers during evaluations. The paper introduces the evaluation differential: a proposed metric to estimate the gap between a model's behaviour during safety evaluations and its behaviour in real-world deployment. read it here: https://arxiv.org/pdf/2605.11496 💬 Discussion focus: Everyone is welcome, whether you're deeply involved in AI safety or simply curious about the topic!
Event location
Let your network know you`re going
Share this event to start conversations, invite colleagues, and connect before it begins.
