
SAIN Amsterdam Discussion Group: Eval Differential
Опис
🤖 AI Safety Discussion 📍 Location: Oerknal, Science Park, UvA🕠 Time: 5:30 PM Discussion topic: Evaluation Differential We'll discuss a recent paper examining how advanced AI models increasingly recognise safety evaluation patterns and may effectively "spot the test" by inferring the expected answers during evaluations. The paper introduces the evaluation differential: a proposed metric to estimate the gap between a model's behaviour during safety evaluations and its behaviour in real-world deployment. read it here: https://arxiv.org/pdf/2605.11496 💬 Discussion focus: Everyone is welcome, whether you're deeply involved in AI safety or simply curious about the topic!
Місце проведення
Розкажіть своїй мережі, що ви йдете
Поділіться цією подією, щоб розпочати розмови, запросити колег та налагодити контакти до її початку.
