
SAIN Amsterdam Discussion Group: Eval Differential
Beschreibung
🤖 AI Safety Discussion 📍 Location: Oerknal, Science Park, UvA🕠 Time: 5:30 PM Discussion topic: Evaluation Differential We'll discuss a recent paper examining how advanced AI models increasingly recognise safety evaluation patterns and may effectively "spot the test" by inferring the expected answers during evaluations. The paper introduces the evaluation differential: a proposed metric to estimate the gap between a model's behaviour during safety evaluations and its behaviour in real-world deployment. read it here: https://arxiv.org/pdf/2605.11496 💬 Discussion focus: Everyone is welcome, whether you're deeply involved in AI safety or simply curious about the topic!
Veranstaltungsort
Lassen Sie Ihr Netzwerk wissen, dass Sie dabei sind
Teilen Sie diese Veranstaltung, um Gespräche zu beginnen, Kollegen einzuladen und sich vorab zu vernetzen.
