
SAIN Amsterdam Discussion Group: Eval Differential
ืชืืืืจ
๐ค AI Safety Discussion ๐ Location: Oerknal, Science Park, UvA๐ Time: 5:30 PM Discussion topic: Evaluation Differential We'll discuss a recent paper examining how advanced AI models increasingly recognise safety evaluation patterns and may effectively "spot the test" by inferring the expected answers during evaluations. The paper introduces the evaluation differential: a proposed metric to estimate the gap between a model's behaviour during safety evaluations and its behaviour in real-world deployment. read it here: https://arxiv.org/pdf/2605.11496 ๐ฌ Discussion focus: Everyone is welcome, whether you're deeply involved in AI safety or simply curious about the topic!
ืืืงืื ืืืืจืืข
ืชื ืืจืฉืช ืฉืื ืืืขืช ืฉืืชื ืืืื
ืฉืชืฃ ืืืจืืข ืื ืืื ืืืชืืื ืฉืืืืช, ืืืืืื ืขืืืชืื ืืืืชืืืจ ืืคื ื ืชืืืืชื.
