Submit event
Tianyi Alex Qiu | Computing Human Ideal Preference, Breaking AI Feedback Loops

Tianyi Alex Qiu | Computing Human Ideal Preference, Breaking AI Feedback Loops

18 Dec 202610:00 - 11:00 America/Los_AngelesOnline0 AttendeesOpen

Description

Foresight Institute’s Computation Group Computing Human Ideal Preference, Breaking AI Feedback LoopsAbstract: How should an AI system assist someone whose beliefs and preferences are still forming while under the influence of such AI systems? Classical approaches based on preference learning and inverse reinforcement learning often mistake transient/instrumental preferences (e.g., money) for stable/terminal ones (e.g., well-being). The latter is stationary while the former changes over time. By assuming the stationarity of both, classical methods are shown to entrench transient/instrumental beliefs and preferences of the human principal, which they would otherwise come to regret. To avoid this problem, we remove the stationarity assumption and propose the alternative goal of learning the human’s ideal preference, i.e., the human belief/preference state that's maximally stable upon sufficient reflection and in face of Socratic counter-persuasion. We show how this target can be formalized and computed in practice, when training language model assistants. Bio: Tianyi co-leads Prevail, an independent group addressing epistemic disempowerment. They develop training/eval interventions (often with humans in the loop) and systemic measures, targeting failure modes like long-term value lock-in and enabling good outcomes like moral/knowledge progress. Tianyi is a researcher at Oxford HAI Lab, 0th-year PhD student at Stanford CS, and former Anthropic AI Safety Fellow. Research projects he led have received two Best Paper Awards: one at ACL'25 and one at the NeurIPS'24 Pluralistic Alignment Workshop. https://tianyiqiu.net/ Computation Group A group of scientists, engineers, and entrepreneurs in computer science, ML, cryptography, and related fields who leverage those technologies to improve voluntary cooperation across humans, and ultimately AIs. Zoom link: https://us02web.zoom.us/j/81623367983

Let your network know you`re going

Share this event to start conversations, invite colleagues, and connect before it begins.

Related events

Artificial Intelligence Platform

from $10
Get tickets