Publications

Also on Google Scholar.

2026

Active Teacher Selection for Reward Learning

Rachel Freedman, Justin Svegliato, Kyle Wray, Stuart Russell
Transactions on Machine Learning Research (TMLR 2026); contributed talks at PRL @ AAAI 2025 and TAIS 2025
2026

Adaptive Pluralistic Alignment: A Pipeline for Dynamic Artificial Democracy

Rachel Freedman
Pluralistic Alignment Workshop (at ICML 2026)
2025

Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback

Stephen Casper, Xander Davies, Claudia Shi, Thomas Krendl Gilbert, Jérémy Scheurer, Javier Rando, Rachel Freedman, Tomasz Korbak, David Lindner, Pedro Freire, Tony Wang, Samuel Marks, Charbel-Raphaël Segerie, Micah Carroll, Andi Peng, Phillip Christoffersen, Mehul Damani, Stewart Slocum, Usman Anwar, Anand Siththaranjan, Max Nadeau, Eric J. Michaud, Jacob Pfau, Dmitrii Krasheninnikov, Xin Chen, Lauro Langosco, Peter Hase, Erdem Bıyık, Anca Dragan, David Krueger, Dorsa Sadigh, Dylan Hadfield-Menell
International Conference on Learning Representations (ICLR 2025); Transactions on Machine Learning Research (TMLR 2023)
2024

Linear Probe Penalties Reduce LLM Sycophancy

Henry Papadatos, Rachel Freedman
Socially Responsible Language Modelling Research Workshop (SoLaR) (at NeurIPS 2024)
2024

Position: Social Choice Should Guide AI Alignment in Dealing with Diverse Human Feedback

Vincent Conitzer, Rachel Freedman, Jobst Heitzig, Wesley H. Holliday, Bob M. Jacobs, Nathan Lambert, Milan Mossé, Eric Pacuit, Stuart Russell, Hailey Schoelkopf, Emanuel Tewolde, William S. Zwicker
International Conference on Machine Learning (ICML 2024)
2023

Active Reward Learning from Multiple Teachers

Peter Barnett, Rachel Freedman, Justin Svegliato, Stuart Russell
SafeAI Workshop (at AAAI 2023)
Best Paper Award Finalist at SafeAI 2023
2022

The Expertise Problem: Learning from Specialized Feedback

Oliver Daniels-Koch, Rachel Freedman
ML Safety Workshop (at NeurIPS 2022)
AI Risk Analysis Award at ML Safety 2022
2020

Choice Set Misspecification in Reward Inference

Rachel Freedman, Rohin Shah, Anca Dragan
AISafety Workshop (at IJCAI-PRICAI 2020)
Best Paper Award at AISafety 2020
2020

Benefits of Assistance over Reward Learning

Rohin Shah, Pedro Freire, Neel Alex, Rachel Freedman, Dmitrii Krasheninnikov, Lawrence Chan, Michael Dennis, Pieter Abbeel, Anca Dragan, Stuart Russell
Cooperative AI Workshop (at NeurIPS 2020)
Best Paper Award at CoopAI 2020
2020

Adapting a Kidney Exchange Algorithm to Align with Human Values

Rachel Freedman, Jana Schaich Borg, Walter Sinnott-Armstrong, John Dickerson, Vincent Conitzer
Artificial Intelligence Journal (AIJ, vol. 283, 2020); AAAI Conference on Artificial Intelligence (AAAI 2018); AI, Ethics and Society (AIES 2018); Participatory ML Workshop (at ICML 2020); Mechanism Design for Social Good (MD4SG 2018)
Outstanding Student Paper Honorable Mention at AAAI 2018