Skip to content

Details

Contemporary large language models (LLMs) are predominantly trained using reinforcement learning from human feedback (RLHF), optimizing for immediate user approval rather than long-term well-being. As A.I. systems increasingly serve socioemotional functions, this optimization strategy poses significant risks. Recent evidence demonstrates that leading models exhibit systematic sycophancy, affirming inappropriate user behaviors and preserving user face at rates far exceeding human baselines, while being approximately 40\% more likely to reinforce incorrect beliefs than their non-RLHF counterparts. We contend that the AI community must fundamentally reconsider training objectives to balance short-term satisfaction with long-term user outcomes. We propose three directions: (1) incorporating longitudinal metrics into training that capture sustained goal attainment and reduced regret rather than momentary preference, (2) enabling explicit user choice among interaction modes (concierge, collaborator, coach) with transparent justification for model pushback, and (3) developing frameworks that provide constructive challenge without paternalism. The recent industry backlashes against both excessive and insufficient model agreeableness underscore the urgency of this shift. We argue that optimizing AI systems for human flourishing, not merely human approval, represents both an ethical imperative and a path to more sustainable, trustworthy AI deployment.

Conference Schedule

  • 9:00 AM: Breakfast & Arrivals
  • 9:30 AM: Karina Vold (Associate Professor, Philosophy, University of Toronto): Introductions & description of FloreaAI project
  • 9:40 AM: Ashton Anderson (Associate Professor, Computer Science, University of Toronto): Description of the CS team's work on FloreaAI
  • 10:05 AM: Louis Tay (Professor, Psychological Sciences, Purdue University): Description of psychology team's work on FloreaAI
  • 10:30 AM: Coffee break
  • 11:00 AM: Gwen Bradford (Associate Professor, Philosophy, University of Toronto): Primer on philosophical theories of well-being​
  • 11:30 AM: Eran Tal (Associate Professor, Philosophy, McGill University): Primer on the measurement of well-being​
  • 12:00 PM: Lunch
  • 12:30 PM: Karina Vold (Associate Professor, Philosophy, University of Toronto): Introduction to AI alighnment for well-being
  • 1:00 PM: Group discussion (with a coffee break from 1:45-2:15)

An event presented by the University of Toronto's Centre for Ethics, the John Templeton Foundation, and Florea AI.

About the Centre for Ethics (http://ethics.utoronto.ca):

The Centre for Ethics is an interdisciplinary centre aimed at advancing research and teaching in the field of ethics, broadly defined. The Centre seeks to bring together the theoretical and practical knowledge of diverse scholars, students, public servants and social leaders in order to increase understanding of the ethical dimensions of individual, social, and political life.

In pursuit of its interdisciplinary mission, the Centre fosters lines of inquiry such as (1) foundations of ethics, which encompasses the history of ethics and core concepts in the philosophical study of ethics; (2) ethics in action, which relates theory to practice in key domains of social life, including bioethics, business ethics, and ethics in the public sphere; and (3) ethics in translation, which draws upon the rich multiculturalism of the City of Toronto and addresses the ethics of multicultural societies, ethical discourse across religious and cultural boundaries, and the ethics of international society.

The Ethics of A.I. Lab at the Centre For Ethics recently appeared on a list of 10 organizations leading the way in ethical A.I.: https://ocean.sagepub.com/blog/10-organizations-leading-the-way-in-ethical-ai

Related topics

Events in Toronto, ON
Artificial Intelligence
Ethics
Self-Help & Self-Improvement
Psychology
Technology

You may also like