Skip to content

Details

Google and other AI companies are hiring philosophers. Why?

In most of human history we have never contemplated the question of how to manage an entity that is much smarter than humanity. Time to start.

This short YouTube clip is a useful primer on some basics relating to the current state of AI before going on to explain the AI alignment problem and why we should worry about this. Along the way it also has some interesting things to say about AI intelligence being very different to our own but that's perhaps for another day.

But back to the hiring of philosophers by the big tech companies. You can't have missed the myriad news stories about how frontier AI models might pose an existential threat to humanity. As the YouTube clip explains, the AI alignment problem requires us to encode not just intelligence but wisdom. We need to deal with the challenges of our "... complex, messy, often contradictory human values". This may require the integration of philosophical thinking as philosophers can bring frameworks for value alignment and can help define ethical parameters for AI decision-making.

Note - I'm not suggesting we tackle the question of whether solving the AI alignment problem is necessary and sufficient to keep us safe from ASI (Artificial Super Intelligence). But it is likely to be important and an interesting topic to discuss as it does throw up some profound questions. Starting with...

  • Q1/ Are we worth saving and if yes (briefly…) what does 'we' mean: This question should help us ground the alignment debate in philosophy rather than pure computer science. The text and questions below dig deeper into the question of who or what is “we” and allows us to focus on the value-alignment problem - whose values, which aspects of humanity, and what kind of future we are trying to achieve or preserve.

What is 'We'?

  • Q2/ The Biological vs. The Structural: Are we aligning AI to protect human biological life, or are we trying to preserve human agency, culture, and consciousness? If an AI preserves human bodies in a state of perpetual, sedated bliss, has it saved "us"?
  • Q3/ The Temporal "We": Are we aligning AI with present human values (which may be deeply flawed) or idealised future human values? How do we account for the moral progress of future generations without allowing the AI to dictate what that progress should be?

Whose Values
If we are aligning an AI, whose morality, culture, or philosophical framework dictates the "correct" alignment features?

  • Q4/ Relativism vs. Universalism: Can we find a set of universal human values (e.g., via the UN Declaration of Human Rights), or does alignment inherently enforce the values of the dominant tech cultures?
  • Q5/ Political Philosophy (Rawlsian Alignment): Should an AI be aligned to a "Veil of Ignorance" standard, where it optimizes for the least advantaged in society?

Those "... complex, messy, often contradictory human values"

  • Q6/ Deontology vs. Consequentialism: Should an AI follow strict, unbreakable moral rules (like Asimov's laws or Kantian duties), or should it dynamically calculate the "greatest good for the greatest number" (utilitarianism)? What happens when a utilitarian AI decides a minor harm prevents a massive catastrophe?

Traditional ethics often assumes that suffering, trial, and error are necessary components of moral growth and human achievement. A hyper-aligned AI designed to maximize safety and well-being might eliminate the bad, but in doing so, strip away what makes human lives meaningful.

  • Q7/ Virtue Ethics: Many core human virtues (courage, empathy, sacrifice) rely on vulnerability, mortality, and imperfection. If an aligned AI eliminates human suffering and risk, does it destroy the very qualities that make human lives worth living? Is it possible or perhaps even essential to align an AI to embody human virtues.
  • Q8/ Perspectivism & Universalism: Is there a coherent, universal human essence to align with, or is "humanity" an aggregation of incompatible perspectives, power dynamics, and cultural traditions? Can an AI align with a species that is not aligned with itself?

What Are We Actually Trying to Do Here with Advanced AI?

  • Q9/ Preservation or Advancement: AI is increasingly an incredibly powerful tool. Are we looking just to preserve humanity (or however we have defined ‘we’) or is AI the means to enable human flourishing?

Wildcard Thoughts and Questions

  • Q10/ What else might help with regards to the AI alignment problem: Should AI get religion? Does AI need a body? Should it know pain?
  • Q11/ The Scope of Moral Patienthood: Does "we" strictly refer to Homo sapiens, or does it extend to future non-human sentient entities, including potential artificial moral agents? If an advanced AI becomes sentient, does "we" expand to include the AI itself?

Related topics

You may also like