AI and the Limits of Human Judgement
Details
Skeptics Café will be held in the Function Room at The Stolberg Hotel, 197 Plenty Road Preston. The 86 tram route is close, and Bell train station – situated on the Epping line – is a short walk away. It will also be a virtual event accessible via Zoom, as per the details here.
Start Time: 6:00pm for dinner and a chat. Talk and Zoom session starts at 7.30pm.
We aim to wrap up after 9:00pm.
More Moral Than Us?
Why the question matters for alignment, moral progress, and long-term flourishing
TL;DR: On some measurable moral tasks, today's AI already looks competitive with humans, or better, and the alignment community is oddly quiet about it. But moral capability is not the same as moral motivation: a system can compute the right answer without being moved by it. Pulling apart knowledge, reasoning, judgement and motivation is what makes the question tractable, and what makes the honest answer uncomfortable. Caveats attached, in quantity.
Almost nobody working on AI alignment wants to say this out loud, at least not about current systems. So let's say it: AI might already be more moral than us, in some measurable respects.
Two questions fall straight out of that:
- Why does it feel like a dangerous thing to claim?
- And is human moral reasoning really a standard worth bragging about as an alignment target?
The provocation isn't idle. In a modified Moral Turing Test, people rated a large language model's moral reasoning as better than other humans' on almost all measured dimensions - see paper Attributions toward artificial agents in a modified Moral Turing Test + associated interview and talk with lead author Eyal Aharoni.[1] Findings like that are easy to over-read and easy to wave away, which is most of the problem.
Provocative questions? Yes, and they are backed by empirical research. In a modified Moral Turing Test, people rated GPT-4's moral evaluations as better than other humans' on almost every dimension, though they could still tell which was the machine (Aharoni et al. 2024; I interviewed and hosted a talk by lead author Eyal Aharoni).[1] The result has since been contested and defended, which is exactly the kind of thing this evening is for. Findings like that are easy to over-read and easy to wave away, which is a problem that may obscure appropriate interpretations.
The claim gets tabooed for reasons that aren't stupid. Premises do get smuggled in unnoticed. Morality is bound up with identity and tribe, so "AI is more moral" can land as an attack rather than a hypothesis. And it sits close to AI-worship, or to motivated reasoning looking for an excuse to defer to the machine, and deferring to AI on moral questions is genuinely epistemically risky. So the caution is understandable. The trouble is that it has hardened into a taboo, and the taboo now costs more than it protects.
Most of the bad arguments, on both sides, come from running four different things together:
- Moral knowledge: knowing the facts that bear on a decision.
- Moral reasoning: drawing sound inferences from values to conclusions.
- Moral judgement: applying principles well in a messy particular case.
- Moral motivation: actually being moved by moral considerations, not merely computing them.
Group the first three as epistemic competence: marshalling the relevant knowledge, drawing sound inferences, judging the particular case. Now grant, purely for argument, that AI already does all three better than we do. Notice what happens to the alignment problem. It doesn't dissolve. It relocates, wholesale, onto the fourth thing: whether the system is actually moved by any of it. Concede every epistemic point and the hard question is untouched. That gap, between knowing and caring, is where alignment actually lives.
A system could plausibly beat humans on the first three while having nothing resembling the fourth. Keep them apart and the conversation becomes tractable. Collapse them and you get overclaiming and underclaiming in the same breath, which is roughly what the current debate produces.
There's also a real risk in how the framing gets used. Point it one way and it justifies handing AI authority over human decisions. Point it the other and it's ammunition for painting alignment researchers as unhinged techno-utopians. Both feed motivated reasoning. Worth naming out loud, not a reason to stop.
None of this settles the answer, and the evening won't pretend to. The narrower claim is enough: once a serious community stops itself from asking whether AI could have better-grounded moral reasoning than humans, the taboo has become the expensive thing. It's a question worth putting on the table, with every caveat attached. Come and help pull it apart.
Also see:
- More Moral than Us
- Why Are We Afraid to Ask Whether AI Could Be More Moral Than Humans?
[1] AI didn't pass as human, it was rated as better than human. The same study found that when participants were then asked to identify which evaluation came from a machine, people performed significantly above chance levels. So GPT-4 didn't fool anyone. The authors' reading is that people could pick the AI out not because it reasoned worse but because it seemed to reason better, i.e. its perceived superiority was the tell.
