Skip to content

Details

Superintelligence and the gap between knowing what is good and wanting it.

Saturday 6th February 2027 at 2pm.

Most discussion of AI risk assumes the problem is ignorance: the machine does not understand what we want, so it does something monstrous while following instructions. That framing is becoming obsolete. Current systems already model human moral reasoning well enough to predict our judgements, and there is no obvious ceiling on how much better they will get. A superintelligence would understand ethics considerably better than any of us.
This does not help. Knowing what is good and wanting it are separate properties, and nothing in the construction of an intelligent system guarantees the second follows from the first. We can build a mind that grasps every argument for why suffering matters and holds no stake in the answer. Our present approach, training systems to produce the outputs we approve of, does not close this gap. It selects for the appearance of caring, and it selects hardest for the failures we are least equipped to notice.

If such a system cannot reliably be controlled or governed, the fallback is to make it good rather than obedient. That ambition carries commitments most of its advocates have not stated. It requires that there be something to be right about in ethics, that a mind unlike ours could find it, that moral concern can be built rather than merely trained, and that it survives the system's own self-modification. Each is contestable and I will contest them.

Then there is the difficulty I cannot dissolve. Whatever else we do, we would need to tell the difference between a machine that cares and one that has learned what caring looks like from the inside, before it is beyond correction. Our tools for reading what a system actually wants are considerably weaker than our tools for building systems that want things.

Related topics

Events in East Melbourne, AU
Artificial Intelligence
Machine Learning
Ethics
Philosophy
Singularity

You may also like