Different day/location! -- Lie Detectors
Details
*** DATE CHANGED TO THURSDAY (this week only) ***
*** LOCATION CHANGED ***
Come to the 5th floor of 100 University Avenue, room 5C (WeWork). Call Mario on 416-786-5403 if you have trouble getting in.
We know that language models can lie. A model may "know" the truth (and would output it correctly in most circumstances) but in some cases still output a falsehood. Detecting this may be very useful, both for keeping models truthful and for learning more about deceptive behaviour patterns in general. (Deception is considered a major category of safety risk in AI).
A recent paper presents a new kind of "lie detector", which surprisingly works without needing access to the model's internal activations. We'll be presenting on this paper (together with some background on deception in AI and why it's important), and there'll be plenty of opportunity for discussion. Is this the right approach to lie detection? Will it be robust as models get smarter? Come with your questions.
You can read the paper in advance if you're curious but it's not required.
