[PDG 446] LLMs Get Lost In Multi-Turn Conversation
Details
Link to article: https://arxiv.org/pdf/2505.06120
Title: LLMs Get Lost In Multi-Turn Conversation
Content: Large Language Models (LLMs) are often evaluated on single, fully-specified instructions, despite being designed as conversational interfaces that handle underspecified user needs. This study uses large-scale simulations to compare LLM performance in single-turn versus multi-turn conversational settings. The results reveal that all tested LLMs perform significantly worse in multi-turn conversations, with an average performance drop of 39%. This decline is primarily due to the models' tendency to make premature assumptions and their inability to recover from early mistakes in the conversation.
Slack link: ml-ka.slack.com, channel: #pdg. Please join us -- if you cannot join, please message us here or to mlpaperdiscussiongroupka@gmail.com.
In the Paper Discussion Group (PDG) we discuss recent and fundamental papers in the area of machine learning on a weekly basis. If you are interested, please read the paper beforehand and join us for the discussion. If you have not fully understood the paper, you can still participate – everyone is welcome! You can join the discussion or simply listen in. The discussion is in German or English depending on the participants.
