Skip to content

Details

Link to article: https://arxiv.org/pdf/2505.06120
Title: LLMs Get Lost In Multi-Turn Conversation
Content: Large Language Models (LLMs) are often evaluated on single, fully-specified instructions, despite being designed as conversational interfaces that handle underspecified user needs. This study uses large-scale simulations to compare LLM performance in single-turn versus multi-turn conversational settings. The results reveal that all tested LLMs perform significantly worse in multi-turn conversations, with an average performance drop of 39%. This decline is primarily due to the models' tendency to make premature assumptions and their inability to recover from early mistakes in the conversation.
Slack link: ml-ka.slack.com, channel: #pdg. Please join us -- if you cannot join, please message us here or to mlpaperdiscussiongroupka@gmail.com.

In the Paper Discussion Group (PDG) we discuss recent and fundamental papers in the area of machine learning on a weekly basis. If you are interested, please read the paper beforehand and join us for the discussion. If you have not fully understood the paper, you can still participate – everyone is welcome! You can join the discussion or simply listen in. The discussion is in German or English depending on the participants.

Related topics

Artificial Intelligence
Deep Learning
Machine Learning
Natural Language Processing
Neural Networks

You may also like