[PDG 493] Moshi: a speech-text foundation model for real-time dialogue
Details
Link to article: https://arxiv.org/abs/2410.00037
Title: Moshi: a speech-text foundation model for real-time dialogue
Track: Speech-Native Conversational AI
Content: Moshi is a real-time speech-to-speech dialogue model that avoids the usual pipeline of speech recognition, text dialogue, and text-to-speech. It can handle overlapping speech, interruptions, emotion, and non-speech sounds by modeling both user and system speech in parallel streams. Its “Inner Monologue” method improves speech quality by predicting aligned text before audio tokens, enabling about 200 ms practical latency.
Slack link: ml-ka.slack.com, channel: #pdg. Please join us -- if you cannot join, please message us here or to mlpaperdiscussiongroupka@gmail.com.
In the Paper Discussion Group (PDG) we discuss recent and fundamental papers in the area of machine learning on a weekly basis. If you are interested, please read the paper beforehand and join us for the discussion. If you have not fully understood the paper, you can still participate – everyone is welcome! You can join the discussion or simply listen in. The discussion is in German or English depending on the participants.
