Skip to content

Details

Link to article: https://arxiv.org/abs/2410.00037
Title: Moshi: a speech-text foundation model for real-time dialogue
Track: Speech-Native Conversational AI
Content: Moshi is a real-time speech-to-speech dialogue model that avoids the usual pipeline of speech recognition, text dialogue, and text-to-speech. It can handle overlapping speech, interruptions, emotion, and non-speech sounds by modeling both user and system speech in parallel streams. Its “Inner Monologue” method improves speech quality by predicting aligned text before audio tokens, enabling about 200 ms practical latency.
Slack link: ml-ka.slack.com, channel: #pdg. Please join us -- if you cannot join, please message us here or to mlpaperdiscussiongroupka@gmail.com.

In the Paper Discussion Group (PDG) we discuss recent and fundamental papers in the area of machine learning on a weekly basis. If you are interested, please read the paper beforehand and join us for the discussion. If you have not fully understood the paper, you can still participate – everyone is welcome! You can join the discussion or simply listen in. The discussion is in German or English depending on the participants.

Related topics

Artificial Intelligence
Deep Learning
Machine Learning
Natural Language Processing
Neural Networks

You may also like