[PDG 492] Neural audio codecs: how to get audio into LLMs
Details
Link to article: https://kyutai.org/codec-explainer
Title: Neural audio codecs: how to get audio into LLMs
Track: Speech-Native Conversational AI
Content: Raw audio is too dense for LLMs to model directly, so it must first be compressed into useful audio tokens. Kyutai explains how neural codecs, especially residual vector quantization, turn speech into discrete token streams that an LLM can predict. These codecs make speech-native LLMs practical, but there is still a tradeoff between audio quality, compression, and reasoning ability.
Slack link: ml-ka.slack.com, channel: #pdg. Please join us -- if you cannot join, please message us here or to mlpaperdiscussiongroupka@gmail.com.
In the Paper Discussion Group (PDG) we discuss recent and fundamental papers in the area of machine learning on a weekly basis. If you are interested, please read the paper beforehand and join us for the discussion. If you have not fully understood the paper, you can still participate – everyone is welcome! You can join the discussion or simply listen in. The discussion is in German or English depending on the participants.
