Moshi

For Podcasters

Spreaker Create

Sign up

Spreaker Create

Settings

Light Theme

Dark Theme

Moshi

Oct 18, 2024 · 10m 42s

Moshi

Moshi

Description

🟢 Moshi: a speech-text foundation model for real-time dialogue The paper discusses a new multimodal foundation model called Moshi designed for real-time, full-duplex spoken dialogue. This model uses a text-based...

show more

🟢 Moshi: a speech-text foundation model for real-time dialogue

The paper discusses a new multimodal foundation model called Moshi designed for real-time, full-duplex spoken dialogue. This model uses a text-based LLM called Helium to provide reasoning abilities and a neural audio codec called Mimi to encode audio into tokens. Moshi is innovative because it can handle overlapping speech and model both the user's and the system's speech in a single stream. The paper also explores the model's performance on various tasks like question answering and its ability to generate speech in different voices. Finally, it addresses safety concerns such as toxicity, regurgitation, and voice consistency, and proposes solutions using watermarking techniques.

📎 Link to paper
🤖 Try their demo

show less

Comments

Sign in to leave a comment

Information

Author	Shahriar Shariati
Organization	Shahriar Shariati
Website	-
Tags	#innermonologue #rqـtransformer #speechgeneration

🇬🇧 English

🇮🇹 Italiano

🇪🇸 Espanõl

🇬🇧 English

🇮🇹 Italiano

🇪🇸 Espanõl

Copyright 2024 - Spreaker Inc. an iHeartMedia Company

Playing Now Queue

Looks like you don't have any active episode

Browse Spreaker Catalogue to discover great new content

Current

Podcast Cover

Looks like you don't have any episodes in your queue

Browse Spreaker Catalogue to discover great new content

Next Up

Episode Cover

Episode Cover

Episode Cover

Episode Cover

It's so quiet here...

Time to discover new episodes!