Open Source

VibeVoice

Expressive, longform conversational speech generation framework.

Source: GitHub Pricing: Open Source
💻 View Code

About This Project

VibeVoice is a novel framework designed for generating expressive, long-form, multi-speaker conversational audio from text. It employs a next-token diffusion framework and an LLM to manage dialogue flow and acoustic details, capable of generating up to 90 minutes of speech.

Tags

asr Audio Generation Text-to-Speech

Reviews & Ratings

Share your experience

User Reviews (0)

No reviews yet. Be the first to share your experience!