HNHacker News
TopNewBestAskShowJobs

watsonmusic

8 karma · joined October 8, 2024

submissionscomments
watsonmusic··on VibeVoice: A Frontier Open-Source Text-to-Speech Model
it's not oss
watsonmusic··on VibeVoice: A Frontier Open-Source Text-to-Speech Model
bonus usage
watsonmusic··on VibeVoice: A Frontier Open-Source Text-to-Speech Model
11labs is facing a real competitor
watsonmusic··on VibeVoice: A Frontier Open-Source Text-to-Speech Model
genius
watsonmusic··on VibeVoice: A Frontier Open-Source Text-to-Speech Model
this model is superb
watsonmusic··on VibeVoice: A Frontier Open-Source Text-to-Speech Model
Microsoft is cool
watsonmusic··on VibeVoice: A Frontier Open-Source Text-to-Speech Model
yes the best
watsonmusic··on VibeVoice: A Frontier Open-Source Text-to-Speech Model
one of the best models built by Microsoft
watsonmusic··on Microsoft releases VibeVoice, generates 90-minute, 4-speaker audio
https://github.com/microsoft/VibeVoice
watsonmusic··on Microsoft releases VibeVoice, generates 90-minute, 4-speaker audio
https://huggingface.co/microsoft/VibeVoice-1.5B
watsonmusic··on Microsoft releases VibeVoice, generates 90-minute, 4-speaker audio
VibeVoice is a novel framework designed for generating expressive, long-form, multi-speaker conversational audio, such as podcasts, from text. It addresses significant challenges in traditional Text-to-Speech (TTS) systems, particularly in scalability, speaker consistency, and natural turn-taking. A core innovation of VibeVoice is its use of continuous speech tokenizers (Acoustic and Semantic) operating at an ultra-low frame rate of 7.5 Hz. These tokenizers efficiently preserve audio fidelity while significantly boosting computational efficiency for processing long sequences. VibeVoice employs a next-token diffusion framework, leveraging a Large Language Model (LLM) to understand textual context and dialogue flow, and a diffusion head to generate high-fidelity acoustic details. The model can synthesize speech up to 90 minutes long with up to 4 distinct speakers, surpassing the typical 1-2 speaker limits of many prior models.
watsonmusic··on Reinforcement Pre-Training
cannot wait seeing how it goes beyond the current llm training pipeline
watsonmusic··on Reinforcement Pre-Training
it could be adaptive. only high-value tokens were allocated with more compute
watsonmusic··on Reinforcement Pre-Training
A new scaling paradigm finally comes out!
watsonmusic··on Reinforcement Pre-Training
14b model performs comparably with 32b size. the improvement is huge
watsonmusic··on Differential Transformer
negative values can enhance the expressibility
watsonmusic··on Differential Transformer
not all hallucinations are creativity Imaginate that for a RAG application, the model is supposed to follow the given documents
watsonmusic··on Differential Transformer
the model is supposed to learn this
watsonmusic··on Differential Transformer
The modification is simple and beautiful. And the improvements are quite significant.
watsonmusic··on Differential Transformer
that would be huge!