2,960 karma · joined January 28, 2019
You would need:
* A STT (ASR) model that outputs phonetics not just words
* An LLM fine-tuned to understand that and also output the proper tokens for prosody control, non-speech vocalizations, etc
* A TTS model that understands those tokens and properly generate the matching voice
At that point I would probably argue that you've created a native voice model even if it's still less nuanced than the proper voice to voice of something like 4o. The latency would likely be quite high though. I'm pretty sure I've seen a couple of open source projects that have done this type of setup but I've not tried testing them.
https://elevenlabs.io/docs/agents-platform/overview#architec...
However, It still doesn't seem capable of producing any of the sounds, like laughter, that I would expect from a native voice model.
The creator posted a little demo of it working with Qwen3 Omni that is quite impressive: https://www.youtube.com/watch?v=5DBFVe3cLto
He didn't include any details regarding how the model was running though
I think ChatGPT has the most lifelike speech with their voice models. They seem to have invested heavily in that area while other labs focused elsewhere.
Are there any open weight models that do? Not talking about speech to text -> LLM -> text to speech btw I mean a real voice <-> language model.
edit:
It does support real-time conversation! Has anybody here gotten that to work on local hardware? I'm particularly curious if anybody has run it with a non-nvidia setup.
The Apple offerings are interesting but the lack of x86, Linux, and general compatibility make it hard sell imo.
Presumably they wouldn't be training on synthetic data produced by anything less than a open frontier model and those are almost exclusively Chinese
Half the dataset being synthetic is interesting. I wonder what that actually means. They say that Datology needed 2048 H100s to generate the synthetic data. Does that mean they were generating data using other open weight LLMs? Seems like that would undermine the integrity of a "US based" dataset.
I was working freelance through late 2023 - mid 2025 and the shift seemed quite obvious to me. Other freelancers, agency managers, etc that I talked to could see it too. The volume of clients, and their expectations, is changing very rapidly in that space.
Now AI makes it unbelievably easy to make those simple but bespoke software packages. The business owner can boot up Lovable and get something that is good enough. The non-software folk generally aren't scrutinizing the software they use. It doesn't matter if the backend is spaghetti code or if there are bugs here and there. If it works well enough then they're happy.
In my opinion that's the unfortunate truth of AI software development. It's dirt cheap, fast, and good enough for most people. Computer's couldn't write software before and now they can. Obviously that is real devaluation, right?