HNHacker News
TopNewBestAskShowJobs

AndreSlavescu

6 karma · joined October 5, 2023

AI researcher

X: https://x.com/andre_slav03 github: https://github.com/AndreSlavescu

submissionscomments
AndreSlavescu··on Speech to Speech Qwen3-Omni visualization tool
Hey everyone!

We at Hathora have recently released our ultra-low latency deployment of Qwen/Qwen3-Omni-30B-A3B-Instruct, one of the leading open source speech-to-speech-capable models.

Platform release:

https://models.hathora.dev/model/qwen3-omni

With the release, it got us thinking, what if we built a visualization tool to see what actually happens when you record audio and get a audio-response back? With that, we introduce our visualization website, that gives a high level overview of the individual pieces present in making speech-to-speech inference possible in qwen3-omni.

Feel free to give it a try, give us any feedback, and give the platform a try!

AndreSlavescu··on Qwen3-Omni-Flash-2025-12-01:a next-generation native multimodal large model
At the moment, no unfortunately. However, to my recent knowledge of open source alternatives, the vLLM team published a separate repository for omni models now:

https://github.com/vllm-project/vllm-omni

I have not yet tested out if this does full speech to speech, but this seems like a promising workspace for omni-modal models.

AndreSlavescu··on Qwen3-Omni-Flash-2025-12-01:a next-generation native multimodal large model
From my understanding of the above problem, this would be something to do with the model weights. Have you tested this with the transformers inference baseline that is shown on huggingface?

In our deployment, we do not actually tune the model in any way, this is all just using the base instruct model provided on huggingface:

https://huggingface.co/Qwen/Qwen3-Omni-30B-A3B-Instruct

And with the potential concern around conversation turns, our platform is designed for one-off record -> response flows. But via the API, you can build your own conversation agent to use the model.

AndreSlavescu··on Qwen3-Omni-Flash-2025-12-01:a next-generation native multimodal large model
Yeah, that's something we currently support. Feel free to try the platform out! No cost to you for now, you just need a valid email to sign up on the platform.
AndreSlavescu··on Qwen3-Omni-Flash-2025-12-01:a next-generation native multimodal large model
We actually deployed working speech to speech inference that builds on top of vLLM as the backbone. The main thing was to support the "Talker" module, which is currently not supported on the qwen3-omni branch for vLLM.

Check it out here: https://models.hathora.dev/model/qwen3-omni