None of inference frameworks (vLLM/SGLang) supports the full model, let alone non-nvidia.
Check it out here: https://models.hathora.dev/model/qwen3-omni
In our deployment, we do not actually tune the model in any way, this is all just using the base instruct model provided on huggingface:
https://huggingface.co/Qwen/Qwen3-Omni-30B-A3B-Instruct
And with the potential concern around conversation turns, our platform is designed for one-off record -> response flows. But via the API, you can build your own conversation agent to use the model.
https://github.com/vllm-project/vllm-omni
I have not yet tested out if this does full speech to speech, but this seems like a promising workspace for omni-modal models.