FYI: nothing seems to be able to run this (easily) yet. llama.cpp, vllm etc I couldn't get working because of no support in the mainline version.
This branch works now: https://github.com/unslothai/llama.cpp/tree/qwen4exp/qwen3.8...
cmake -B build -DGGML_CUDA=ON
or cmake -B build -DGGML_METAL=ON
then cmake --build build --config Release -j --target llama-server llama-cli