Since ollama has diverged from llama.cpp, it will take a bit of time for ollama to support multi-modality. If you're using plain llama.cpp it looks like a PR has already merged for this model with vision and audio support:
They've actually gone back to (a lightly patched) llama.cpp with the 0.30 release a few weeks ago, and have now vendored-in an up to date release. Needless to say this is great news for both projects!
Just use llama.cpp or Unsloth Studio which wraps it, I don't know why anyone use Ollama anymore.
I switched from llama.cpp to vLLM because of prompt cache bugs in qwen/gemma models
This is a good starting issue with a bunch of linked/related
Highly recommend just dropping Ollama. You can download binary releases of llama.cpp for every platform and run them trivially in 5 seconds. Ollama serves no purpose other than to take open source work and rebadge as its own, while providing inferior functionality