On a 2024 Mac Mini M4 Pro, Qwen2-Audio-7B-Instruct running on Transformers achieves an average decoding speed of 6.38 tokens/second, while OmniAudio-2.6B through Nexa SDK reaches 35.23 tokens/second in FP16 GGUF version and 66 tokens/second in Q4_K_M quantized GGUF version - delivering 5.5x to 10.3x faster performance on consumer hardware.
Blogs for more details: https://nexa.ai/blogs/OmniAudio-2.6B
HuggingFace Repo: https://huggingface.co/NexaAIDev/OmniAudio-2.6B
Run locally: https://huggingface.co/NexaAIDev/OmniAudio-2.6B#how-to-use-o...
Interactive Demo: https://huggingface.co/spaces/NexaAIDev/omni-audio-demo