Whisper: Nvidia RTX 4090 vs. M1 Pro with MLX
owehrens.com
owehrens.com
I have a 3090 and an M1 Max 32GB and and although I haven't tried Whisper the inference difference on Llama and Stable Diffusion between the two is staggering, especially with Stable Diffusion where SDXL is about 0:09 seconds 3090 and 1:10 minute on M1 Max.
Several people working on mlx-enabled backends to popular ML workloads but it seems inference workloads are the most accelerated vs generative/training.
We may just be thankful that this particular bit of marketing never caught on for CPUs.
Apple M1 Max has 32 GPU cores, each core contains 16 Execution Units, each EU has 8 ALUs (also called shaders), so overall there are 4096 shaders. Nvidia RTX 4090 contains 12 Graphics Processing Clusters, each GPC has 12 Streaming Multi-Processors, and each SM has 128 ALUs, overall there are 18432 shaders.
A single shader is somewhat similar to a single lane of a vector ALU in a CPU. One can say that a single-core CPU with AVX-512 has 8 shaders, because it can process 8 FP64s at the same time. Calling them "cores" (as in "CUDA core") is extremely misleading, so "shader" became the common name for a GPU's ALU due to that. If Nvidia is in charge of marketing a 4-core x86-64 CPU, they would call it a CPU with 32 "AVX cores" because each core has 8-way SIMD.
I think it's got the same total FP ALU resources as zen3, and shows how register width and ALU resources can be completely decoupled.
70B in particular is indeed a significant compromise on the 4090, but not as much as you'd think. 34B and down though, I think Nvidia is unquestionably king.
I'm no expert, but to me that sounds like a recipe for bad performance. Does a 70B model in 2-bit really outperform a smaller-but-less-quantised model?
I woukd say 34B is the performance sweetspot, yeah. There was a long period where allow we had in the 33B range was llamav1, but now we have Yi and Codellamav2 (among others).
Whisper is no exception: https://github.com/Vaibhavs10/insanely-fast-whisper
SDXL is actually an interesting exception for Nvidia because most users still tend to run it in PyTorch eager mode. There are super optimized Nvidia implementations, like stable-fast but their use is less common. Apple, on the other hand, took the odd step of hand writing a Metal implementation themselves, at least for SD 1.5.
What I like to say is (generally speaking) other implementations like AMD (ROCm), Intel, Apple, etc are more-or-less at the “get it to work” stage. Due to their early lead and absolute market dominance Nvidia has been at the “wring every last penny of performance out of this” stage for years.
Efforts like this are a good step but they still have a very long way to go to compete with multiple layers (throughout the stack) of insanely optimized Nvidia/CUDA implementations. Bonus points nearly anything with Nvidia is a docker command that just works on any chip they’ve made in the last half decade from laptop to datacenter.
This can be seen (dramatically) with ROCm. I recently took the significant effort (again) to get an LLM to run on an AMD GPU. The AMD GPU is “cheaper” in initial cost but when the dollar equivalent (to within 10-30%) Nvidia GPU is 5-10x faster (or whatever) you’re not saving anything.
You’re already at a loss unless your time is free just to get it to work (random patches, version hacks, etc) and then the performance just isn’t even close so the “value prop” of AMD currently doesn’t make any sense whatsoever. The advantage for Apple is you likely spent whatever for the machine anyway, and when you have it just sitting in front of you for a variety of tasks the value prop increases significantly.
Also, as you hint whisper.cpp certainly isn't one of the fastest implementations of whisper inference out there. Perhaps a comparison between a pure PyTorch version running on the 4090 with an MLX version of Whisper running on the M1 Pro would be fairer. Or better yet, run the whisper encoder on ANE with CoreML and have the decoder running with Metal and Accelerate (which uses Apple's undocumented AMX ISA) using MLX, since MLX currently does not use the ANE. IIRC, whisper.cpp has a similar optimization on Apple hardware, where it optionally runs the encoder using CoreML and the decoder using Metal.
I haven't tried the hardware/software/framework/... of the article, but I have an opinion on this exact topic.
The provided context is n earlier version of hardware where known implementations perform drastically differently, an order of magnitude differently.
That leaves the question why that specific tool exhibits the behavior described in the article.
Not sure if this is a hardware limitation or just unoptimized MLX libraries but I find it hard to believe they would have just ignored this very prominent use case. It's more likely that convolutions use high precision and much larger tile sets that require some expensive context switching when the entire transform can't fit in the gpu.
So I'd be very careful about your intuition on whisper performance unless it's literally the same software and same model (and then the comparison isn't very meaningful still, seeing how we want to optimize it for different platforms).
We use whisper in production and this is our findings: We use faster whisper because we find the quality is better when you include the previous segment text. Just for comparison, we find that faster whisper is generally 4-5x faster than OpenAI/whisper, and insanely-fast-whisper can be another 3-4x faster than faster whisper.
Edit: I see it doesn't yet support CPU inference, should be interesting once it's added.
Ideally it also exposes that parameter to the user.
Speed comparisons seem moot when quality is sacrificed for me, I'm working with very poor audio quality so transcription quality matters.
Insanely fast whisper (god I hate the name) is really a CLI around Transformers’ whisper pipeline, so you can just use that and use any of the settings Transformers exposes, which includes beam size.
We also deal with very poor audio, which is one of the reasons we went with faster whisper. However, we have identified failure modes in faster whisper that are only present because of the conditioning on the previous segment, so everything is really a trade off.
Just call the pipeline with:
result = pipe(sample, generate_kwargs={"num_beams": 5})
Side note, the insanely fast whisper readme gives benchmarks on an A100 but only the FA2 lines were. The rest were on a T4 looking at the notebooks/history. Turing doesn't support FA2 so the gap should be smaller with it, but based on the distil-whisper paper CTranslate2 is probably still faster.
TensorRT-LLM might be faster but I haven't looked into it yet.
It's enabled by default with the latest Transformers version, so just make sure you have:
* torch>=2.1.1
* transformers>=4.36.0
I just reran the notebook with 4.36.1 (minus the to_bettertransformer line) but it was slower (the batch size 24 section took 8 vs 5 min). Is there something I need to change? Going back to 4.35.2 gives the old numbers so the T4 instance seems fine.
/s
Edit: OK I took the bait. I downloaded the 10 minute file he used and ran it on my 4090 with insanely-fast-whisper, which took two commands to install. Using whisper-large-v3 the file is transcribed in less than eight seconds. Fifteen seconds if you include the model loading time before transcription starts (obviously this extra time does not depend on the length of the audio file).
That makes the 4090 somewhere between 6 and 12 times faster than Apple's best. It's also much cheaper than M2 Ultra if you already have a gaming PC to put it in, and still cheaper even if you buy a whole prebuilt PC with it.
This should not be surprising to people, but I see a lot of wishful thinking here from people who own high end Macs and want to believe they are good at everything. Yes, Apple's M-series chips are very impressive and the large RAM is great, but they are not competitive with Nvidia at the high end for ML.
ETA: actually it's unclear from the article if the whisper optimizations were done by apple engineers, but it's definitely an optimized version.
There is also cpp Whisper (https://github.com/ggerganov/whisper.cpp) which seems to have it’s own kind of optimizations for Apple Silicon - I don’t think this was the one used with Nvidia during the test.
I installed following the official docs and found it much, much slower, although I sadly don't have a 4090, instead a 3080 Ti 12GB (just big enough to load the large whisper model into GPU memory).
pipx install insanely-fast-whisper
pipx runpip insanely-fast-whisper install flash-attn --no-build-isolation
To transcribe the file: insanely-fast-whisper --flash True --file-name ~/Downloads/podcast_1652_was_jetzt_episode_1289963_update_warum_streiken_sie_schon_wieder_herr_zugchef.mp3 --language german --model-name openai/whisper-large-v3
The file can be downloaded at: https://adswizz.podigee-cdn.net/version/1702050198/media/pod...I just ran it again and happened to get an even better time, under 7 seconds without loading and 13.08 seconds including loading. In case anyone is curious about the use of Flash Attention, I tried without it and transcription took under 10 seconds, 15.3 including loading.
Another question that's only slightly related, but while we're here...
Using OAI's paid Whisper API, you can give a text prompt to a) set the tone/style of the transcription and b) teach it technical terms, names etc that it might not be familiar with and should expect in the audio to transcribe.
Am I correct that this isn't possible with any released versions of Whisper, or is there a way to do it on my machine that I'm not aware of?
I was rather hoping for a guide of just how to either adapt classic whisper usage or adapt one of the optimised ones like faster-whisper (which I've just set up in a docker container but that's used up all the time I've got for playing around right now) to take a text prompt with the audio file.
The 4090 is an absolute beast, runs extremely quiet and simply powers through everything. DCS pushes it to the limit, but the resulting experience is simply stunning. Mine's coupled to a 7800x3d which uses hardly any power at all, absolutely love it.
For ex ctranslate optimized whisper “implements a custom runtime that applies many performance optimization techniques such as weights quantization, layers fusion, batch reordering…”
Intuitively, I would agree with your conclusion about Apple’s M-Series being impressive for what they do but not generally competitive with Nvidia in ML.
Objectively however, I don’t see concluding much with what’s on offer here. Once you start changing libraries, kernels, transformer code, etc you end up with an apples to oranges comparison.
That said, for practical purposes, the ready availability of Nvidia-optimized versions of every ML system is a big advantage in itself.
In any event, it’s super cool to see such huge leaps just in the past year on how easy it is to run this stuff locally. Certainly looking very promising.
Also, the price difference between a prebuilt 4090 PC and a M2 Ultra Mac Studio can buy a lot of kilowatt hours.
Wouldn't the high end for Nvidia be their dedicated gear rather than a 4090?
Or having a better image of performance in your head when you buy new hardware.
It's just a blog article. The time and effort to make it and to consume it is not in the range of millions there is no need for 'more'.
Personally, I have already a RTX 3090, so for me it’s interesting if a M3 would be a noticeable upgrade (considering RAM, speed, library support)
The super optimized code for a network like this is one test case. General networks are another.
I mean, you're probably right, but MLX framework was released just a week ago so maybe we don't really know what it's capable of yet.
The power:perf ratios seem about equal although, so it's really up to apple to release equivalently sized silicon on their desktop to really have this be a 1:1 comparison.
Could anyone who understands this hardware better than me chime in on the complexities of bringing unified ram/vram to PC, or reasons it hasn't happened yet?
I think that not using optimizations allows this to be a 1:1 comparison, but if the optimizations are not ported to MLX, then it would still be better to use a 4090.
Having looked at MLX recently, I think it's definitely going to get traction on Macs - and iOS when Swift bindings are released https://github.com/ml-explore/mlx/issues/15 (although there might be some C++20 compilation issue blocking right now).
There are some rare exceptions (like GPT-Fast on AMD thanks to PyTorch's hard work on torch.compile, and only in a narrow use case), but I can't think of a single one for Apple Silicon.
To me the news here is how well the Mac runs without needing that additional hardware/large power draw on this benchmark.
The post here is exactly one for Apple Silicon. It compared a naive implementation in PyTorch which may not even keep 4090 busy (for smaller/not-that-compute-intensive models having the entire computation driven by Python is... limiting, which is partly why torch.compile gives amazing improvements) to a purposedly-optimized one (optimized for both CPU/GPU efficiency) for Apple Silicon one.
I think Apple's API would be as popular as CUDA if you could rent their chips at scale. They're quite efficient machines that don't need a lot of cooling, so I imagine the OPEX of keeping them running 24/7 in big cloud racks would be pretty low if they were optimised for server usage.
Apple seems to focus their efforts on bringing purpose-built LLMs to Apple machines. I can see why it makes sense (just like Google's attempts to bring Tensor cores to mobile) but there's not much practical use in this technology right now. Whisper is the first usable technology like this, but even my Android phone can live translate spoken text into words as an accessibility feature, I don't think Apple can sell Whisper as a product to end users.
I don't think so, in the sense of a hand-optimized CUDA implementation. This just using the MLX API in the same way that you'd use CUDA via PyTorch or something.
Otherwise you need a whole bunch of custom mac mini style racks and management software which really increases costs and lead times. If you don't believe me, look how expensive AWS macOS machines are compared to linux ones with equivalent performance.
You can beat these benchmarks on a CPU; 3-4x realtime is very slow for whisper these days!
>At the time of writing this comparison convolutions are still some of the least optimized operations in MLX.
I think the main thing at play is the fact you can have 64+G of very fast ram directly coupled to the cpu/gpu and the benefits of that from a latency/co-accessibility point of view.
These numbers are certainly impressive when you look at the power packages of these systems.
Worth considering/noting that the cost of m3 max system with the minimum ram config is ~2x the price of a 4090...
I spend an hour or two, trying to run figure out what I need to install / configure to enable it to use MLX. Was getting cryptic Python errors, Torch errors... Gave up on it.
I rented VM with GPU, and started Whisper on it within few minutes.
Repos like https://github.com/SYSTRAN/faster-whisper makes immediate sense on why it's faster than the original implementation, and lots of others do so by lowering quantization precision etc (and worse results).
but this one, it's not very clear how. Especially considering it's even much faster.
It's a shame faster-whisper never landed batch mode, as I think that's preventing folks from trying ctranslate2 more easily.
I looked at https://github.com/thomasmol/cog-whisper-diarization and https://about.transcribee.net/ (from the people behind Audapolis) but neither work that well -- crashes, etc.
Thank you!
It shouldn’t be so hard since many apps have this. But what is the most reliable way right now?
M3 MAX GPU -> 10 TFLOPS
It is 8 times slower than 4090.
But yeah, you can claim that a bike has a faster acceleration than Ferrari, because it could reach the speed of 1km per hour faster...
They just shipped 1.0 of the Ryzen AI Software and SDK. Alleges ONNX, PyTorch, and Tensorflow support. https://www.anandtech.com/show/21178/amd-widens-availability...
Interestingly, the upcoming XDNA2 supposedly is going to boost generative performance a lot? "3x". I'd kind of assumed these sort of devices would mainly be helping with inference. (I don't really know what characterizes the different workloads, just a naive grasp.)
This is again not native PyTorch so there's still room to have better RTFX numbers.
I'm trying to decide between the two. I figure the M3 Max would crush the 4070?
Matching performance while using less power is impressive. Using less power while also being slower not so much, though.
So where is the - presumably much cheaper - 65 W 7950X?
His point is that M-CPUs being somewhat competitive but much more efficient is not as stunning since you're comparing a CPU tuned to be at the most efficient speed to a CPU tuned to be at the highest speed.
Similarly a 4090's power consumption drops dramatically if you underclock or undervolt even slightly, but what's the point? You're almost definitely buying a 4090 for its raw speed.
And because I don't want a software limit that may or may not work.
And your concerns about "a software limit that may or may not work" are completely at odds with how their power management works.
It’s a choice each user can make if they care more about efficiency. Tasks take longer to complete, but the total energy consumed for the completion of the task is dramatically less.
The 250 W space heater vacuum cleaner soundtrack mode should be opt in rather than opt out. Same for video cards.
AMD isn’t going to offer a huge discount for a 65W 7950X for the reasons discussed elsewhere: they don’t need to.
“I can afford it” is still wasteful. Most people wouldn’t even notice the speed difference in 65 W mode. Even if they can afford the space heater mode.
Edit: The article has been updated to compare against a faster implementation for the Nvidia GPU and no longer makes that claim.
The is another one that uses huggingface's implementation, but I haven't tried it since my spec doesn't support flash-att2 https://github.com/luweigen/whisper_streaming
If you compare whisper on a mac with Mac optimized build Vs on a pc with a few NON-optimized NVIDIA build The results are close! If nvidia optimized is compared, it’s not even remotely close.
Pfft
I’ll be picking up a Mac but I’m well aware it’s not close to Nvidia at all. It’s just the best portable setup I can find that I can run completely offline.
Do people really need to make these disingenuous comparisons to validate their purchase?
If a mac fits your overall use case better, get a Mac. If a pc with nvidia is the better choice, get it. Why all these articles of “look my choice wasn’t that dumb”??
Anyway the middle button is “refuse all” according to my phone, not sure how accurate the translation is or if they’ll shuffle the buttons for other people.
It is poor design to have what appear to be “accept” and “refuse” both in green.