But this is just one feature of FFmpeg. Usually the heaviest CPU user is encode and decode, which is not affected by this improvement.
It's interesting and good work, but the "94x" statement is misleading.
But this is just one feature of FFmpeg. Usually the heaviest CPU user is encode and decode, which is not affected by this improvement.
It's interesting and good work, but the "94x" statement is misleading.
https://news.ycombinator.com/item?id=42042706
Talk about stacking the deck to make a point. Finely tuned assembly may well beat properly optimized C by a hair, but there's no way you're getting a two orders of magnitude difference unless your C implementation is extremely far from properly optimized.
GCC for AVR is absolutely abysmal. It has essentially no optimizations and almost always emits assembly that is tens of times slower than handwritten assembly.
For just a taste of the insanity, how would you walk through a byte array in assembly? You'd load a pointer to a register, load the value at that pointer, then increment the pointer. AVR devices can load and post-increment as a single instruction. This is not even remotely what GCC does. GCC will load your pointer into a register, then for each iteration it adds the index to the pointer, loads the value with the most expensive instruction possible, then subtracts the index from the pointer.
In assembly, the correct AVR method takes two cycles per iteration. The GCC method takes seven or eight.
For every iteration in every loop. If you use an int instead of a byte for your index, you've added two to four more cycles to each loop. (For 8 bit architectures obviously)
I've just spent the last three weeks carefully optimizing assembly for a ~40x overall improvement. I have a *lot* to say about GCC right now.
https://x.com/FFmpeg/status/1852913590258618852 https://x.com/FFmpeg/status/1850475265455251704
I'll bet money, sight unseen, that poster above is right its used for HEVC. I'll bet even more money its not some massive out of nowhere win, hand-writing assembly for popular codecs was de rigeur for ffmpeg. Thrust of the article, or at least the headline, is clickbait-y.
But the rest of it is on the CPU because GPU cores aren't any good at largely serial things like video decoding. So it doesn't matter.
I'm surprised to hear this, considering GPUs are often used for video encoding and decoding. For Nvidia cards, this is called NVDEC, and AMD/Intel both have corresponding features for their video cards.
Any kind of decompression isn't fully parallelizable. If you've found any opportunities, that means the compression wasn't as efficient as it theoretically could be. Most codecs are merciful and eg restart the entropy coder across frames, which is why the multithreaded decoding in ffmpeg is able to work.
(But it comes with a lossless video codec called ffv1 that doesn't allow this.)
There's no reason you can't bundle the same kind of ASIC on a CPU too, and indeed Intel does do that with QuickSync. For video game capture/screen recording though (which is a big part of what people tend to do with NVEnc) it might be a bit more convenient for the chip to be on the GPU? I don't know, not a GPU expert.
Yeah there's usually a fast path which copies the framebuffer directly to the encoder internally, so the huge uncompressed frames never have to be transferred over the PCI Express bus.
Intel has QuickSync on their CPU cores and that has reigned supreme for many years. That's not a GPU block.