Intel Gen12/Xe Graphics Have AV1 Accelerated Decode – Linux Support Lands
phoronix.com
phoronix.com
It gives the false assumption that AV1 is being Hardware Accelerated in Intel Xe GPU which often means dedicated decoding block to decode the video codec in lowest energy usage possible.
https://github.com/intel/media-driver/blob/master/README.md#...
"AV1 hardware decoding is supported from Gen12+ platforms."
GPGPU implementation of AV1 would be a useful thing though. Anyone know of implementations?
What they don't have a SKU for quite yet is a 8C16T U-class (15W) part like the AMD Ryzen 7 4800U. https://www.amd.com/en/products/apu/amd-ryzen-7-4800u
https://www.cpubenchmark.net/compare/Intel-i7-10610U-vs-AMD-...
If your workload is running a single-app that may be underused but particularly in home-office mode I'm constantly doing VCs, having 4 or 5 browser tabs that are heavy, Office for documents, etc. It doesn't matter if all those are a single thread, I'd be filling up the 8 cores and probably taking advantage of the 16 threads from SMT as well.
Those are the typical units that externs get from customers' IT, when it isn't some kind of cloud based VM.
And regular consumers don't even know what they own, rather what the guy at the store or some relative has given as advice.
https://www.cpubenchmark.net/compare/Intel-i7-10610U-vs-AMD-...
AMD wins on price as well. That people don't usually make good buying choices is not an argument about CPU performance.
I thought we have learned by now that it isn't the best tech that wins.
As such most developers only bother to use what they already know and take very little effort for adding any form of parallelism or concurrency to their applications.
Android and UWP since the start have taken architecture decisions that forbid synchronous code, because both companies came to the conclusion that if that would be available, the developers would write single threaded code as they have been doing for years, so they took that option out of the platform.
I am yet to see anyone using one of those AMD CPUs that get so high praises on HN.
If you look at A100 compared to V100 for e.g. FP32 FMA performance (not tensor). 14.1TFLOPS -> 19.5 = +38%, for 2x transistors (16->7nm), +35% SMs and 250W->400W is not that great. Note that NVIDIA uses boost clock for all A100 numbers and seems not have published any base clock so far. So their is a chance that actual sustained A100 performance is lower.
Turing GPUs have rather large dies. TU102 754nm vs GP102 471nm. So comparing them as is, isn't quite fair.
On the CPU front Intel used to use rather small dies for consumers (and even use die shrinks to just cram more chips onto a waver -> more $$$), but now that AMD forces their hand, they are giving in. But of course a lot this area goes into extra cores, not single threaded performance (diminishing returns there).
I feel like one would expect this, given that the rendering engines of any given generation, are architected under specific assumptions about the optimal relative "shape" of their graphics pipeline.
The clearest example being old game consoles. You could write a SNES emulator for the SuperFX coprocessor, that used the host's GPU to render the polygons, but the rendering would not go any faster than it does on the SNES, because the draw commands are just being trickled out as the physics-engine work necessary to update their positions gets done in fits and starts per scan-line. That trickling-out was necessary on a SNES, because there wasn't time to run all that logic during VBLANK; but in the modern era, we have the opposite problem — that the logic could all be completed during VBLANK (with tons of room to spare), but instead is being "dragged out" across 512 HBLANKs, such that the GPU only gets the full picture of the completed scene at the last moment. (Despite the recent source-code leak, rewriting StarFox to make it render at 60FPS or more will not be a trivial process.)
The same thing is true, to a varying extent, of all legacy renderers. They're written for graphical pipelines that just don't match the one we have now. The one we have now is "wider" — more parallel — in so many places, but if it's just being used to recapitulate a long, serial set of fixed-function legacy pipeline stages, that width doesn't help it any.
AFAIK, with each generation, Nvidia has increased raw performance by 20-30%, which is significant.
On the other hand, the wall of physical limits is getting closer (which may restrict the chip size), but until then, GPUs are faring very well.
GPUs functionalities also have a very different nature (as you point out), but this can play well for the user. Realism depends on many functionalities, which have plenty of headroom for improvement (in the sense of hardware support); see ray tracing, which supposedly, is going to be significantly faster on Ampere.