* through a data leak at Gigabyte
* through a data leak at Gigabyte
I've been hoping AVX-512 goes away. Intel doesn't even seem to be embracing it very well, probably due to the huge increase in power consumption.
I look forward to RISC-V Vectors, which allow scaling the hardware without changing the instruction set.
Do the SIMD units have similar register duplications? If so, reducing the duplication count would do pretty much that (Reducing the area at the cost of lower performance).
Even if not, the large general purpose register duplication count likely means that the SIMD registers are actually a lower proportion of the register file on chip than the register width would imply, making it less of a saving.
But generally I agree, I'm much more interested in the extra instruction vocabulary and things like masking lanes than the increased register width. I also wonder at what point the cut off of when a use case would be worth pushing to a more dedicated accelerator (like a dsp or gpu).
Alder Lake more or less fixed this according to Anandtech, with AVX512 using less than the rated turbo power of the CPU and running at the same frequency as non AVX code: https://www.anandtech.com/show/17047/the-intel-12th-gen-core...
Unfortunately, AVX-512 isnt officially supported on Alder Lake, so it might be another generation before its widely available.
It's actually there in Alder Lake's performance cores and can be used on at least some motherboards if the efficiency cores are disabled.
A shame because the ADL AVX512 units are much more efficient.
Wouldn't an illegal instruction fault suffice to flag a thread running on a mismatched core and allow it to be moved to a suitable one? Or do they lose state when the fault occurs?
X86 doesn’t make that easy.
I wouldn't be surprised if the x86 would let us down here however. Wouldn't be the 8367th time.
Then there is the issue with new instructions that get added.
Yes, you need OS support, but then again, you need support for the ThreadDirector anyway.
OTOH, this could allow us to have very heterogeneous processors in a single system, each with an ISA tailored to a kind of task and the OS knowing which core can run which program.
Linux and Windows both have partitioned schedulers, available through per-process and per-thread affinity masks. So it doesn't seem like it should be all that hard to implement. Its "just" a matter of expressing the process's (thread's?) requirements.
Beyond that, AVX-512 is not POR. It's not validated, checked for IEEE accuracy, and YMMV on whether it works at what frequency with the correct outputs.
Source: AnandTech.... where I wrote about it :)
if you wrote it, it I wouldn't count it as a source /s
isn't that than a bit risky to put something into the bios which might disable intels advances? and especially if not even intel tested it correctly? I mean as a consumer what would happen if I enable it and something breaks?
It's also possible it's buggy.
A possible workaround would be to use the invalid instruction fault to look for a core that implements that instruction.
That, however, wouldn't work if they were completely different, as in a couple GreenArray cores and a couple x86 ones.