AMD Gives Details on EPYC Zen4: Genoa and Bergamo, Up to 96 and 128 Cores
anandtech.com
anandtech.com
* through a data leak at Gigabyte
It's actually there in Alder Lake's performance cores and can be used on at least some motherboards if the efficiency cores are disabled.
A shame because the ADL AVX512 units are much more efficient.
Wouldn't an illegal instruction fault suffice to flag a thread running on a mismatched core and allow it to be moved to a suitable one? Or do they lose state when the fault occurs?
X86 doesn’t make that easy.
I wouldn't be surprised if the x86 would let us down here however. Wouldn't be the 8367th time.
Then there is the issue with new instructions that get added.
Yes, you need OS support, but then again, you need support for the ThreadDirector anyway.
OTOH, this could allow us to have very heterogeneous processors in a single system, each with an ISA tailored to a kind of task and the OS knowing which core can run which program.
Linux and Windows both have partitioned schedulers, available through per-process and per-thread affinity masks. So it doesn't seem like it should be all that hard to implement. Its "just" a matter of expressing the process's (thread's?) requirements.
Beyond that, AVX-512 is not POR. It's not validated, checked for IEEE accuracy, and YMMV on whether it works at what frequency with the correct outputs.
Source: AnandTech.... where I wrote about it :)
if you wrote it, it I wouldn't count it as a source /s
isn't that than a bit risky to put something into the bios which might disable intels advances? and especially if not even intel tested it correctly? I mean as a consumer what would happen if I enable it and something breaks?
It's also possible it's buggy.
A possible workaround would be to use the invalid instruction fault to look for a core that implements that instruction.
That, however, wouldn't work if they were completely different, as in a couple GreenArray cores and a couple x86 ones.
I've been hoping AVX-512 goes away. Intel doesn't even seem to be embracing it very well, probably due to the huge increase in power consumption.
I look forward to RISC-V Vectors, which allow scaling the hardware without changing the instruction set.
Do the SIMD units have similar register duplications? If so, reducing the duplication count would do pretty much that (Reducing the area at the cost of lower performance).
Even if not, the large general purpose register duplication count likely means that the SIMD registers are actually a lower proportion of the register file on chip than the register width would imply, making it less of a saving.
But generally I agree, I'm much more interested in the extra instruction vocabulary and things like masking lanes than the increased register width. I also wonder at what point the cut off of when a use case would be worth pushing to a more dedicated accelerator (like a dsp or gpu).
Alder Lake more or less fixed this according to Anandtech, with AVX512 using less than the rated turbo power of the CPU and running at the same frequency as non AVX code: https://www.anandtech.com/show/17047/the-intel-12th-gen-core...
Unfortunately, AVX-512 isnt officially supported on Alder Lake, so it might be another generation before its widely available.
Let's do it so that the executives don't have to.
"Blood?"
"Too gross, and not enough variations."
"Rubies?"
"People would find a way to confuse the CPUs with the programming language, and besides, if Twitter digs up those tech demos of ours we're totally gonna get cancelled."
"Romans?"
"Sure, Romans are cool, and they conquer things, like we're going to conquer Intel. I like it."
"How about emperors?"
"Fine in principle, but they each did at least a few scandalous things... besides, we probably don't want the "history class" vibe, or the energy of the weird uncle who always wants to talk about Roman battle tactics at dinner."
"Yeah, fair enough. Cities, then? I visited Milan the other year..."
(30 minutes later)
"Cities it is!"
https://www.champagne.fr/en/comite-champagne/bureaus/bureaus...
Come on, Bergamo? For real? I guess it could sound "cool" to non-Italian speakers, but here in Italy I think that Bergamo is one of the most forgettable cities ever. Which is sad, because its historical city centre is very beautiful, but the rest of the city is kind of dull and boring.
Of course, once this hits the market it won't actually be sold as the EPYC Bergamo, but rather the EPYC 9873 or whatever
https://www.folklore.org/StoryView.py?story=World_Class_Citi...
For example: https://en.wikipedia.org/wiki/Alder_Lake
Apple does the same thing with recent macOS releases, which are all places in California.
So it's something they've been doing to some extent for almost 20 years.
Now, with Zen 4 and Zen 4c they'll have two different single-thread performance versions to play with.
Seems silly, does AMD really need two different/incompatible sockets for a 16-64 core CPU with 8 channels of DDR4-3200?
It just divides the market, decreases volume per product, and makes everything more expensive. The TR (and TR pro) motherboards and CPUs are low volume speciality items, priced a a premium, and seem to have no advantages over the epyc flavor.
Hopefully we get much cheaper cloud computing because of it.