AVX10/128 is a silly idea
chipsandcheese.com
chipsandcheese.com
Two decades ago I put a Pentium III and a Celeron on to the same dual processor motherboard (using a Celery to Nintendo cartridge adapter), and said, "Look! Different processors can run at the same time in the same system. Perhaps we should have a way for processes to express a desire to use specific CPU features, and if they request features not available on all CPUs, we'll keep them on the ones that do." Sigh.
Intel should've told OS writers years ago that heterogenous CPU features would be a thing soon, so prepare for it. Then we could've easily had AVX512 on big cores and AVX256 on little cores, and it'd've all be just fine.
Sure, old apps might not then have that flag, but old apps aren't using avx512 either.
As mentioned elsewhere, I'm not sure this is hardware design being forced to adapt due to software implementation difficulties, so much as avx512 as a whole may not be worth the hardware area going forward. Even on the larger cores.
Note that even simple things like memcpy, which every single program out there will use, will use AVX512!
But really I think the entire ecore/pcore split for avx512 is academic, as I'm not sure tying features and flexibility to wider registers makes sense even on pcores, as I'm not sure the hardware area cost is worth the benefit. I honestly wouldn't be surprised if newer architectures don't have the wider registers options on even their larger cores.
You're chasing a pretty small market IMHO of people who have datasets large enough to benefit from large registers and wider alus, but not so large it's worth it to pass it over to an even more specialized accelerator.
The benchmarks that tend to show 512-wide simd benefits may get even bigger benefits from running on a GPU.
Why not just determine the right cpu to run on by examining the arch of the binary? Waiting for an instruction failure seems ridiculous.
If the arch is incorrect, it is a bug, and program will crash on illegal instruction. Ie like if you had an ARM binary that was incorrectly set to x64 and was run on x64.
What if the JIT generates new instructions dynamically and initially there are no AVX512 instructions but later on in the process' lifetime there are?
For this reason, most binaries built with modern Visual C++ are technically using AVX instructions even when compiled for SSE only. It doesn’t mean these binaries gonna fail without AVX support, it only means they’re capable of using AVX when available.
I'd think you could do some pretty good heuristics, like if the thread hit P-core instructions in the last time slice, don't schedule it on an E core for the next one. When a time slice starts on a P-core, leave the specialty instructions disabled, so you can monitor usage --- if you trap, enable it, and return; if that cost is still too high (which it might be), keep track of how many time slices in a row hit the trap, and maybe enable the P-core instruction preemptively for a few slices.
Or just, give the program more information and let it decide. If you've got some cores with avx and some without, maybe the process wants to schedule only on the avx cores. Or maybe it can schedule some threads on any core and others need an avx core. As long as the possible permutations at run time aren't too crazy, it's reasonable.
I'm in the trenches with AMX right now, and it's not without its challenges. When Nvidia rolled out the `wmma` intrinsics for tensor cores years ago, they felt a bit unconventional. But as technology evolves, these new extensions seem to be increasingly idiosyncratic, challenging to navigate, and hyper-specialized.
Hardware improvements significantly outpace compiler optimizations. So hand-written kernels are becoming the norm again :)
There is definitely a growing cottage industry for writing high-performance kernels using this increasingly rich set of intrinsics. I find myself writing a lot more intrinsics code than I used to.
See that bug for the example what gonna happen https://bugs.chromium.org/p/chromium/issues/detail?id=121838...
TLDR: Windows ABI conventions say xmm6-xmm15 vector registers are non-volatile. When a function uses them, it needs to backup old data, and restore before the return. Linux ABI conventions say all vector registers are volatile and don’t need to be preserved. Chromium has non-trivial pieces written in assembly, and crashes.
When writing C or C++ with intrinsics, registers are managed by the compiler. Modern ones follow these ABI conventions very carefully.
Every Android phone and ChromeOS device is its own special snowflake in hardware configuration, regarding RISC, ARM and Intel CPUs, or GPUs for that matter.
The runtime would basically have to detect which AVX-512 instruction has failed at any point in the compiled code, map the processor state to AVX-2 registers and jump to an equivalent AVX2 code path. It would probably be easier to build the implementation on software transactions that can be restarted at coarse granularity, i.e. something like STM.
However ART is also a great piece of engineering at JVM and CLR level, even though Dalvik was quite lame, even when compared against Nokia and Sony-Ericson J2ME implementations.
ART is another level.
Nowadays you have an handwritten Assembly interpreter for faster startup, a JIT compiler, that also gathers PGO data, and an AOT compiler that picks up that data when the device is idle taking all the required optimizations.
To take this even furter, those PGO files are uploaded into the Play Store, and shared across compatible devices when applications are downloaded, not only allows jumping straight into JIT with PGO data, all devices collaborate gathering PGO data that covers all common use cases across the application audience.
And yes this also takes into account the ARM architectures.
On the JVM side, IBM and Azul have similar offerings with cloud based JITs for their implementations.
Even SVE must use the same vector width across all cores, which makes it unlikely that >128b SIMD will ever show up on consumer cores in the near future: https://gist.github.com/zingaburga/805669eb891c820bd220418ee...
https://developer.arm.com/documentation/102484/0001/The-Cort...
https://developer.arm.com/Processors/Cortex-A720
https://developer.arm.com/Processors/Cortex-A520
You will notice that for e.g. Asymmetric MTE isn't supported across all of them, and the differences in microarchitecture are enough to have impact in the kind of instructions being used.
But it's worth pointing out that X4, A720 and A520 support asymmetric MTE [1] - just the A720 page you linked to fails to mention it.
[1] https://www.arm.com/blogs/blueprint/memory-safety-arm-memory...
https://sky.cs.berkeley.edu/events/database-seminar-helena-c...
Very interesting work about associative processing: “The Associative Processing (AP) paradigm leverages content-addressable memories to realize bit-serial arithmetic and logic operations, via sequences of search and update memory operations over multiple memory rows in parallel.”
[1] Sean Parent: "Now What? A vignette in three parts" https://www.youtube.com/watch?v=iGenpw2NeKQ&t=1382s
[2] Keynote: The Tragedy of C++, Acts One & Two - Sean Parent - CppNorth 2022 https://youtu.be/kZCPURMH744?t=2924
Current CPUs already do this, and have been doing so for quite some time. And AVX10/128 doesn't alleviate it either.
Me: "Page one of the unified standard lists its many variants, which seem to be the same number that existed before?"
Intel: "We've renamed them all! See! Unified!"
ARM SVE2 and RISC-V Vector 1.0 are vector extensions. AVX is just SIMD with marketing.
So either SVE2 isn't vector due to lack of VL, or AVX-512/AVX10 is vector if VL isn't necessary (the only difference between AVX and SVE2 being the scalable registers, which Cray-1 doesn't have).
SVE and RVV are no different in that regard. Note that RVV's VL doesn't alter the vector length - it's just an alternative way of masking (which can easily be implemented in SVE).
If setting VL is enough to qualify (which seems to be such a minor differentiating factor), then I'd argue that SVE also qualifies since you can use WHILELT to do exactly the same thing as setting VL. SVE2.1 basically makes the whole thing identical, with its support for predicate-as-counter.
SVE2.1 predicate-as-counter seems interesting, but I can't find any overview of it (only specific instructions/intrinsics), would you happen to know of any?
I doubt I know any more than you. ARM's info on the matter is a little scarce, and I struggle to understand their documentation at times.
The Atom-derived cores did not support AVX for quite a long time. Even the Tremont microarchitecture (i.e. Elkhart Lake launched in September 2020 and Jasper Lake launched January 2021) does not support AVX. Only from the Gracemont microarchitecture on (launched in November 2021), there exist Intel Atom processors that support AVX/AVX2.
With the Gracemont microarchitecture, this "principally" changed, but until now Intel has not announced any processor of this microarchitecture that it branded "Pentium" or "Celeron":
> https://en.wikipedia.org/w/index.php?title=Gracemont_(microa...
Some Atom line processors were branded Pentium/Celeron. The low end big core SKUs were also branded Pentium/Celeron (and often had features disabled).
Additionally, because it consumes half as many register bits as the 512 bit profile, it could be a lot more economical (and power efficient) to implement on cores that are not the high end server parts. Keep in mind ARM (even the top of their line) is still stuck on 128.
So I think the case for this is a good one.
If memory bandwidth is a significant factor in your execution time, doubling the SIMD vector size is going to give you far less than the 100% theoretical speed boost (before taking into account thermal throttling, of course). Furthermore, the execution penalties due to things like remainder loops or vector misalignment (every misaligned 512-bit vector crosses cache lines) are comparatively more painful. The fancy new features ameliorate those downsides a lot and furthermore make vectorization profitable on a larger fraction of your code.
And finally, if the limiting factor of your code's execution time is the number of FMA units you can ram data through, then at this point, you've already likely switched to using a GPU to do that instead of the CPU.
It is a SIMD extension, despite Vector era (RISC-V Vector 1.0, ARM SVE2). Far from relevant.
>desktop CPUs
Will too be RISC-V. Not far from now.
Some chips (likely from Intel and AMD) will offer "x86 acceleration" to run legacy software, for a transitional period lasting a few years.
FUD unless you give specific examples and evidence backing them.
Many "seemingly insane/backwards decisions" in RISC-V are good decisions backed by large bodies of data.
Refer to Computer Architecture: A Quantitative Approach[0].
0. ISBN-13 978-0128119051
Though to be fair, it all depends on the hardware. Windows on ARM struggled not just because it wasn't x86 but because nobody was making powerful-enough ARM processors. Now that those processors are starting to arrive, it's supposedly going to take off in the next few years. …Probably only on the low end, though. But you never know.
And RISC-V? Maybe someday.
Is it really that high? Where do these figures come from?
However, 90% of them are Macbooks, which means the share of the ARM laptops which run Windows, Linux or ChromeOS is 1.4%.
Historically, bottom feeders won in IT. unix was a bottom feeder againts the IBM os'es and VMS. x86 was a bottom feeder architecture against mainframes, server class hardware, and workstations. Linux was a bottom feeder against the real UNIXes like SunOS and HPUX.
Bottom feeders get picked as cheapest option, then client demand of better features slowly pushes them upwards. Meanwhile the top concentrates on its most profitable customers, showing disdain for bottom feeders until they've grown out of control.
The last 5 Tandem owners are probably still out there somewhere, showing all of us what real uptimes look like.