Measuring energy usage: regular code vs. SIMD code
lemire.me
lemire.me
GPUs take this even further by going even wider, with each "core" typically executing 1024-bit SIMD operations (i.e. 32x FP32, 64x FP16, etc). CPUs have roughly settled at 128-bit (most ARM) or 256-bit (most x86) with a little bit of 512-bit (x86 with AVX512) which grew out of Intels earlier aborted attempt at making a dedicated GPU.
I don’t think inter-core communication is too relevant when comparing vectored and non-vectored on a single core, but definitely would be when batching across multiple cores.
(Presumably doing this would only be possible in the context of a single-threaded uninterruptible unikernel workload with system-management-mode functions disabled — but I’m sure a lot of “one powerful single-core SoC”-type embedded systems would be happy to make that trade off!)
Right now, AFAICT, even when executing a uOp stream from L0 cache, several "planning" systems are still active — juggling caches around and deciding routing between them; renaming registers; switching between power "license" states; etc. None of these decisions are explicit even in uCode for modern CPUs, so they have to be made, over and over again, even when running from uOp cache. Which means the silicon that's making these decisions can't ever go dark, even when running from uOp cache.
A "nanocode" would be a version of microcode that burns in all these "planning"-system decisions — and which can thereby put all the "planning" silicon to sleep. There would be "nanoOps" for explicitly wiring registers to other registers or cache buffers, for adding precise numbers of delay cycles, etc. And these would all happen at Nth-of-a-cycle-precise points in the execution stream, indicated by either entire nOp instructions, or pragma bits on nOp instructions.
(Given this, to usefully function, these "nanoOps" would either likely need to be bytecode and be decoded+executed at some ridiculously high frequency relative to the regular CPU clock — or they would need to be some arcane parallelized packing of what was originally a serial stream of uOps, such that you get VLIW nOps where all the pragmas you want to have happen for the next CPU cycle are indicated in the instruction-word at once as individual signal-line bits. Just like the 6502 PLA output "instruction word", actually!)
I know this probably sounds like nonsense — the instruction stream would be so much more bloated than microcode that it probably wouldn't be worth it. But that assumes an instruction stream that needs to live in RAM and get shuttled through layers of caches to reach the decoder. But what if the instruction stream could be loaded into a reserved SRAM area within the CPU itself — an SRAM that would effectively act as EEPROM (in that it would be loaded from CPU NVRAM on boot); and which you could directly drop the instruction pointer into (in nCode decode mode), skipping RAM+caching entirely?
I mention this, because I've always had a vague hypothesis that Intel x86 CPUs specifically already have something akin to this implementation of a "nanocode" + reserved SRAM area that holds some of it — specifically for the purpose of programming custom dynamic RTL for instructions after CPU release to hotfix CPU errata. It's something they had to learn the necessity of over and over again, after releasing CPUs with bugged instructions, all the way back to the FDIV bug.
If esoteric hardware interests you I would recommend IBM system/360 as reading material. Sadly hardware like that has fallen out of fashion and we've been stuck with the x86 and arm dichotomy. if anyone else has interesting/newer non-x86 hardware to link I'm more than happy to delve into that.
A. you do byte-level processing instead of float words;
B. you use embedded, IoT, and other low-energy devices.
A few years ago I've compared Nvidia Jetson Xavier (long before the Orin release), Intel-based MacBook Pro with Core i9, and AVX-512 capable CPUs on substring search benchmarks.
On Xavier one can quite easily disable/enable cores and reconfigure power usage. At peak I got to 4.2 GB/J which was an 8.3x improvement in inefficiency over LibC in substring search operations. The comparison table is still available in the older README: https://github.com/ashvardanian/StringZilla/tree/v2.0.2?tab=...
Edit: found it (I think) I can't remember where in the presentation but he does mention energy efficiency as a proxy for performance. https://m.youtube.com/watch?v=kLiwvnr4L80
Edit 2: about 26:40 he starts talking about energy use in the context of performance.
Loading (fig IV)
Location Energy (pJ = pico Joules)
L1 64 pJ/Byte
L2 121 pJ/Byte
L3 254 pJ/Byte
RAM 1250 pJ/Byte
Adding integers with SIMD (fig VI)
428 pJ/op, for 8 Byte/op; this means:
53 pJ/byte
So it takes actually 20 times to fetch data from RAM than to add it to something! And most often this is also the source of latency. Generally energy is linear to the distance signal has to travel, and that's the same for latency.That's why successful data structures are sized a tiny bit under L1/L2 sizes! (BTree chunks, ring buffers).
If you've been following hardware, it's all about putting RAM closer to compute at the moment with chiplets at the moment.
[1] https://tu-dresden.de/zih/forschung/ressourcen/dateien/abges...
isn't this another formulation of the race-to-idle thinking?