Intel's Lion Cove Architecture Preview
chipsandcheese.com
chipsandcheese.com
It's also interesting to see that FPGA integration hasn't gone far, and good vector performance is still important (if less important than integer). I wonder what percentage of consumer and professional workloads make significant use of vector operations, and how much GPU and FPGA offload would alleviate the need for good vector performance. I only know of vector operations in the context of multimedia processing, which is also suited for GPU acceleration.
The best case scenario here is if you can have the compiler do all the heavy lifting but more realistically you’ll end up having to make developers switch to a whole new programming paradigm.
- Folly: https://github.com/facebook/folly/blob/main/folly/container/...
- Benchmark, and verify that the code is hot.
- Rewrite from Python, Ruby, JS into a systems language (if necessary). Honorary mention for C# / Go / Java, which are often fast enough.
- Change to better data structures. Bad data structure choices are still so common.
- Reduce heap allocations. They’re more expensive than you think, especially when you take into account the effect on the cpu cache
Do those things well, and you can often get 3 or more orders of magnitude improved performance. At that point, is it worth reaching for SIMD intrinsics? Maybe. But I just haven’t written many programs where fast code written in a fast language (c, rust, etc) still wasn’t fast enough.
I think it would be different if languages like rust had a high level wrapper around simd that gave you similar performance to hand written simd. But right now, simd is horrible to use and debug. And you usually need to write it per-architecture. Even Intel and amd need different code paths because Intel has dumped avx2.
Outside generic tools like Unicode validation, json parsing and video decoding, I doubt modern simd gets much use. Llvm does what it can but ….
And this would highly depend on how ubiquitous they’ll become and how standardized the APIs will be so you won’t have to target IHV specific hardware through their own libraries all the time.
Basically we need a DirectX equivalent for general purpose accelerated compute.
Some cores will be high-performance, OoO CPU cores.
Now you make another core with the same ISA, but built for a different workload. It should be in-order. It should have a narrow ALU with fairly basic branch prediction. Most of the core will be occupied with two 1024-bit SIMD units and a 8-16x SMT implementation to hide the latency of the threads.
If your CPU and/or OS detects that a thread is packed with SIMD instructions, it will move the thread over to the wide, slow core with latency hiding. Normal threads with low SIMD instruction counts will be put through the high-performance CPU core.
I think it's reasonable for the non-SIMD focused cores to do so via splitting into multiple micro-ops or double/quadruple/whatever pumping.
I do think that would be an interesting design to experiment with.
Obviously there are compromises in terms of bandwidth but it’s also a lot easier to mix into a broader program if you don’t have to send data across the bus, which also gives it other potential use-cases.
But, if you take the CUDA lane idea one step further and add Independent Thread Scheduling, you can also generalize the idea of these lanes having their own “independent” instruction pointer and flow, which means you’re free to reorder and speculate across the whole 1024b window, independently of your warp/execution width.
The optimization problem you solve is now to move all instruction pointers until they hit a threadfence, with the optimized/lowest-total-cost execution. And technically you may not know where that fence is specifically going to be! Things like self-modifying code etc are another headache not allowed gpgpu too - there certainly will be some idioms that don’t translate well, but I think that stuff is at least thankfully rare in AVX code.
Static analysis would probably work in this case because the in-order core would be very GPU-like while the other core would not.
In cases where performance characteristics are closer, the OS could switch cores, monitor the runtimes, and add metadata about which core worked best (potentially even about which core worked best at which times).
The key part is that now there are far more use cases than there were in the early dozer days and that the current main CPU design does not compromise on vector performance like the original AMD design did (outside of extreme cases of very wide vector instructions).
And they are also targeting new use cases such as edge compute AI rather than trying to push the industry to move traditional applications towards GPU compute with HSA.
This is in part (major part IMHO) because few languages support vector operations as first class operators. We are still trapped in the tyranny that assumes a C abstract machine.
And so because so few languages support vectors, the instruction mix doesn’t emphasize it, therefore there’s less incentive to work on new language paradigms, and we remained tapped in a suboptimal loop.
I’m not claiming there are any villains here, we’re just stuck in a hill-climbing failure.
C or Java have no concept of `a + b` being a vector operation the way a language like, say, APL does. You can come closer in C++, but in the end the memory model of C and C++ hobbles you. FORTRAN is better in this regard.
It is always possible to inline assembler in C, and present vector operators as functions in a library.
Otherwise, R does perceive vectors, so another language that performs well might be a better choice. Julia comes to mind, but I have little familiarity with it.
With Java, linking the JRE via JNI would be an (ugly) option.
And CUDA is Nvidia specific.
There are other libraries (e.g. OpenMP, Intel's oneAPI) and languages (e.g. SYCL) that do let the same code be run on either CPU or GPU.
Asking sincerely: what’s specifically so interesting about that? That is what I would naively expect.
Hardware designers are adding a lot of speciality hardware, they're just not putting it into the core, which also makes a lot of sense.
https://www.researchgate.net/figure/Architectural-specializa...
One with SMT, for server CPUs, and one in which all circuits related to SMT are completely removed, for smaller area and power consumption, to be used in hybrid CPUs for laptops and desktops, together with Skymont cores.
I like these architectures but don't expect it to happen. Every attempt to produce a CPU that works this way has done poorly because software developers don't know how to use them properly. The hardware companies have done the "build it and they will come" thing multiple times and it has never panned out for them. You would need a killer app that would strongly incentivize software developers to learn how design efficient code for these architectures.
In many ways barrel processing and hyperthreading are remarkably similar to VLIW - they're two very different sides of the same coin. At the end of the day, both work well for situations where multiple simultaneous instruction streams need to access relatively small amounts of adjacent data. Realtime signal processing, dedicated databases, and gaming are obvious applications here, which I think is why VLIW has done so well in DSP and hyperthreading did well in gaming and databases. Once the simultaneous instruction streams are accessing completely disparate parts of main memory, it all blows up.
Plus, in a multi-tenant environment, the security issues inherent to sharing even more execution state context (and therefore side channels) between domains also become untenable fairly quickly.
Of course, the big elephant in the room is security - timing attacks on shared cores is a big problem. Sharing anything is a big problem for security conscious customers.
Maybe it's the case of the server leading the client here.
Another benchmark I've done is .onion vanity address mining. Here it's about a 20% improvement in total throughput when using hyperthreading. It's definitely not useless.
However, I didn't compare to a scenario with hyperthreading disabled in the BIOS. Are you telling me the threads get 20-50% faster, each, with it disabled?
RAM fetch latency is what happens on CPU level.
And like you said, the vulnerability consideration makes HT a hard sell for workloads that would benefit (hypervisors).
That i7 is a huge downgrade from what I was running before (long story) so I'm looking forward to Arrow Lake and I like everything I've read. In addition to removing HT, they're supporting Native JEDEC DDR5-6400 Memory so XMP won't be necessary. I've never liked XMP/Expo...
And since hyperthreading only increases IPC by 30% (according to Intel), they're planning on making up the loss of threads with more E-cores.
But we'll have to see how that turns out, especially since Intel's first chiplet design (the Core Ultra series 1) had performance degradations compared to their 13th Gen mobile counterparts
The only issues with these architectures is that they are priced as "high-end server stuff" while x64 is priced like a commodity.
While there are also other tasks where SMT does not bring advantages, for the compilation of a big software project SMT does bring an obvious performance improvement, of about 20% for the same Zen 3 CPU.
In any case, Intel has said that they have designed 2 versions of the Lion Cove core, one without SMT for laptop/desktop hybrid CPUs and one with SMT for server CPUs with P cores (i.e. for the successor of Granite Rapids, which will be launched later this year, using P-cores similar to those of Meteor Lake).
On the high utilization end, stuff like offline rendering or even some realtime games, would have significant performance degradation when HT/SMT are enabled. It was incredibly noticeable when I worked in film.
And on the low wattage end, it ends up causing more overhead versus just dumping the jobs on an E core.
For most of the HT's existence there weren't any E cores which conflicts with your "never" in the first sentence.
The difference is that now low wattage doesn’t have to mean low performance, and getting back that performance is better suited to E cores than introducing HT.
Saying "no" doesn't magically remove your contradiction. E cores didn't exist in laptop/PC/server CPUs before 2022 and using HT was a decent way to increase capacity to handle many (e.g. IO) threads without expensive context switches. I'm not saying E cores are a bad solution, but somehow you're trying to erase historical context of HT (or more likely just sloppy writing which you don't want to admit).
We could politely discuss it or you can continue being rude by making accusations of sloppy writing and denials.
It was a product of its time, a way to get cheap multi-cores when getting real cores was too expensive for regular consumer products.
Besides the security issues, for high performance workloads they have always been an issue, stealing resources across shared CPU units.
Performance is still a reason. Anecdote: I have a pet project that involves searching for chess puzzles, and hyperthreading improves throughput 22%. Not massive, but definitely not nothing.
SMT is a crutch. If your frontend is advanced enough to take advantage of the majority of your execution ports, SMT adds no value. SMT only adds value when your frontend can't use your execution ports, but at that point, maybe you're better off with two more simple cores anyway.
With Intel having small e-cores, it starts to become cheaper to add a couple e-cores that guarantee improvement than to make the p-core larger.
Apple does a lot better. I'm not sure about newer chips, but M1 has the same 192kb of L1, but with the 4 cycle latency of Intel's tiny 48kb cache.
And, you know, stop the security vulnerability bleeding.
Amateur question: is that due to the recent advances in interconnect or are they designing multiple versions of this chip?
A Ferrari will beat a tractor on every test bench numbers and every track, but I can't plow the field with a Ferrari, so any improvements in tractor technology is still welcome despite they'll never beat Ferraris.
If you're talking apps before say 2015 (10 years ago), they can be emulated on ARM faster than they ran natively. That rules out 95% of the backward compatibility argument.
Most more recent apps are very portable. They were written in a managed language running on a cross-platform runtime. The source code is likely stored in git so it can be tracked down and recompiled.
Over 15 years of modern smartphones has ensured that most low-level libraries have support for ARM and other ISAs too as being ISA-agnostic has once again become important. Apple's 4 years of transition aren't to be underestimated either. Lots of devs/creatives use ARM machines and have ensured that pretty much all of the biggest pro software runs very well on non-x86 platforms.
Yes, some stuff remains, but I don't think the remaining stuff is as big a deal as some people claim.
Non-embedded POWER implementations are around 1000 opcodes, depending on the features supported, and even MIPS eventually got a square-root instruction.
> The key operational concept of the RISC computer is that each instruction performs only one function (e.g. copy a value from memory to a register).
and in fact that page even mentions at https://en.wikipedia.org/wiki/Reduced_instruction_set_comput... that
> Some CPUs have been specifically designed to have a very small set of instructions—but these designs are very different from classic RISC designs, so they have been given other names such as minimal instruction set computer (MISC) or transport triggered architecture (TTA).
13900K lost couple of % in single tread performance, which lead to 14900K being so overclocked/overvolted that it lead to it being useless for what it's made for - crunching numbers. See https://www.radgametools.com/oodleintel.htm.