AMD Is Working on Its Own Hybrid x86 CPU: Patent Filing
extremetech.com
extremetech.com
For me the term Hybrid CPU means a CPU that has more than one core design that runs different architecture. But can see how this could apply if you omit that last caveat, albeit this is power management.
What would be good, is an advanced in design tools to accommodate asynchronous(clockless) CPU design over synchronous(clock fixed ala most CPU's and all X86 today). But then the design overheads really are the limiting factor, many (not all) could be reduced with advances in design tooling. Though be mindful, how long and the state of open source FPGA tooling and how long those have been around.
[EDIT] Fixed synchronous /asynchronous in the wrong place before anybody noticed. And added some clarity)
It's like we wouldn't call a purely electrical car with two different specced motors inside a hybrid.
Perhaps "asymmetric cores" would be a better name?
Asynchronous means that all your busses have to carry, and your logic units calculate, readiness information. That's extra work to be explicit in detail about something that's implicit in clocked designs where it's all pre-calculated and summarized into the (externally documented) maximum frequency.
As I understand it, it makes sense in battery-powered devices. But does it advantages in desktop or server machines?
I vaguely recall reading a blog post - many years ago, though - by a laptop-battery-lifetime-nerd who claimed that running the CPU at a higher clock speed actually saved energy, because it got done faster and could go back to a low-power sleep mode which consumed a lot less energy than the "awake" CPU running at a lower clock speed.
Was this ever accurate? Is it still? I do not know, so I'd appreciate it very much if someone who does know something about this topic could shed some light on this.
"Ten years ago, I asked Intel’s smartphone SoC designers why they were relying on DVFS — Dynamic Voltage and Frequency Scaling — to keep Atom’s power consumption competitive rather than big.Little. According to Intel’s architects, DVFS was competitive with what big.Little could deliver when one considered the silicon die space requirements and the overall power savings. That may have been true at the time — Medfield was generally competitive with the midrange Cortex-A9 devices it was intended to compete against — but it doesn’t seem to be true today."
I assume the idea is that if you're using the LITTLE cores, you turn off the big cores entirely (or at least some of them).
If you have low-priority or very short tasks to run, then you can wake just a few small cores, and that's way less expensive than to wake big cores.
An other advantage is that it can help progress background tasks more efficiently: if you run background tasks on small cores then you don't need to suspend tasks which need the big cores or migrate them between threads to run background tasks.
User tasks want to use the rush to idle because it improves responsiveness. Most background tasks don’t really care.
High performance cores trade power efficient for execution speed by using lots of transistors to analyze the code and optimize on the fly (literally a semi-programmable core in itself to run your core faster). In addition, power consumption is linear with clocks in theory, but not at all in practice.
Let’s consider AMD specifically. In order or semi in order cores are much narrower. All the speculation bits go away. The extra load/store complexity (very significant) also disappears because the core simply can’t use it. SMT will also go away. As it adds a ton of duplication, that also reduces size (10% if I remember Intel’s claims from a few years ago). Simple cores get much better performance per watt and per transistor, but at the expense of time. Makes a lot of sense if time isn’t critical.
On desktop, a zen 3 core at 4.5GHz will use around 12w. Dropping frequency just 200MHz will lower power usage per core by almost 2w. (from Anandtech).
In the space of that one core you mention, you could fit 4 complete in order cores with lots of room to spare. They could all run at a couple GHz and still use less than half the 3.8w you mention.
(https://overclock3d.net/news/cpu_mainboard/intel_s_hybrid_bi...)
Yes - because some tasks are highly parallel and are better suited to running on lots of small cores, and some tasks are highly sequential and are better suited to running on a single bigger core. big.LITTLE is the compromise that suits both.
See Amdahl's Law, Gustafson's Law, etc.
Of course, laptop and desktop CPUs got that feature as well, so that idea still holds true today.
It's kind of like Auto stop/start on engines - saves more fuel than it wastes.
In that case "rush to idle" is not that useful, and waking a big and expensive core (not to mention ramping up its frequency) leads to much higher power consumption for limited gain.
In fact recent big.LITTLE scheduling has been leveraging this: in the oldest model the SoC or kernel would activate either the big or the little cluster[0] depending on what it believed was useful, then the CPU got split "vertically" so the kernel would see each pair of big and little cores as a single vcore, and whichever of the real cores was best for the workload on that vcore would get selected.
More recent big.LITTLE implementations have gone fully heterogenous, if you have 4 small and 4 big cores the system sees and shows 8 actual cores with different performances, and the scheduler tuning selects which tasks go where.
[0] big.LITTLE clusters are the "banks" of respectively big and little cores
AFAIK different C-states dictate that per-core and different caches power on.
Given that, surely this can be extended further with ALUs, FPUs etc only being powered on as necessary? I assume that is already the case no?
Or is this more towards having smaller denser cores with fewer interconnects and components for even higher power efficiency?
Somebody once suggested on HN that slower cores would execute unsafe code (e.g js) and they'd have all those speculation mechanisms turned off, meanwhile fast core would execute local code (e.g ffmpeg) with all those speculative mechanisms ON
that was interesting idea, I'm curious how does it reflect reality
It always seemed odd to me that my whole machine pays the cost for spectre mitigation when the only processes that need it (to me) are JavaScript vms in the browser.
Even my electron apps don't need it, they have access to the filesystem which is far better than any spectre attack.
You should have a syscall to opt into spectre protections. Like "I run untrusted code". Then the OS scheduler can treat it differently, keep the hyperthread only for other untrusted processes like that, flush the cache when scheduling something else on that core, etc.
If they do, then any process can run on any core. But even the little cores have to support fancy SIMD stuff and have hundreds of registers.
If they don't, then compilers need to be aware that simd or other features could suddenly vanish anytime. There needs to be logic for checking CPU capabilities at the start of every loop and switching to a different implementation depending which core we're currently running in. There probably also needs to be some kind of exception handling in the OS to emulate unimplemented instructions for backwards compatibility.
I'm not sure that's the best way to handle this, but I don't think I'm qualified to make a judgment one way or the other.