AMD's Mild Hybrid Strategy: Ryzen Z1 in Asus's ROG Ally
chipsandcheese.com
chipsandcheese.com
There are also desktop APUs using a combination of Zen 4 and Zen 4c cores (Ryzen 3 8300G, Ryzen 5 8500G), as well as other mobile CPUs doing the same (Ryzen 3 7440U, Ryzen 5 7545U).
I do a lot of highly parallelized image processing and preprocessing for machine learning and I don't want any cores holding other cores back.
It is pretty common for a CPU to have different clock speeds depending on how many cores are active (not to mention AVX clocks, so clock speed is just a thing that bounces around). This doesn’t seem drastically different.
I have more reading to do on the AMD side, so if you have more info than I take this with a grain of salt, but at least on the Intel side of the camp, the performance of a fully saturated E core is only slightly worse than its P core buddy, with 2x E cores meeting or exceeding the work output of a single P core + its hyperthread.
Even if the only purpose for the E cores is to run OS and other background workloads while your big processing job is getting unfettered access to the P core array, I imagine you'd see pretty great gains here.
So imagine you set a given budget for silicon die size and chip power. You might then have 3 design possibilities. Say 8 big cores, 16 little or 12 in hybrid. The hybrid design should be able to match the all-big in single/lightly threaded loads, while beating the all-big in highly multithreaded. These tradeoffs of course work best in form-factors that are constrained by cost or battery power. Which is where you see amd using them in consumer. Or by space/power where you see AMD giving choice of small/dense core options in Epyc.
The only such example I'm aware of is Intel's recently-launched Meteor Lake laptop chips that include an extra two low-power E cores on a separate chiplet from the rest of the CPU cores, running at half the performance of the regular E cores and with no L3 cache, so whether they're used or not only makes a few percent difference to embarrassingly parallel workloads.
The "same number of threads" bit is throwing me off. A hybrid config will generally have more CPU cores allowing it to handle more threads for a chip of the same size, so testing with a fixed thread count can only tell you that the performance of each small core is less than the performance of a large core. But maybe the examples you have in mind are complicated by SMT/HyperThreading (which Intel has on P cores but not E cores, while AMD has it on both)?
Bad design: see N cores, split work into N partitions, start N threads, each thread processes one partition.
Good design: see N cores, start ~N threads, each thread takes 1/M of work at a time, M>>N.
This is really not true. With a decent motherboard and sufficient cooling, you can force the big cores to stay virtually locked at peak clock speeds. Der8auer does it for a Ryzen 8000 APU here just as an example: https://youtu.be/VNYx72Elgss?t=335
But if you have two different kinds of cores with different clocks and no control over the scheduler it is not a good compromise.
I have an Intel N100-based mini PC. As far as I can tell the N100 is not hybrid... because it only uses E-cores.
They were all horrible to use. They constantly throttled to the point that I eventually ripped the dell open and repasted everything inside of it including de-lidding the cpu. I had to refoil and do a bunch of other stuff you wouldn't have to do on a desktop to repaste it was honestly nerve wracking and I've done all sorts of custom desktop pc water loop pasting stuff. Point being that it was an awful thing to have to do but the xps ran SO bad I did it. I gave the dell to my mom and I think I gave that samsung away too, I def don't have it anymore. And then I was done with intel ulv procs.
That dell xps 13z was the prettiest, nicest form factor laptop I've used and it SUCKED and it was so sad.
2 months ago I bought an ally and I was absolutely terrified that it would run the same as those.
It runs GREAT. It feels almost like using my amd ryzen 7700x, obviously not the same power but it's snappy and I don't mind using it as a laptop in bed at all. They did a great job. I have mine with a case on it that has a laptop keyboard/trackpad and it holds the screen up in bed, and I just use a travel mouse with it. I do wish the screen was larger though, so I wonder if I'd be happier with the Legion.
As far as I'm aware, laptop CPUs haven't been packaged with an integrated heat spreader at any time since the ultrabook category was defined, so I have no idea what you are describing here.
edit: Oh, I found a 2022 post about here on HN. I guess I didn't de-lid it. I took pictures of everything for the tutorial but no idea where that'd be, reddits search for dell 13z is worthless. Wow I didn't have that laptop very long. I absolutely hated using it, I wouldn't even use it in bed, and I DID get it replaced at one point hoping it was a bad laptop.
https://news.ycombinator.com/item?id=31097858
Narrator voice: He did buy the $3500 macbook in the end
edit2: Tutorial https://old.reddit.com/r/Dell/comments/vqz9wn/xps_13_2in1_93...
This is giving me terrible flashbacks I even broke a few things, there were hidden screws holding board down you had to remove from the underside somehow or break everything when you lifted the board up. I had NO idea they were there and everything was super fragile.
On the bright side, phoronix said the detachable controller drivers are (about to be?) merged into one of the latest kernels. That's pretty remarkable.
What do you mean by this? Like you think it's better as a productivity or media consumption device? I never use my Ally for watching anything because the screens so small and I have an ipad, I kind of figured I'd use it more if I got the legion instead.
I imagined using the device, with its stand, on my tummy. Mostly reading but also hosting stuff. It's about 650 grams without the controllers, so not ideal, but better than my 1.4 kg laptop. Better than a phone.
Minisforum is about to release a high end Ryzen 8000 14 inch tablet real soon. Still a kilo, though.
The perf impact of this or any hybrid design is also blunted by the power/cooling limits on non-hybrid designs: you can't afford to blast all the full-size cores on the chip at the max theoretical frequency at once anyway. Then, at the other end, lightly-threaded workloads can run mostly or only on the larger cores, so they're not heavily impacted either.
If there's a way this ends up fun for consumer folks that follow this stuff, it'd probably be by allowing products w/more cores. In server chips using Zen 4c (Bergamo), more cores is explicitly the focus. On their desktop platform, one CCD of full-sized cores and one of smaller cores would allow a 24C two-CCD package. That would add another "flavor" of CPU, the way X3D is another "flavor", but I suspect there are applications where the extra throughput could help. No particular hints that this is part of their plans, to be clear.
The Zen4c core has a comparable size to the Gracemont E-cores, so they don't really need a separate design.
I hope AMD would at least double the channels on desktop and laptop. How else are they going to feed the future beefy APUs?
Will be interesting to see AMD try this in consumer processors, with frame time graphs.
edited
NITS not bits. My bad.
BFI stands for "Black Frame Insertion" and it enhances motion resolution in LCD/OLED/Sample and Hold displays by inserting a black frame or blackout between regular frames. This has the result of reducing motion blur by shortening the time each frame is visible, making fast-moving content appear sharper and clearer.
It effectively mimics in a way older CRT monitors, improving the perception of motion. Which for me is huge, because I still game on CRT's on the regular and LCD's abysmal handling of motion is a major detriment to me of my enjoyment of such things, especially for titles that were designed and produced for displays that had much clearer motion.
It's a niche use case i'll admit, and i'd never claim otherwise, and it's not for everybody. But its funny that without that capability, id trade it in for an OLED deck tomorrow.
Sadly OLED is not able to escape the problems inherent to sample and hold displays when it comes to the blurring of motion.
Old mate Mark from Blurbusters (An expert in the field) does a far better job of explaining why then I can:
It gets you really close, but not all the way to the motion clarity of a CRT. It's not at the point yet where I would remove my CRT's from my desk, but another few years of brightness increases and resolution increases in <27" and under OLED's, and I could be almost convinced to finally retire them.
Funnily enough VR headsets might be some of the best displays for retro due to their ability to do really good low persistance strobing without something called crosstalk.
Another drawback I forgot to mention is that if you couldnt perhaps tolerate 60hz on a CRT, you may find the same with BFI/strobing, so again, it rules it out for a lot of people.
Normally it’s closer to 25% off and that is still uncomfortable to view compared to plasma or CRT, unless you are blessed with a golden visual system in your brain
If you see one now you'll see how bad it is.
Only the latest super high refresh OLEDs are approaching that capability and the software is not done. Nobody is simulating the rays right now. Pretty niche, so progress is slow.
I also use BFI with my OLED and sure, BFI at 240hz is better but it’s not available to me in a handheld.
According to AMD, both the big core and the small core use the same RTL design, but a different physical design, i.e. they use different libraries of standard cells (optimized either for high speed or for low area and low power consumption) and different layouts in the custom parts.
For a given target frequency, the synthesis tool will always use the most efficient transistors it can. And the result is a mix, using the few available types. But the highest the frequency, the higher the proportion of faster and bigger transistors in the mix.
This is the bird's eye view and very simplified, but hopefully enough to get the idea ;)
I think this post just enlightened me to the EE involved in chips than I've learned over 15 years.
Intel goes with here are some real beefy cores who can do anything , here are some weaker core who can do only some task.
AMD goes here are half of the cores who can go real fast, here are half core who must remain slower, but everyone can do everything.
In theory, Intel could have better perf if optimized for, while AMD could have better perf with any generic random app out there... As long as the OS has enough hint to put the right app on the right core, and bothers to do it.
On the other hand, AMD only really has their Zen series of cores to use, but they rely more than Intel on automated layout tools so they can more easily port designs to a different fab process or do a second physical layout of the same architecture on the same process.
This isnt true any more as of Intel's current CPUs (Meteor Lake). Both P and E cores support the same instruction set, including AVX10.
[1] Back in the day due to the intrinsic capacitance of the transistors themselves. These days more because bigger transistors are further apart leading to more line capacitance.
FWIW Intel supposedly “fused off” avx-512 in alder lake though I don’t think that was what was actually done, physically speaking.
Otoh, modern power management involves clock gating --- turning off the clock in specific clock domains that aren't being used at the moment; having fewer clock domains makes that less granular and potentially less effective.
Other's points about individual transistors being smaller for a lower frequency design also applies. There may be other complementary benefits from lowering the frequency target too.
But note, it's not magic. The Zen4c server parts, where design area had been most disclosed, use a lot less space per core, and for L1 cache, but L2 and L3 cache take about the same area per byte as on Zen4.
A lot of what Apple started doing with A12 ( or even planned well before A12 ) and finally known to the PC world with A14 / M1 will be coming to Zen 5 and Zen 6.
Apple have been on this design for at least 3 iterations or 6 years. Zen 5 will be a ground up x86 design with wide decode by default ( i.e not something bolted on ). It will probably take some learning curve for them as well. The only other wide decode stage design I am aware of is POWER10. But then we never have enough information about it to read through. And it isn't clear how well the latest x86 ISA fits in compared it AArch64 which is fairly clean without AArch32 baggage.
The Anandtech article [1] on M1 goes over this. But again none of these are new. If you follow the Apple CPU Core design trend. And there are insane amount of smaller details that matters. And the industry are all trending on similar wide decode designs in terms of high performance core, from Cortex X4 and the up coming X5 to Qualcomm's Oryon core. I would envision in 4-5 years time, unless Apple has any other breakthrough. All Performance CPU Core design to be converging into something fairly similar from a high level. This also echo what was witnessed by a Chip And Cheese article. ( Which I dont have time to find out the link so people will have to do some digging themselves :) )
[1] https://www.anandtech.com/print/16226/apple-silicon-m1-a14-d...
I'm sure intel & AMD know fully well that there's lots of benefit to decoding wider, it's just significantly more expensive to do so on x86 than it is on aarch64 and thus they spend silicon where it's more effective to do so.
Apple has just designed the CPU cores with the highest IPC among all known until now.
The highest IPC is the best choice for minimum energy consumption in a CPU with few cores, like a mobile phone or light laptop CPU, because it allows the same performance at a lower clock frequency.
A higher IPC requires more transistors and more money spent for the design of the core. Apple has been rich enough to afford both the design cost and the cost of being able to use exclusively more advanced manufacturing processes.
When the IPC is increased to very high values, the area and the power consumption are increased by greater factors than the increase in performance.
Therefore, for a CPU with many cores, like a server CPU, the highest IPC is not the best. There will be a lower IPC that is better, because it allows packing more cores in a given chip area and a given thermal design power, leading to a higher aggregate performance for multi-threaded tasks.
The optimum IPC depends on the characteristics of the CMOS manufacturing process. Unfortunately, the design rules of the up-to-date CMOS processes are secret and in recent years the CPU manufacturers have published much less information about their designs than it was customary decades ago.
Therefore, based on the publicly available information it is impossible to estimate which would be the optimum IPC for a CPU with many cores.
Nevertheless, it is likely that the IPC of the big Apple cores is too high and even the IPC of the big Intel and AMD cores is likely to be too high.
The IPC of the Zen 4c cores may be closer to optimal in the current technology, but the companies which design CPU cores probably cannot afford to explore enough of the design space to determine which would be a really optimum IPC.
Weird comment. 'Optimum' can only be defined with reference to a given task. Otherwise it's like talking about the correct top speed for a car – F1 racer, or a little runaround?
If economic criteria like money spent for the design are used for optimization, than it becomes pretty much impossible to compare the competing companies, because almost every CPU design might be considered to have an optimum IPC, no matter how bad it is, because it is the best that could be achieved within their budget, by their design team.
Because Apple can sell expensive products, the die area and the cost of the manufacturing process had little importance for them. Because their products use only few cores, the area of one core was not constrained, because the total die area remained below the process limits, when the core area is multiplied by the number of cores.
So their main optimization criterion has been the ratio between performance and energy consumption for a single core, without area per core and power dissipation per core constraints. In this case the IPC is not constrained, there is no optimization problem for determining the IPC, the higher IPC, the better. Therefore, unlike the designer of cores for a server CPU, Apple has designed the cores with the highest IPC that they could achieve within the project schedule and budget, in a given TSMC process.
On the other hand, the customers, except for the biggest customers who might receive samples for in-house testing before purchasing, cannot estimate accurately those costs, because the vendors avoid to provide adequate information, but the customers must depend on things like published benchmarks and extrapolations from older systems.
So the CPU vendors have the incentive to skew their designs to get good results in some popular benchmarks, even when this policy results in designs that are inferior at better criteria.
For example, for many years Intel has included the secure hash extensions only in their Atom CPUs, even if they would have been more useful in their other CPUs. This was caused because Geekbench included a SHA subtest, the ARM CPUs already had SHA instructions and Intel wanted for their Atom CPUs to reach Geekbench scores competitive with ARM. Only after Zen has also added SHA, the Intel Core and Xeon CPUs have also implemented it.
The desktop CPUs spend a lot of resources for obtaining the best scores in the most popular single-threaded benchmarks.
While a responsive computer is very desirable, the truth is that whenever a computer with a 5-GHz CPU reacts slowly to user input, that is guaranteed to be caused by badly designed software.
The multi-threaded performance is much more fundamental, because it has physical limits determined by the current CMOS technology and it is also the true limit for productivity when the computer is used with really optimized software.
What I mean, is that targeting popular single-threaded benchmarks in order to promote sales can result in spending more resources and in much more complex CPU cores than in the case when the cores would be optimized only for objective criteria and they would be exploited by better software.
Depends on the workload. Some are inherently serial and cannot use many cores effectively, these will benefit from a bigger core w/ more IPC.
Does Zen4+Zen4c have a "Ryzen Thread Director" of sorts? Everything I've heard up to this point suggests it does not, relying completely on the OS scheduler to just figure it out (spoiler: it won't).
AMD might make great hardware, but they consistently seem to give nary a damn about software.
The optimal scheduler behavior is: if a task needs more performance than a Zen 4c core can provide, run it on a full-sized Zen4 core. If it doesn't need that extra performance, run it on a Zen4c core to save power.