AMD Zen 3 Ryzen Deep Dive Review
anandtech.com
anandtech.com
It's really impressive to me how big of an improvement AMD has made under the exact same constraints as last generation. Same process (even down to the same PDK) though generally higher yields probably let them choose higher bins. Same exact chiplet size to mount to the same substrate. Same power availability. Same IO die with same interfaces and memory controller. Yet they achieve +19% IPC and higher frequency with just design changes. I wish there were more detailed information available about what day-to-day engineering work goes into these design changes. Speculating:
- Designing better workload and electrical simulations to make better decisions when evaluating changes. - Running lots of simulations to determine the optimal size and arrangement of each cache. - Improving CPU binning processes and data, improving on-die sensors, and improving boost algorithms to reliably run closer to the edge. - Tweaking parameters of automated layout algorithms to reduce area of specific functional units. - Improving the algorithms implemented in logic for various processes.
And the Linux performance is looking great too: https://www.phoronix.com/scan.php?page=article&item=ryzen-59...
The difference between TDP and Peak Power for AMD is around 35W, while for Intel it is up to 140W.
When choosing a PSU that difference can have a real impact on system stability.
The 5600X looks fantastic at 65WTDP. Add the unavoidable price drop of Zen 2(+) on top of that. Buyers remorse is real.
I bought a Ryzen 7 2700X a few months before the release of the Zen 2 / 3700X. And yeah, it's faster for nearly the same price (though due to timing I paid ~$230 so it's not insignificant when compared to the launch day pricing.)
There's "always" more, more, more performance for the dollar in tech, and there is no perfect price/performance product. The Ryzen 5 3600 was and still is a great price/performance product and that doesn't change because a more expensive, higher performing chip is released.
You can make decisions at a point in time, or if time is irrelevant, you can keep waiting for new milestones... and then make a decision in that point of time instead.
For me, I waited nearly 10 years to replace a 3+ Ghz quad-core CPU, so the 4+ Ghz octa-core was a nice upgrade (albeit still not earth shattering) and I am very glad to have it. I love the recent increase in performance growth over time, but it will likely be quite a few years before I upgrade my desktop again. Of course, I'm just one guy, and I'm glad many enthusiasts and computer-using professionals will benefit from all this advancement!
You'd end up with more 'residual' computing power (when your computation is done, you have the hardware lying around, and now it's more powerful than if you had bought on day 0), and it would probably cost less in energy as well.
I thought those types of applications would actually be able to take good advantage of all those cores.
SysInternal's https://docs.microsoft.com/en-us/sysinternals/downloads/core... actually captures this detail (probably there are better apps, but it was nice to inspect machines quickly). The tool measure some relative slowdown.
The same with the upcoming Macbooks set to be demoed on the 10th. If I want one, I'll wait for at least second gen.
Regarding AMD CPUs this is an unusual situation, because 3rd gen or 4th gen may be optimal, which is highly uncommon. This is because next gen will run on faster ram. The bottleneck on the Zen 2 processors was cache delays. The bottleneck on the Zen 3 processors is ram delay. This shows to me there is more improvements to come. However, unfortunately, this also means cost will go up quite a bit in newer generations with smaller gains making it not as worth it. I don't game much and I do data science work, so processing is done in the cloud, so I have virtually no reason to upgrade my nearly 10 year old CPU as odd as that may sound.
Zen and Zen+ did NOT have an I/O die yet. Zen2 added the I/O die, and now Zen3 improves the compute die (without changing the I/O die).
My general preference: buy Tock generation (but Intel recently no Ticks for desktop, AMD does Tick+Tock on Zen2)
Most of the time I would advise against waiting but right now there's a clear advantage to being on the other side of this shift.
Just waiting on the radeon reviews before I decide on a new GPU to get into 1440p gaming.
I will be going for 3900X for my next machine. Probably waiting for the RTX 3080 Ti I now see being leaked, or a regular 3080 (not interested in the 3090).
Amusingly, my thoughts of upgrading were initially triggered by wanting new storage. I currently have the original 250GB SATA SSD I put in originally, along with a handful of hard drives from various older systems. Motherboard doesn't have an M.2 slot and I was thinking, hmmm...I bet there are a lot of things I can't upgrade at this point.
Waiting on these new Zen 3 CPUs to be available, but not in a huge rush since I also can't get my intended 3080, so we will see which comes into stock first. It will be a nice bump to a new gen CPU, DDR4, and a few gens of GPU once I can actually buy the dang parts. If they were around I would've built the thing a month ago.
Unfortunately when I switched from the GTX 980 to the RTX 2060 Super all PCIE 3.0 lanes were being used (ASRock Z77 Extreme 6) so had to put the riser card into a PCIE 2.0 slots which means I am not getting the full performance.
I mean, the 5600X is about double the price of the 3600, depending on when you got the latter.
I keep my desktop for a long time, so it seems to be "okay" to do. For reference, upgrading from launch year 2500K. But still..
Would a 3700X have been even nicer? Probably, but when I bought my new CPU last year it was almost twice the price for two extra cores. So I do feel your temptation, but I don't think you would regret the 3700X.
Oh well...
> It's really impressive to me how big of an improvement AMD has made under the exact same constraints as last generation.
It's 8-core chiplets now, but yeah same die size according to marketing.This image shows the change more clearly: https://images.anandtech.com/doci/16214/Zen3_arch_19.jpg
You can see it's the same overall amount of "stuff" in the chiplet. It went from 8c / 32MB L3 per chiplet to... 8c / 32MB L3 per chiplet. But it's now not divided into 2 smaller chunks as it was.
Make sure to read the subreddit, if it happens to be that one. The CPU is fine, but dGPU integration is, uh, interesting.
The biggest issue I've had is with making sure it turns off, and /r/zephyrusg14 is somewhat helpful there. On Linux it's more straightforward, but currently needs a few kernel patches.
X-Plane is indeed super scriptable, which is used by various IHV’s to auto test new drivers and hardware. We also use it extensively internally in our CI testing suite.
Edit: To the OP, I can provide performance benchmarks from an i7-6700k vs Ryzen 3900X for your Dad. It’s unfortunately not super apples for apples. What I can say is that the massive L3 cache is doing wonders for X-Plane, and the greater core count is definitely going to be more relevant in the future as well.
FS2020 is pretty - lots of eye candy, but the flight model simulation isn't as realistic as X-Plane, which historically has put correctness above eye candy. FS2020 is also very buggy still (including crash to desktop after restarting a flight, etc), although getting better.
That, and there are not many plugins, add-ons and aircraft models available yet for FS2020, although it's growing. The default built-in aircraft have many missing features, buttons, autopilot modes, etc - which prevent folks from using these as "study level" simulations.
As of right now, serious simmers still use X-Plane/Prepar3d, but maybe that changes as updates roll out for FS2020.
We recently completed our move away from OpenGL, which served as the base graphics library for a good 20 or so years. It took forever, partially due to how deeply ingrained GL-isms were in the code base. But we are now all fancy and modern with Vulkan and Metal, which allows us to do some really cool and long overdue things with the sim that were previously almost if not completely possible. So that's quite exciting for me, not to mention that it got quite a bit faster in the process. I can't speak for the other parts of the sim since I almost exclusively work on the graphics and game engine bits of X-Plane, but I think most everyone in the company is pretty excited for the future.
Also, and this is purely from a flight sim user perspective, competition is always good. Just like it's good if AMD and Intel compete with each other. The only winner is the end user.
Best of luck developing X-Plane, I wish I could help you guys justify the Linux support, but unfortunately I've not got the enthusiasm or free time to commit to a flight sim just now.
Side question, will the large cache on RDNA 2 close the gap with Nvidia?
The problem is that the AMD driver has a bug with scheduling mixed workloads. Sometimes the GL side gets to execute first, sometimes follow up Vulkan work executes first. This leads to incorrectly composited rendering and really annoying flickering. And despite trying to have this fixed for over a year and many many emails with AMD engineers, we are nowhere near close to a fix for this. We are at the point where we are even considering integrating Zink just so we can ship an enjoyable add on experience to users.
Not every user has that problem, it depends on the hardware and software that runs (ie a very classic race condition). But I can't in good conscience recommend AMD over Nvidia here. And I hate that, because the Big Navi cards look amazing.
In short, even a 5600X will exceed an Intel 10900k.
Will Intel bring out a new architecture that will rise above all this as they did with the Core2Duo (actually typing this upon one of those cores), well - they need to.
But for me the shift on CPU's does seem to be large and faster cache access and latency. Say Intel released the same chips but added a very large SRAM layer - that for some applications would swing things for sure, not cheap but again - Intel controls and runs their own fabs, and whilst AMD could equally just do that, would the costs in them doing that with TSMC be just as comparable? There is money to be made in fabricating or TSMC would not be doing so well financially of all this. But that is a cost/income that Intel are tapped into with their own fabs.
However it pans, competition has been good for the desktop x86 folks and so glad AMD back in the game and going strong.
I imagine they will do that on the GPU side as well at some point.
The ones that don't have full 8 (well enough) working cores just use them for 5600X (6 core) or 5900X (12 cores) if there is 6 or 7 good cores.
The fully working chiplets you can bin for 5800X or 5950X (or for threadripper/Epyc for the very best ones)
Which they did for Broadwell but then immediately abandoned it. Anandtech recently did a retrospective on it: https://www.anandtech.com/show/16195/a-broadwell-retrospecti...
(although at the time of writing this Anandtech's site is dead so pop that into your cached lookup of choice)
(I should have said on package, not on-chiplet, too late to edit)
Of course, these are enormous and ridiculously expensive dies.
So, if you can use ~6x the area of eDRAM, you can have almost no power draw SRAM.
https://ark.intel.com/content/www/us/en/ark/products/190887/...
The other aspect is that they will act as a heat spreader at the die level, so whilst not best use of space - given all the factors - not something could easily dismiss and better a waste of space than a waste of a chip who's only crime was to have a non-functioning GPU upon it.
10nm was / is a disaster failure and even if could somehow fenagle it into a usable state by next year will be trounced by tsmc 5nm zen 4. They need to cut their losses and leapfrog or they will stay technologically irrelevant going forward.
Afaik he is not a chip architect anymore - more of organizing engineering teams to perform to their best ability and asking the right questions.
The AMD guy who designed Zen architecture was also responsible for the ill-fated Bulldozer design.
Those scare quotes are kinda rude, and whatever you're trying to insinuate with them is wrong.
Clearly a lame PR explanation considering his high profile stature at Intel (he is going to save us) and the suddenness of it all.
https://www.extremetech.com/extreme/311664-intels-jim-keller...
I see they used single quotes .. my bad.
You still haven’t explained what was so suspicious about him leaving.
I think the implication in the parent comment is that things were so bad at intel that he couldn't fix and chose to leave instead. Whether that's true or not I don't know. But the rumours I've read are not good.
Intel responded later with the Core 2 Duo, which not only met AMD on its playing field but beat it. For anyone who is wondering how Intel did this, here is an in depth article on the topic: https://www.anandtech.com/print/10525/ten-year-anniversary-o...
The CD burner issues were because they required hard real time control to not screw the disc up.
Jim Keller was certainly key in getting Zen up & going, but he left in 2015 so his impact on Zen 3 would be rather minimal to say the least.
Combine that with the timing of Intel struggling hard with new process nodes and there you go. A bit of AMD got better, a bit of Intel got worse, and now you've got this landslide result.
AMD's execution looks meticulous by comparison, and they deserve all the praise and profits resulting from it. But their win is inflated by Intel completely imploding. Until next year they're merely iterating on 2015 technology - their Skylake architecture - on the desktop and in servers at least. This is not to say that Intel's 2021 chips will be competitive with this (they won't be) but if Intel with better management would have been at this point in, say, 2018 instead of 2021 this would have played out very differently.
So before they faceplanted with 10nm Intel first stumbled with 14nm. 2 generations in a row of issues.
Seems like a classic textbook case of a large coorporation not understanding how to change their ways, ending up being beaten by a more agile smaller competitor. Oh well, hindsight is 20/20.
Highly doubt that he did so in a technical role, maybe in a PM like role, but not a design lead. The team lead was Suzanne Plummer, architect Michael Clark.
I was saying many times, if you hire a JK level cadre, it's best to not spend his very expensive time just doing technical work.
Although AMDs and Dr. Lisa Su's achievement are not to be underestimated, I doubt the 7nm processes would be as mature as it is right now without Apple.
Apple's extra cash certainly didn't hurt, but TSMC wasn't struggling before Apple came along, either.
And prior to 7nm the other major fab, Global Foundries, as perfectly competitive with TSMC despite not having any Apple money.
Clearly TSMC is heavily diversified and gets to pick their clients. Apple has also been the primary (volume) launch partner on both 7nm and 5nm, which to me indicates how much TSMC values that partnership. Imagine the slam dunk Nvidia would have had if the RTX 3000 series was on 5nm.
They wanted more margins? But note that Nvidia does still use TSMC's 7nm for their largest & most expensive dies. The A100 is TSMC 7nm at a staggering 54 billion transistors on 826mm² of silicon. Nvidia retains the largest die manufactured on TSMC's 7nm. By a lot. The next largest would I think be Navi 21 at 536 mm² and 26.8 billion transistors.
> Apple has also been the primary (volume) launch partner on both 7nm and 5nm
Apple's die sizes & transistor counts are comparatively tiny. The first use of a new fab is pretty much always a small, low-power die. That's what yields best when yields are the lowest.
See also why Snapdragons are also among the first to launch on a new TSMC node, despite those SoCs having probably the lowest margins of the high-volume stuff TSMC makes.
Apple could have made the A14 much bigger and faster than the A13, but they chose to cram more chips on the wafer, which I speculate they did to have more wafers available for their bigger upcoming iPad and Mac chips.
I think you have it backwards on margins. Per square mm it’s clearly more expensive to produce an A100 die than an A13 or Snapdragon SoC due to the wasted die space and need for golden samples on the Nvidia side of things where they’re not selling any cut down dies in GeForce cards.
TSMC sells wafers, not functioning dies. They don't care how you slice it up or how bad you yield as a result.
It’s rumoured that Nvidia’s deal with Samsung was for working dies, not total wafers.
Let's just go back to 2014/15. And Look at Intel's Roadmap ( Or Slides ).
2016 - 10nm - CannonLake - First Gen 10nm Process. Industry leading Density.
2017 - 10nm+ - Icelake - New Architecture.
2018 - 10nm++ - TigerLake - Optimisation[1].
2019 - 7nm - Sapphire Rapids [2]
Most people would expect the ETA year to be +1. As 14nm itself was delayed.
TigerLake has around 15%+ IPC. That would put it ahead of Zen 3. The Intel 7nm is roughly between TSMC 5nm and 3nm. Had Intel shipped that in 19/20, along with Sapphire Rapid. AMD wouldn't even lead, they could barely compete.
So to answer your question. It wasn't AMD did some magic. They have been executing, and shipping their product on schedule. ( Although in tech, that in itself is pretty damn impressive ). All while Intel was.... um..... I dont know what they were doing.
10m fiasco has lots of other knock on effect as well. But I guess this is good enough for a ELI5 post :)
[1] This Process - uArch - Optimisation name came in later part. In 2014/15 people were still talking about Tick Tock.
[2] Sapphire Rapid was always a Server Specific design on the roadmap. But things changes bit. Along with Intel's plan of AVX512
Intel's next generation cpu node was supposed to come out this year, and it's been pushed back (possibly due to covid) until the end of 2021 or even into early 2022. Meanwhile, AMD says they have 5nm around the corner.
Unlike Intel who have micro improvements, each generation has been a big leap in performance so it’s been exciting for consumers.
Intel did something similar with their core2duos when AMD was on top around 12-13 years ago.
Those processors were such a leap over the Athlon XP 3200+ I had before and it runs a Linux desktop perfectly fine.
What are you using it for? Floppy disks were already very dead by 2008, let alone 2020.
I know there are alternatives to using a floppy to transfer to these ancient systems but for me it works okay.
For any practical purposes, Intel is now behind.
His previous stint on AMD was when he invented AMD64.
I think the 2 key high level choices AMD made was betting on TSMC, and betting on the chiplet + memory controller design.
We only can guess what they can be?
I suspect he would have left no matter what though, because Intel is so arrogant they wouldn't listen to him in the manner in which they would to.
That was the great thing about AMD; when you're already at rock bottom, you're humbled enough to listen to people.
AMD either did well or went broke which forces management to take painful but important decisions.
source?
A big part of his success story, was him having this great opportunity to develop his skills.
Microelectronics is a career suicide in comparison to software today.
I'd say 100 to 1 sieve off rate for people getting into the industry, to ones who actually get to do serious architecture work would be too small. The ratio is likely to be much larger.
They have decades of X86 experience(they invented it) and literally more money than God(they probably spend more on office supplies than AMD on R&D).
How do you have everything going for you and still lose?
Sure, but AMD has decades of X86 experience, too. They were, after all, the second source supplier of the the Intel 8286, the AMD version being called the Am286. And AMD did, after all, invent the x86_64 ISA that we're all now actually using, while Intel was over fucking around with the utter failure that was Itanium.
They then bought a new chip startup called NexGen which became the K6/K7 later the Athlon.
AMD benefited by TSMC basically already investing in smaller nodes for Apple, combined with new manufacturing innovations caused by their innovative chiplet/infinity fabric designs. This enabled them to glue more cores together at higher yields and therefore winning massive benefits in multicore performance.
Intel had to worry about constraints caused by both the architecture combo-ed with manufacturing. I'm also not sure if they were even looking at something like chiplets at the time.
Complacency
End users lamented the lack of real improvements, but Intel continued to sit on their laurels because it was the easy thing to do - they simply didn't need to do anything drastic. Which it pretty sad - with different leadership with technical vision, who knows where computing power would be today? But that's business...
Intel took it too far though, and didn't seem to see AMD catching them up in their rearview mirror. IMO it serves them right, and I'm so happy there is a underdog challenger that is outpacing them in so many areas. The desktop I bought last year has an AMD processor - the first in my household in many years, but hopefully not the last.
And AMD invented AMD64 ;)
This is starting to become a problem. It's happening to the new Nvidia GPUs, and the new AMD CPUs, and it will probably happen to the new Radeons later this month. At this point it will probably be impossible to purchase any of these products at retail before January.
I'm not sure what Nvidia, AMD, and their retailers can do about this situation, but it's pretty bad.
This aligns with reports of retailer inventory, where availability of high end models was extremely limited, but there were many more of the mid-tier models.
I could easily have seen the bidding wars on the first supply of 3000 series GPUs having people buying them for $1500+ easily with the 3070 selling for $1k at least.
For the real high-value users they have completely different products for sale to capture that margin at even higher prices (Epyc, Tesla).
Also, most of the high bidding on the 3000 series were anti-scalper bots putting in bids they had no intention of honoring.
<<<
PDEP/PEXT Parallel Bits Deposit/Extreact 300 cycle latency 250 per clock 3 cycle latency 1 per clock
It’s worth highlighting those last two commands. Software that helps the prefetchers, due to how AMD has arranged the branch predictors, can now process three prefetch commands per cycle. The other element is the introduction of a hardware accelerator with parallel bits: latency is reduced 99% and throughput is up 250x.
>>>
It talks about pdep/pext and links them to prefetchers - what? Weirder yet "is the introduction of a hardware accelerator with parallel bits" all that's happened, and backed up by earlier info in the article, is that these were trapped+emulated before, hence their hideous cost, and now they're in hardware[0]. It comes across as if someone didn't understand that.
[0] calling it a 'hardware accelerator' is just peculiar.
Certainly AMD done well with the CPU cache changes and looking at their GPU, they leveridged that appraoch again very well it seems.
If Intel was to release the same chips and lob on a large SRAM buffer - things may flip again.
However it pans out - this real and credible competition is good for innovation and many area's of x86 have stagnated for so long that the whole Ryzen momentum has done wonders and totally flipped the positions of Intel and AMD.
If you're compiling really large projects and want to cut those times down it could be worth it though.
https://videocardz.com/95980/amd-ryzen-5000-vermeer-zen3-rev...
The idea is to increase yields by putting the bigger chips in the middle, and smaller chips around the edges—perhaps for different customers entirely…
And "yield" in this case is not how many chips can be placed on a wafer, but how many of them will not have defects in critical areas.
TSMC themselves don't care how many chips gets made from a single wafer.
Does AMD advertised this change as a "Eco-friendly" like Apple? It's reasonable for 8 cores or upper SKUs.
They also save a couple of bucks by doing this.
Zen 3, of course, will be different again.
2. Higher clock rates (exceeding the advertised frequency this time)
3. Power consumption basically the same
4. Lower latency for most operations (notably 5 -> 4 cycle FMA)
5. Same process node, same die size
6. Slightly higher price
This generation they've also moved to simplify further. The mobile chips with the same number were always one gen behind the desktop (mobile 4xxx ~= desktop 3xxx), but now they've skipped desktop 4000 both are back in sync.
Also Intel's naming scheme is getting complicated too. Base vs K vs KF Vs F. 10xxxgY vs 10xxxH/U on mobile are two entirely seperate tenth gens with lots of overlap.
I don't know what AMD embedded is compare to r3
Comparing NUCs, which is better: the BXNUC10I7FNK1, or BXNUC10I3FNK1?
I've got a i5 3350P. Quick, what does the P stand for?
edit: it's certainly fair to criticize the manufacturers for having bad sku names, but it really doesn't matter that much in the big picture. most people that buy computers neither know nor care about the specific CPU. especially with laptops, the decision is usually made for you by the OEM and you decide based on what size SSD and display you want. if you do actually care what part you get and are in a position to choose, you should make your decision based on benchmarks, not a sku name.
10900K, 10900KS, 10900 vs. 5600X, 5800X, 5900X, 5950X
I see where confusion could creep in.
(i|Ryzen )[3579] (1 or 2 digits to indicate generation)(3 digits for model)(optional suffixes)
Of course, since Intel is pretty much stalled for the last few years, it's a little easier now.
I feel like I learn them during the purchasing/comparison period then never need to look again unless I'm repeating that cycle.
Not really. Look at core count at the old I7 vs. the new. If anything Intels naming is even worse.