How did AMD Ryzen get 50% faster in two years?
lemire.me
lemire.me
Zen5 is incredible when you're able to make the most of it. I’m super excited about Zen6.
Does it use vector? What can you hit with SMT disabled?
I think it’s only able to exceed 4.0/thread with SMT off because of a uops cache? From what I’ve read the Zen5 front end only had a 4-wide instruction decode per thread.
I know nothing about workload optimization, and I'd like to know more - right now I feel like the Good Burger gif. "Yeah, I know some of these words".
Increasing instructions-per-clock is all about minimising program branching (essentially 'if' statements). Because a CPU core can execute instructions faster than main memory can fetch em. It's a fun game to look at an 'if' statement and figure out how you could instead make it an arithmetic operation :).
Maximising cache hits is all about how you structure and access data. For example, if a cache entry is n bytes long, you want to ensure your struct is smaller than n bytes. Having very consistent access patterns can also help (e.g. arrays-of-structs vs structs-of-arrays).
This field is super deep, it's very fun to learn about!
So there could be a "instructions completed" counter and a "cycle" counter, and before starting the benchmark, you record the current value of both counters, then after the benchmark, record the final values, and compute the Instructions Per Cycle (IPC) as the difference in instructions completed over the difference in cycle count.
As for optimization, at a high level it's about minimizing the amount of time any part of the CPU is waiting for other parts of the CPU. The specifics require a lot of background knowledge about how modern CPUs work, more than can fit in a post, but if you're interested, topics to read about include:
### CPU cache hierarchy
CPUs store copies of data from RAM in smaller, faster memory physically closer to where the computation happens, so it's available more quickly. The CPU decides what values to store in the cache, and gives only limited control to the program, so an optimal program needs to be careful to not make the CPU make bad caching decisions (including for synchronizing the cache between multiple threads of execution).
### Instruction pipelining and out-of-order execution (a.k.a. "superscalar" execution)
Modern CPUs operate like an assembly line. A new instruction can start executing before the previous instruction(s) finishes. CPUs also have redundant hardware, so multiple instructions can be in progress at the same step of the pipeline. But there are limitations; sometimes the input to one instruction depends on the output of the previous instruction, so the whole pipeline stalls until the result is ready. An optimal program orders its operations to avoid these stalls as much as possible.
### Branch predicition
When the code to execute depends on the result of a computation, like in an `if` statement, we call it a branch. Branches can stall the pipeline, because the CPU doesn't know what instructions to execute next until the current result is ready. However, to mitigate this, modern CPUs predict which code-path will be taken when a branch is reached and begin executing the associated instructions immediately. If the prediction is right, the pipeline stall is avoided, but if it's wrong the pipeline state has to be restored to what it was before the wrong branch started executing, which is even more expensive than a stall. Usually, the CPU predicts branches correctly, so branch prediction is a net gain. Optimal programs need to understand how the CPU makes these predictions and make their branches as predictable as possible.
### Single Instruction, Multiple Data (SIMD)
Some programs do the same operations to each item of a set of data. CPUs have so-called SIMD instructions to accelerate this by, unsurprisingly, performing the same operation to multiple items at once. For example, they could add 4, 8, or even up to 64 pairs of numbers at once, depending on the instruction set and the range of the inputs. Suitable programs are optimized by arranging their inputs and operations so that SIMD instructions can be used — either directly or by being written in a way that a compiler can translate individual operations to SIMD operations.
I've said in another comment that any optimisation I'd be doing would be looking at the entire stack, and CPU/GPU optimisation would only be as a learning exercise. I'm a tiny bit familiar with caching and instruction pipelines thanks to college, and branch prediction thanks to Spectre-class bugs. I want to find some time to dive deeper, now.
[1] https://en.wikipedia.org/wiki/Sheldon_Brown_(bicycle_mechani...
I eventually gave up and turned off memory context restore and now I just deal with the minute plus (!) time to Post.
Not sure if amd memory controllers are just garbo or what but I’m going back to intel next chance I get.
Up to date on bios updates?
DDR5 has been kind of a mess imo, just in general.
I feel you though, regardless of the cause, that is an extremely frustrating place to be.
It is, but isn’t it all warranty in your scenario?
It would be painful to claim though, as what part is at fault?
There are numerous recent stories of companies refusing to make good on their warranties and instead offering customers a refund for the original price paid. One example: https://www.tomshardware.com/pc-components/hdds/toshiba-refu...
As a bonus Granite Rapids is super power efficient and runs much cooler than my previous TR and Ryzens. Intel seems to be doing good things again! My only minor complaint is that P2P doesn’t work on my multiple GPU setup because every PCIe5x16 lane has its own dedicated root to the CPU, but all that bandwidth is useful for MoE models that are offloaded to RAM.
A good alternative is ADT, they make excellent boards, there is a 4 and a 5 slot PCIe expander with a 88096 on it that works extremely well. I have two of these connected to my rig with their own power supply and four GPUs in each (and another two in the main machine). It's not exactly a portable affair (to put it mildly). Note that these won't work in a standard PC case due to the slot spacing so you'll have to rig something for that yourself (I use 2020 + some custom 3D printed fixtures).
Not sure I’ll need to upgrade my computer ever again :D
One is a tool, the other is a speculative financial asset.
So in my view, this situation is almost exactly analogous to stock market speculation.
My current plan is to stick with my existing computers for a couple more years and hope this blows over. Time will tell.
They (in theory) could padlock them so the drives and ram couldn't be stolen?
https://francisuniverse.wordpress.com/2017/10/07/the-turbo-b...
I put 256G in it wondering if I would ever need that much RAM and now it looks way too small. Highly frustrating. I wonder if the companies that made these deals realize that they've just given an entire industry a reason to look for alternatives. It's not like they wouldn't have sold their product otherwise.
The web app that's hosted on it deals with lots of images and text - over 100mil of each deduplicated, embedded, simhashed - it does the hashing, and the embedding in real time during ingest. Just handles it.
These things are absolutely insane.
That's pretty much irrelevant since the AMD's FX arch's issues weren't that SW at the time wasn't using all the 8 cores. Intel dropped the Core 2 Duo and Quad into the era where most SW was still stuck in single threaded for a long time and those CPUs still ripped single-threaded SW tasks regardless.
Here's the big reasons why the FX sucked back then and why they still suck today in the multi-thread SW era:
Instead of discrete, fully independent cores, AMD grouped processing units into "Modules" where each module contained two integer execution units, but they had to share critical resources like one FPU, the instruction fetch/decode pipeline, and the L2 cache so when both "cores" inside a module were heavily taxed especially with math or physics-heavy calculations (like in videogames), they choked fighting over shared hardware.
AMD designed Bulldozer with a very long pipeline, betting they could sacrifice efficiency per clock cycle in exchange for extraordinarily high clock speeds(a-la Intel Pentium 4) but the IPC was so bad that an FX core was often slower clock-for-clock than AMD’s previous-generation Phenom II chips and also their power consumption exploded.
FX processors were plagued by high cache latencies and an inefficient memory subsystem as another bottleneck.
So unless you're into collecting vintage CPUs as display pieces, this one definitely belongs in the e-waste pile instead of burning electricity, because it did not age like wine with the adoption of SW multi threading like people were hoping.It blows my mind that AMD watched Intel try to do basically the same thing only a few years prior with NetBurst, and fail so badly that they had to scrap that entire evolutionary branch and start over – and AMD still went and did it again themselves anyway.
That's hugely debatable and depends on SW workloads and the SMT implementation + CPU pipeline design.
In SMT the execution engines, ALUs, FPUs, and caches are completely shared. When one thread stalls waiting for RAM, the second thread sneaks into the idle execution units. At best, SMT yields a ~10% to 20% throughput boost over a single thread.
>Because that description makes the 8 core Bulldozers sound exactly like a 4 core with SMT.
It's not the same thing. Bulldozer arch sits between a true 8-core and 4-core + SMT implementation.
Exactly, which is a significant benefit for how marginal the costs are.
> It's not the same thing. Bulldozer arch sits between a true 8-core and 4-core + SMT implementation.
Then surely it should be even better than 4 cores with SMT?
If you're going to argue that the problem with Bulldozer was it's weird semi-SMT solution, you need to explain how it would've been better without it (aka as a regular quad core). Because even if it just gets the 10-20% performance improvements from being a form of SMT it would be better to have it than to not. And if you have lots of integer unit-bound threads, it should be even better than that.
I explained all the bottlenecks of the architecture in a comment above, that the issue was more than 4-core +SMT instead of true 8 cores. Please read it.
Now the very long pipeline and high memory latency are obviously significant issues with the architecture but those seem disconnected from the 4-core+SMT issue? I'm not questioning those issues at all, it's just not the part of your comment which interested me
EDIT: okay so in this comment: https://news.ycombinator.com/item?id=49809017, you explain that there's actually a fairly large part of a core that's duplicated, not just two integer units. If each "core" gets its own integer unit, register file and L1 cache, you're actually paying a ton of die space for it, unlike SMT which is "free". I can totally get how that can be a terrible trade-off for most workloads if it all ends up mostly starved due to front-end/FPU/memory throughput.
AFAIR, a bulldozer 'module' has what is exposed to a core as two CPUs, but, per everything above, is two integer cores, one shared FPU core, and depending on the version of the arch, possibly shared fetch/decode/other resources between all of that. Also AFAIR the decoder sucked as far as being able to feed both the integer cores, and the integer cores were more anemic compared to what was in, say, a K10H Phenom.
Having the same count of units (4 FPUs on the chip) did not mean having the same throughput. It's a HW bottleneck, not something AMD could fix via the OS's kernel allocation and scheduling of resources to the CPU to be able match Intel.
In strictly integer 4-8 thread benchmarks, yeah, AMD was often tied to Intel's 4C+SMT.
Bulldozer’s design didn't lose because the concept of sharing an FPU between two threads is worse than SMT. It lost because:
Intel's FPU was natively twice as wide (256-bit vs. split 128-bit).
AMD's write-through L1 cache caused catastrophic write contention in L2.
AMD's L2 and L3 caches had double to triple the access latency of Intel's.
A single shared 4-wide decoder couldn't feed an FPU and two integer units simultaneously.This is exactly the kind of thing I've seen ChatGPT do way too much FWIW, arguments silently change after push-back without acknowledgement. I think you're either a bot, or (more embarrassingly) using ChatGPT to formulate your arguments.
AMD was also having to deal with the fact GloFo split off and was relying more on general 'bulk' lithography, which kneecapped them for some time especially due to yield issues on the FX series and overall cost of that deal.
Intel also very quickly after, released Sandy Bridge and aggressively scaled it up and down; the 2500K was so cheap yet powerful I know of at least one setup that ran for a decade an only got replaced because they needed to upgrade to windows 11 for compliance-esque reasons. My own 2500K I replaced in 2017-2018-ish, only because either the motherboard took an unfortunate dive and it was easier to replace both at once.
FWIW, I did do a cheapie FX build in 2015ish for my then-girlfriend as a DVR and light gaming/emulation 'under the TV box', and it did the job well for the price, but it definitely wasn't anything amazing.
It was a tough time for AMD for sure. I think the 'split' between the Cat cores (Bobcat/Jaguar) also hurt them from a resource standpoint, although one could argue that it also kept them alive to recover (i.e. Jaguar in XBox One and PS4 being a volume contract part) [0]. They did a lot of moves that caused short term pain (that glofo spinnoff helped pay off the ATI Acquisition AFAIR) but helped them become the company that is still surviving today.
[0] - One odd side note, I still find it odd that they never did a dual channel Jaguar laptop part. I still ask whether it was because it would have made the FX look that bad...
Yeah that's hyper-threading intel was doing it as well and all modern CPUs do it as well. Where AMD dropped the ball, was they did not disclose that in their marketing as clearly as they should.
All CPUs today are marketed as x cores 2x threads, back then some AMD marketing genius in their infinite wisdom put 8 cores on the box, instead of the honest 4 cores with hyperthreading.
I had a pilledriver one, it was a perfectly good cpu, I would buy it again. If I remember back then it was the best overall performance per dollar, the alternatives if I remember correctly were i7-39.. and i7-38.. and were at best 50% more expensive for 10-15% more performance.
The Piledriver would win the consumer bang/buck mindset back then because of the 6-core part was reasonably priced and unlocked for overclocking, so people would overclock them to beat the more expensive (locked?) 4-c/8-t Intels at a lower price, but that ignored the costs of massive extra power draw(100+ W) over the Intel, the need for beefier more expensive coolers and power supplies, more expensive AMD motherboards with beefier MOSFET power delivery stages built to withstand the higher power draws of the Piledriver, so in the end the actual bang/buck gain of the AMD system wasn't remotely as big as people were making it out to be, they were just happy to get a "6-core" AMD cheaper than a 4-core Intel thinking more cores = more "better", same how having more mega-herz was also more "better" a decade before that.
The Team AMD VS Team Intel wars on forums on these topics were wild back then.
Media encoding is actually an FPU workload, and a pretty brutal one at that.
Media encoding might not use much floating point arithmetic, but it does use massive amounts of packed integer SIMD. And all SIMD instructions (both integer and floating) execute on the shared FPU, not the integer unit. It's only scalar integer instructions that execute on the integer unit.
Which leads me to believe that Bulldozer's shared FPU is not a bottleneck at all. Most evidence seems to point to the shared frontend being the primary bottleneck (which is why steamroller puts some effort into duplicating the instruction decoding, for some pretty large IPC wins)
They did just fine in parallel workloads, so I think this is not accurate. The design scaled just fine. The problem was that each core was weak.
Depends how you define "doing just fine in parallel workloads". The contemporary competition from Intel that was 4-core + SMT was beating AMD's 8-core FX CPUs in most real-world tasks and benchmarks at the time. The 8-core AMD broke even and rarely won only in >4-thread strictly integer benchmarks and some >4-thread media encoding tasks/benchmarks. So if you wanted a prosumer media encoding workstation a budget then yeah, the AMD was better, but for most real world task, it really wasn't.
>The design scaled just fine. The problem was that each core was weak.
Can you elaborate and be more exact? What you wrote is technically vague and doesn't mean anything in technical dissection/terms.
The myth that each two-core module functioned more like one core with hyperthreading would suggest that these CPUs would have much higher per-core performance when lightly loaded than when fully loaded. That is not what happened. Each core was crummy even when lightly loaded, but under full load you would have eight crummy cores, which would beat four Intel cores on a lot of workloads.
The only time contention was a serious problem was with workloads that were dominated by floating point, which were relatively rare.
By that definition it definitely was not scalable.
>But we did in fact see a roughly 8x speedup with eight threads,
Care to share a source? Because AFAIR there definitely was no 8x linear speedup with 8 threads even in benchmarks, let alone in real world use cases. The only benchmarks where those 8 threads would scale best and beat Intel were archival compression/decompression and media encoding. At everything else Intel wiped the floor with it.
>The only time contention was a serious problem was with workloads that were dominated by floating point, which were relatively rare.
Many real-world compute workloads, especially gaming related, are floating point.
This argument has never made very much sense to me. Yes, the decision to group cores into modules with some shared resources did introduce a bottleneck and result in lower performance than having isolated cores. But this didn't make the processor worse than if it didn't have the additional cores at all. These processors were at their best on highly parallel workloads. AMD shipped far more cores than Intel at the same price. The contention for the shared front end was not serious enough to make up for the core count advantage. How is the module architecture an explanation for why the design failed, when the bottleneck only becomes relevant in situations where the design is winning?
It's hard to find benchmarks from 15 years ago. But Phoronix finds what I describe: the design doesn't scale quite as well as a "true" eight-core design, but still scales better than the competition of the time which had lower core counts. Though it does seem straightforwardly bad at some workloads.
https://www.phoronix.com/review/amd_bulldozer_scaling/7
It's as FlowingRiver said. If we were to release the same parts again with modern software, Bulldozer would be a far stronger competitor to Sandy Bridge. FX aged better than Intel's designs of the same era. But Intel's designs were better to start with, so I'd still take the Core.
The reason they sucked is that they ran hot and their per-core performance was terrible.
So in a physical heavy computation task, involving double precision math (where the AMD design of share FPU units should penalize most), the FX cores where happy churning numbers with the expected speedup Vs a single core/thread version of the code. So stop saying that FX cores sucks at multithread. They fucking worked fine on that kind of tasks.
Steamroller did see 30% IPC improvements over Bulldozer (all the design work would have been done before they switched to Zen), but AMD canceled the full FX version, and only ever shipped the APU version of Steamroller (with only 2 modules, aka 4 threads).
If they had shipped a Steamroller FX cpu, the generational improvements would have looked similar to many of the generational improvements that Zen received... but didn't really matter as Bulldozer started so far behind.
This only goes back one year, to Linux 6.18. Wins would be even bigger if we go back another year. Also, this isn't tracking any of the rest of the improvements in userland: it's just the kernel. Some newer GCC, and upcoming new x86-64v3 targets will all have some pretty nice wins too.
Great days to be on open source. And it only ever gets better.
First zen 3 November 5, 2020 with desktop processors. First desktop Ryzen 9000 processors on August 8, 2024. So the generations were about 4 years apart.
>Compared this to the article "How did AMD Ryzen get 50% faster in two years?" The Zen 3 uArch being used in the article came out in 2020. So it is more like AMD got 50% faster in 5 years.
Not implying anything here, but my comment was actually before yours, this article was reposted via HNs 2nd chance pool, https://news.ycombinator.com/pool
From Geekerwan and many other Chinese sources. The M6 CPU Core is more like an M5 die shrink.
Edit: Here is the Geekerwan Video. Absolute insane improvement that a lot of current media and press seems to have missed.
It has the highest base clock of any modern AMD chip (4.3 GHz), plus 16-cores, which is already more than most workloads need.
Can’t wait to see what Zen 6 brings early next year. Rumors say 24-cores and another big jump in single-core performance.
All stated for gaming, that is.
It’s even difficult to find CPU benchmarks that don’t overemphasize 1080p and eSports scenarios.
My upgrade from 5600x3D to 9850x3D felt kind of dumb at the time, but I decided to do it because Micro Center’s bundle deals are so far below market pricing.
I was shocked at how much better it is. Benchmarks and FPS don’t really show things like micro stutters and little performance wrinkles like that. I’m not even sure 1% low FPS counts capture it.
A great example game for this is Oblivion Remastered. Upgrading my CPU alone with the same GPU took away the environment loading slowdown almost entirely.
The benchmark will tell you that I didn’t gain any FPS during gameplay but every time I open a door into the new environment my CPU is positively impacting the experience.
I would have said the exact same thing you are saying until I actually experienced upgrading to the best on the market. For the record, this is the first time in my life I’ve actually owned the best CPU on the market for gaming.
My old advice would have been to buy one of those sweet spot cheaper mid-range gaming CPUs, but my newer advice is really if you’ve already spent all that you’re willing to spend on a GPU (I have a 9070XT, my only upgrade paths are insanely expensive), buy the highest gaming CPU on the list for gaming benchmarks (e.g., I wouldn’t go crazy with a 9950X3D2 since it doesn’t have any gaming improvements above the 9850X3D).
And the thing about CPUs is they’re not insanely expensive like GPUs. We are talking a price delta of $200 between this beast of a CPU and something way more middling.
Given the massive increase in hardware prices lately, we are going to have to get by with mid range or even low end hardware for a lot longer. Studios will be forced to actually optimise games to not run like shit on a sub $8000 PC.
Every game should be targeting the switch 2 and steam deck in terms of power.
I think PS5 is a perfectly reasonable spec to be targeting, which was the higher end of mid range about 5 years ago, and still significantly out performs Switch 2 (which significantly out performs steam deck). I’m sure the GTA sales will demonstrate that the install base at this spec is wide enough for broad commercial success.
I grew up with gaming systems that got rapidly better and while it can be argued that we are at a diminishing returns plateau of gaming horsepower, I don’t think the solution is to never get hardware upgrades ever again, even in the current RAM-constrained environment.
GTA VI is coming out with FPS maximum of 30FPS on the PS5 including PS5 Pro. That’s a subpar experience that is limited by hardware.
We saw with much of the PS5’s lifespan that many of its cross-platform games came out on PS4 and original Switch with minimal compromise.
Not enough of the install base owns those higher end solutions, so for the most part someone with a better performing rig is getting better FPS, higher resolution, and not much else. Even gaining additional draw distance is rare these days, and higher resolution textures are basically impossible to see.
If the console is going to survive rather than decline as a whole I think that Microsoft and Sony will need to get more serious about innovation rather than just making a samey spec bump at a high price that doesn’t motivate their users to upgrade. Nintendo already proved that basic and rather mild level of innovation can excite the market (original Switch).
It's relatively easy to allow powerful hardware to just run the game at 4k 120fps highest settings while weaker hardware runs exactly the same game at lower graphic settings.
I think it's limited by their software. The hardware buys them flexibility. Nobody is pushing it like the ps2 days.
Having a powerful rig is the brute force way around the game’s performance problems.
My PC wasn’t $8000. My 9850X3D, 32GB of RAM, and motherboard cost $800 in a bundle. Yes, purchased in 2026 during the RAM crisis. I was also able to sell my previous CPU, RAM, and board for around $500.
Obviously this is not a budget build but about half the cost is the graphics card. I would have saved minimal amount of money if I had gotten a worse CPU.
Does the base clock even mean anything beyond something that may or may not vaguely indicate the performance? Its lower than what these chips can sustain with sufficient cooling from my experience. It's higher than what the chips will clock down to when not sufficiently cooled.
They aren't the maximum or minimum but I think they are the only clocks with the specific guarantee "if you match X, you must be able to get Y". The max boost clocks usually aren't too hard to reach if you're trying though, there is just no guarantee you will in all workloads or how much it takes to get there.
Let's not discuss battery life and overall size, however it helps the thermals.
I do a lot of CPU bound engineering and scientific compute on the computer and it chews through it.
I also bought the entire computer for $2500+tax new from Microcenter in December and it's appreciated 25% in value since then.
I recently upgraded from that to a Core Ultra 7 that I got from work (e-waste recycling). It was not my intended upgrade path, and I hope Lisa isn't too mad lol (I kept my Radeon though).
How the hell does a high end CPU that's barely 2 years old end up in e-waste recycling? Was Richie Rich using them or something?
Sort of: there were about half a dozen in micro PCs in a load picked up from a hospital. I was kinda bummed that they were BIOS password locked (as were the other PCs from there). The model didn't have a password reset jumper (or any reset mechanism outside of contact Dell support with proof of purchase), so I couldn't sell them as whole systems, and stripped them for parts.
We get DDR5-based systems from time to time, but they're understandably rare. Most of the time, they still work. My rule: any day I come across DDR5 is a good day.
Are you by any chance in some super-rich part of the US? Because then no wonder your healthcare is so expensive when hospitals treat high end PCs as single use disposables. Here in Europe I sometimes see hospitals and doctors practices using 10+ year old PCs and still rocking the old 17-19 inch 1280x1024 CCFL LCD monitors from like the mid-2000s.
Regardless, I'd love to just run into high end PCs being thrown away here and pick them up for free, but where I live I see people barely throwing away their value-line Dell/HP Core 2 Duo towers with audacity to ask for 25+ Euros for their e-waste with the classic "no lowballs, I know what I got" attitude.
I don't understand it either. Two scenarios make sense to me:
Somewhat probable: everyone in a department got upgraded, no exceptions. 'This is a really nice PC, but the boss says it has to go.' (There were about 100 i5 9th gen mini PCs (still OK), and 500 t640 thin clients in that same load.)
Less probable: some department got closed and cleaned out.
> Are you in some rich part of the US?
Pittsburgh, Pennsylvania. Not particularly rich (rust belt), but healthcare is a significant part of the local economy. About 90% of healthcare here seems to run on HP, but Dells can appear.
> Here in Europe I sometimes see hospitals and doctors practices using 10+ year old PCs and still rocking the old 17-19 inch 1280x1024 CCFL LCD monitors from like the mid 2000s.
My stomach turns seeing the dozens of 24" probably 1080p monitors we scrap every day. Pretty much all are significantly scratched when they come off the truck, otherwise reselling them might be good business.
Here in Europe all that is 100% resold either locally, or further to balkans/eastern-europe due to much lower purchasing power making the resale efforts of ~5 year old HW economically viable. I check the local "craigslist" often and I notice when a business is clearing house due to a flood of used HPs/Dells/Lenovos from a single user. However the list prices aren't remotely palpable to be good deals. I assume the person/business hired for clearing the house is marking them up significantly hoping to squeeze a big win from a mark.
Some of the tech companies I worked for here would first auction their older HW internally at every upgrade cycle to the employees before selling what was left to a clearing company, which makes me feel we're being robbed to be offered by employers to bid for their e-waste when in the US people find better stuff in the trash for free. Europoor indeed.
>otherwise reselling them might be good business
I assume those businesses don't bother reselling 1080p monitors and 2 year old PCs, and instead just throw them away, because the price of labor is so high in those parts of the US, that paying a full-time employee to take care of such resale tasks will cost them more money than they expect to make from the sale, so it's assumed to be cheaper and less hassle to just throw everything in the trash instead of wasting time on the used market dealing with tire-kickers just to gain what is essentially peanuts money for those businesses' bottom line.
Am I close with my assessment?
I'm not sure how prevalent this practice is, but IT equipment is sometimes on a depreciation schedule. If an employee starts selling the company's PCs and monitors, an accountant will get very angry, because he told the government that the stuff is literally worthless, but it apparently isn't because it was sold for something. IANAL, but I'm fairly sure that's tax fraud. Once it's turned over to someone else (my company) for free, the valuation resets, or something.
None of what you're describing is available here at all. I sometimes see Europeans willing to ship here (already very few) offer heavily used and scratched ThinkPads and Dells for money they weren't worth when they were new. Add at least $80 on top for shipping, and no warranty, you get what you get.
On the local market, you will find those Core 2 Duo with precisely the attitude you're describing, but at 2-4 times the cost.
So I buy everything new, which is expensive with our salaries, but we have no choice. At least new hardware is available now thanks to selling over the internet and improved logistics.
10 years ago you could only buy what was offered by local shops, which was always at least two generations behind, and at least twice as expensive as it was in the EU (forget the US).
For example, if things continued like this to this day, I'd estimate the newest CPU you could buy right now would be something like Ryzen 3600, at twice the cost it was in your country when it just came out. If we're lucky, maybe 7600 would already appear at a similar overprice.
I'm sure many parts of the planet are even in a worse position than we are.
I fully believe you, but I want to _understand_ you.
My company's policy is that retired laptops are destroyed. (I suppose properly wiping them or removing storage devices should be enough but for some reason it's not considered sufficient.)
There are indeed lots of second-hand former company laptops being sold, and some companies even specialize in refurbishing and selling them. I don't know what percentage ends up being recycled but it's definitely not 100 %. I'd assume the more security-oriented the organization gets, the more likely it might be to avoid recycling devices intact.
With that said, I haven't seen two-year-old high-end devices getting retired. But if there are only a few of them, and a large fleet is being renewed, I suppose it can make sense from a fleet management point of view to renew the few perfectly current ones as well.
I suspect that had something to do with it. The Core Ultras came out of Dells, and as mentioned, the industry around here uses HP almost exclusively. (The rest of that load was HP, only about 1% were those Dells.) Some dogmatic system administrator probably said "We we are only supposed to use HP, get those Dells out of here!"
They probably got some new garbage software where because of modern terrible programming practices, the requirements are very high.
Thus they need to replace the "old" hardware.
DDR5 is literally twice the price per GB. It's better for sure, higher bandwidth but waaaay more expensive. The 9800X3D supports much faster and even more expensive DDR5 than the 7800X3D too (DDR5-2667 to DDR5-5600).
Even really good DDR5 could be had for cheap; I remember paying like 300$ for 96GB of high-quality M-die hynix DDR5.
Edit: dug in it, no, because cache isn't addressable and it's directly managed by the cpu.
Zen 3: 2020 release
Zen 4: 2022 release
Zen 5: 2024 release
The author is disingenuous by looking at only the X3D variants. The Zen 3 X3D variant came out late.
For comparison, in the same time frame, the M4 is 52% faster in ST and 72% faster in MT than M1.[0] If you go by what you can physically buy in stores since Zen 3 was released to now, Apple's chips have gotten 83% faster in ST and 143% faster in MT.[1]
[0]https://browser.geekbench.com/v7/cpu/compare/433673?baseline...
[1]https://browser.geekbench.com/v7/cpu/compare/435139?baseline...
To my eyes, your message looks kinda like, "The term 'bicycle' doesn't cover it anymore, things have changed so much since they were introduced. We should call them something else, maybe 'bicycle pro'? 'super bicycle'?" No it's just a bike, and no, modern CPUs are just CPUs which look scalar at the ISA level but have ILP
> The number of transistors is way up, by about 50%, from roughly 11 billion to 16 billion. Most of the extra transistors went into the core, not the cache.
From https://en.wikipedia.org/wiki/Moore%27s_law
> Moore's law is the observation that the number of transistors in an integrated circuit (IC) doubles about every two years, with minimal increase in cost.
And while in the past, the doubling of transistors DID double performance, it's no longer true.
Better thermals maybe? Less throttling?
I’d also be weary of single core scores,Geekbench is known to be favoring specialized instruction sets like AVX or encryption extensions more and more as the version number progresses. Though I don’t know if GB6 is also like this I wouldn’t be surprised it it was caused by it.
One of the biggest improvements to the Ryzen 7 X3D series (e.g. 9800) is that the caches have been moved from one side of the die (to the other), which places the major heat source closer to the heat sinks.
SO yes, less heat throttling.
Power is heat, and the amount dissipated is the square of the voltage- a processor that is reliable at less voltage means you can get a lot more frequency in the same heat envelope.
Yet reliability decreases as frequency goes up, and that can only be stabilized by adding more voltage- so the faster you run the processor, the more voltage you ultimately have to give it, so the power/heat produced grows exponentially until you can't get rid of the heat fast enough (at which point your only option is to actively cool the chip).
This is overclocking 101.
Note that classic overclocking was viable because of arbitrage- buying a processor, pushing it to the point it got too hot, and stress-testing it at that temperature to ensure reliable operation. Processor manufacturers all do that formal verification at the factory now as they'd be uncompetitive otherwise, especially in laptops.
still on 5800x (non-x3d)
My hope is on Huawei, which has recently developed a stacked chip architecture that reminds me of HBM, and I'm wondering if that's step 1 of China entering HBM production.
I don't know enough about hardware, so I might be saying absolute nonsense, happy to be corrected so I can learn where I'm wrong
And it shrank.
2020 Zen 3 AMD Ryzen 7 5800 - 1852
2024 Zen 5 AMD Ryzen 9600X. - 2914
2027? Zen 6 - Est 3200?
And compare this to Apple.
M6 4000+
M5 3600+
M4 3300+
Even the Snapdragon X2 Elite X2E is better at 3300.
The gap isn't exactly shrinking. And I would love x86 to prove me wrong. But even a hypothetical Zen 7 in 2028 ( 2029 for consumer ) may only catch up to Apple's M5 released in 2025.
Although, if you count EPYC, which you might as well with the prices Apple charges for their top-end models, AMD still beats Apple with ease at multi-processing.
It's a real shame that these processors are made by Apple (and Qualcomm, I suppose). They'd be my go-to chip if they weren't.
Apple has been iterating on a yearly basis, A17, A18, A19 and A20 are all different core design. Some of them just so happens to come with new node. On the other hand Zen 5 to Zen 6 took 24 to 30 months.
I mean, RTX 5080 has nearly 1TB/s, 5090 has nearly 2TB/s. Maybe you are talking about CPUs / unified memory platforms? I agree nobody else does it better. But for LLMs, GPUs can still be significantly faster than even the most advanced Apple silicon on the planet. TTFT in particular is super inferior with Apple, for now.
That's probably also the reason Apple had to reluctantly give into Nvidia servers for the initial rollout of Siri AI, though they claim to use trusted computing extensions to reach an acceptable level of privacy. (I do not trust that nearly as much as the Apple Silicon nodes)
It'll be amazing five years or whatever down the line to see Apple reaching those figures. They seem to be heading in that direction lately.
The x86 folks lost the ball completely. For workstation/ mobile needs you go to arm. For matmul at scale you go to gpus.
I guess windows gaming with separate gpu is the only remaining market for them.
Unless Nvidia releases a motherboard with an arm cpu on top of their RTX cards, and then that market is gone too.
There is no chance the sustained single core perf is better than a 150=200+w ryzen with a big noctua cooler on it. The power draw alone would kill the battery in a few minutes(?), and the heat would make it catch fire. lol
A20 Pro scores around 4725 at 8.9w. Geekerwan's review showed that the iPhone 18 pro could dissipate as much as 6.4w during long gaming tests. Cutting power by 50% likely still keeps around 80% of the clockspeed.
As 9950x scores around 3400-3450 in geekbench, or around 28% slower.
There's a very good chance that the iPhone single thread performance is genuinely faster even during sustained loads.
So whatever your argument is, it doesn't make much sense.
The M6 is based on A20 Pro and it sustains its performance in a Mac Mini indefinitely - likely longer than Ryzens.
>Single core performance still worse than an iPhone.
I can believe the core with active cooling can sustain good performance at least, but not in the IPhone.