The Great CPU Stagnation
databasearchitects.blogspot.com
databasearchitects.blogspot.com
During that era for the most part Intel's i7 prosumer CPUs started with 4 cores with the Bloomfield Nehalem chips in 2008 (which at the time were awesome and a game changer) and ended with 4 cores with the Kaby Lake-S in 2017. It really only changed in 2017 with AMD Ryzen forcing Intel to actually increase core count.
2008 Nehalem benchmark: https://cpu.userbenchmark.com/SpeedTest/778/IntelR-CoreTM-i7...
2017 Kaby Lake-S benchmark: https://cpu.userbenchmark.com/Intel-Core-i7-7700/Rating/3887
When I compare the two, it shows an effective 20% speed increase, although microbenchmarks show a 50% increase. That is a stagnation.
During that era it felt like a lost decade. I don't miss it.
Not because anything I was currently doing with my computer was becoming too slow, but because I wanted to do new things (VR). A shame, I wanted to run that thing into the ground.
I've got my old system sat spare, I'm not sure what to do with it.
Just lost it to a fire and not sure what to replace it with, alas. Too busy in the aftermath to justify building and parting stuff independently, so that probably leaves me with an off the shelf or a workstation + GPU combo.
I know you said it’s been ruled out, but I sometimes wonder if I should worry more about aging lithium batteries in older devices… the oldest ones (e.g. 2000s handhelds) are now old enough to be retro-cool and therefore worth keeping for nostalgia, but they’re also a bit scary.
Not only was that the opposite of computing trends, and just there to spite AMD, the new system had results like an i3-9350KF being overall faster than an i9-9980XE.
Use Geekbench instead of Cinebench or Userbenchmark.
Would you advise the same if the colors were reversed?
The majority of people here are AMD fanbois, so it's sometimes hard to discern whether something is rooted in objectivity or subjectivity.
There's a reason why x86 CPU users will only want to compare Apple Silicon using Cinebench. It's because Cinebench uses Intel Embree, which is hand optimized for x86 instructions. It's like testing Ryzen or Core CPUs on software optimized for ARM instructions, and then making a conclusion on how fast they are for x86 applications.
Use Geekbench.
You also didn't answer my question: Would you likewise advise against Cinebench if the colors were reversed? As it is, you're advising against Cinebench just because it favors Intel.
Today, it seems like CPU competition is a live again with Intel, AMD, Apple, Ampere, ARM, Graviton, RISV, Qualcomm, etc.
Even with the latest and greatest 4k, HDR, wide colour, display, no app should ideally use more then 256 MB of memory by those standards, unless it's even more complex.
That is to say, a theoretically ideal OS running on the A16 could run every mainframe program in existence until the late 1970s, simultaneously.
I still regularly use old systems for fun. I grew up with Macs, so that’s the point of reference. When I use these old systems, I pine for a few specific things I’m used to on newer systems—but it almost feels like nitpicking.
The things that I really want from a computer are pretty basic. Like good, consistent copy/paste and drag ’n’ drop, good autosave, good file browser, that sort of thing. It seems like new, half-baked stuff got dropped in our laps before the basics really got perfected.
It also kind of drives me up the wall when you see what was being done on computers in the 80's and 90's on sub 100Mhz processors and realizing just how much efficiency has been lost in the name of ease. Excel shouldn't need to use all 12 threads of my CPU and visibly take time to sort less than a 1MB of data but here we are.
I have been working on an essay regarding Permacomputing for the last few weeks and it can be kind of difficult to summarize at times. The closest I can get is that it is part retrospective about want worked in the past but with the direct goal of implementing the goals long term sustainable systems that do not require large external inputs.
Maybe there is the possibility to trim down a package like that into something akin to what we had in the 90's but, funnily enough, there is a lot of that 90's legacy that was stacked onto, it looks like there is just too much legacy in it to make that a viable path.
Microsoft gets pressure from their hardware partners to keep up the hardware replacement cycle, and they themselves of course get a cut of that via OS licenses.
Phone manufacturers seemingly invented "Always On Displays" to also cause people to start to think they needed a new phone or battery just a year/18 months into owning their device. Christ what a waste of power for such little value delivered! My wife said her new S23+ would barely last a day - so I turned off the always on display option, and now it lasts 2.5/3 daysish of normal usage.
It's a slightly different story with laptops, where battery life is an important feature. You'd think battery life would be important for phones too.
Also nice that they've finally added an option on Samsung phones to only charge to 85%, but it would be nice if it had smarter options like on iOS. Eg, an option to charge up to 100% just before you expect to wake up, so you have the full charge but it doesn't sit at 100% all night degrading the battery.
Upgrades will be forced via software obsolescence rather than hardware performance.
This is why I think the Linux/Free-Libre software folks would do well to focus a bit more on optimizing performance as there is a lot of room for improvement that will not need to be forced onto people.
Instead of their high-powered development machine with the latest CPU, tons of high speed memory, and the fastest SSD; they would get to experience what many of their customers have to endure on slower hardware with capacity constraints.
Nothing spurs optimization like seeing first hand how your code creeps along on slow hardware.
I5 Gen 3, 500g HDD and 8gig of ram.
Yes, outlook runs fine for the secretary, visual studio not so fine for debugging.
What I was saying that if the developers ONLY run their software on their high-powered computers and never try it on slow hardware, they generally resort to the 'it runs fine on my machine' response when customers start complaining about performance.
We had few surprises when some newbie dev noticed that the site doesn't work quite as well outside of 1Gbit connection with 1ms ping to the app server...
That is exactly what I do with my desktop product. I test it on really crappy hardware first.
On the Ps3 side, I had completely forgotten just how hamstrung it was in terms of memory, the OS and the load times on games are SLOOOOOW! And yet, they got Uncharted 3 and Crysis 3 out of that thing.
One of the most impressive feats I have seen would have to be Daytona USA 2 in the arcades. That thing is running on a single 166Mhz PowerPC 603 and a Real3D GPU that still needed all the polygon setup done on the CPU. How they got that out of that hardware is just beyond me. It is optimised to the max!
I guess you would get a similar state in some embedded systems. The folks at NASA working on the Mars Rovers have a set target and are usually targeting mid 90's MIPS or PPC processor so they have to be fixated on speed and performance.
Like it or not, computation is cheap and developers are expensive. Code isn't worth optimizing until it either becomes a bottleneck, or you are running it on thousands of machines.
Can’t you write closer to optimal code to begin with? With experience can’t you start with learnings of the past? A lot of things are easy and convenient to do now but people like you just parrot about not needing to optimize so as to not do it at all.
Maybe instead of thinking if it’s expensive or cheap think about actually wanting to do something decent? Or is everything fake and just a transaction?
The problem is programs that are running on thousands or millions of machines but don't get optimized, because the customer runs the program and pays the costs and it's hard for them to blame any program in particular.
This is a well-known fallacy. There's no guarantee your performance problems have a single bottleneck. In fact, more often than not, your entire program is poorly thought and the only way to fix it is a full rewrite, with the associated risks.
> When I was teaching, I often used this metaphor: suppose you’re writing some system, you decide that you should avoid premature optimization, so you take the usual advice and build something simple that works. In this metaphor let’s pretend that your whole program is a sort. So you choose a simple sort that works. Bubble Sort. You try it out and it functions perfectly. Now remember Bubble Sort is a metaphor for your whole program. Now we all know that Bubble Sort is crap, so you have to eventually change to Quicksort. Hoare likes you more now. So how do you get there? Do you just, you know, “tune” the Bubble Sort? Of course not, you’re screwed, you have to throw it all out and do it over. OK, except the greater-than test, you can keep that. The rest is going in the trash.
> But you got valuable experience, right? No, you didn’t. Anything you learned about the Bubble Sort is worthless. Quicksort has entirely different considerations.
> The point here is that a small bit of analysis up front could have told you that you needed a O(n*lg(n)) sort and you would have been better served doing that up front. This does not mean you have to microtune the Quicksort up front. Maybe down the road you’ll discover that part of the sort (remember this is a metaphor) should be written in ASM because it’s just that important. Maybe you won’t. There will be time for that. But getting the right key choices up front was not premature. There is a suitable amount of analysis that is appropriate at each stage of your product.
https://ricomariani.medium.com/hotspots-premature-optimizati...
The only real way forward that isn't a temporary workaround seems finding a new type of semiconductor that has lower overall resistance than silicon. Whoever figures out how to dope graphene and produce wafers without defects will probably make trillions.
The energy is a mix of leakage current and active current. Leakage current can be thought of as resistance - it's how much current flows through a transistor that's off. This can be better based on the material, but gets harder with smaller transistors. (Thinking about quantum tunneling as a resistance is good to get intuition, but not good enough to help solve the problem. A material with a lower bulk resistivity will not help here.)
Active current is based on capacitance. Each FET has a little capacitor that needs to be charged and discharged every time the logic is switched - that adds up. Lowering the capacitance of each FET would reduce the energy required to switch it, but generally comes with bad tradeoffs. High-k dielectrics increase the capacitance, all other things being equal. But all other things are not equal, and they are used to create better performing FETs with lower power leakage.
Those are the same thing. Or at least close enough as makes no practical difference. Only an extremely tiny fraction of the power used but a CPU is becoming anything other than heat.
Intuitively, you might think of it like this: to charge a capacitor (or transistor) up to a certain voltage, you need a fixed number of electrons. That number of electrons will always pass through the resistor and generate heat based on their energy. Even if you change the resistor value, its still the same number of electrons, and the same amount of energy.
What does change with resistance, though, is the time over which the power is dissipated. In practice, you have to make sure the resistors are small enough such that you can achieve your desired clock speed.
There are actual resistive losses too, but they're mainly related to power delivery.
[1] https://en.wikipedia.org/wiki/RC_circuit#Time-domain_conside...
To add a bit of troll physics here (but I'm told computers using this sort of principle actually exist), why not then channel those electrons to a boost converter that pipes them back into VCC, recycling most of the current? Theoretically a 99% power usage improvement, minus what the converter loses to heat, and that can be as low as 10%.
Re: channeling electrons - what you've described doesn't quite make sense. Fundamentally, if you're taking an electron at ground or 0V potential, and changing it's potential to VCC, it requires energy that comes from somewhere. The battery (or power supply) is doing exactly that. As the electrons flow back to the ground, the battery "recharges" them up to VCC potential.
What you can do, though, is put circuits in series between supply and ground. That way the electrons flow through the "top" circuit, do their thing, then flow through the "bottom" circuit. There's no free lunch though, as the voltage across each circuit will be reduced. Nonetheless, this is a common technique for low power analog circuits, and one I've used in the past. It's just not practical or worth it in digital circuits like a CPU.
Modern processors are very careful about this, and actively turn off the supply voltage to large parts of the die to prevent extra leakage current. The funny-but-appropriate name for this is "dark silicon" https://en.wikipedia.org/wiki/Dark_silicon
Damn TIL, I never would've expected that. But I guess it makes sense to use a few of the older, larger transistors that don't leak as much to power off a section of the smaller leaky ones while they're not performing any operations.
I understand leakage will go up if we increase voltages to support higher switching speeds, but aren't there still a lot of losses that happen with logic transitions and reduce when the states are stable, even if voltages are held constant?
I realize it we can't move charges around for free, but in some fantasy superconducting-fet logic circuit, wouldn't the power consumption be reduced? I.e. much of the waste is resistive losses while charging and discharging those gates.
Not really. It makes more sense to think about it as filling and emptying capacitors. You are charging the gate capacitance up to the supply voltage, then dumping that charge to discharge the gate to 0 again. The energy of each capacitance that gets charged and dumped is CV^2/2, which happens for each logic transition.
> I realize it we can't move charges around for free, but in some fantasy superconducting-fet logic circuit, wouldn't the power consumption be reduced?
If there was no resistance when distributing charge, it would help a bit, but not enough to change the clock frequency by more than 20%, assuming that the fantasy superconducting-fet had normal leakage and gate capacitance.
I guess I am entertaining the idea of an idealized Maxwell-demon CMOS circuit, if we could bounce the charge between gates with very little work to just pump the charge back and forth.
If you had a lossless bidirectional voltage converter circuit for each gate capacitance, then you could charge the capacitor from the supply and discharge it back into the supply, removing any switching losses.
As the sibling comment says, you would need voltage converters running both ways to avoid this waste.
Asynchronous clockless designs might also drastically cut the power budget but those have failed to find adoption for some reason.
Clockless designs mean that you compute readiness information on the fly instead of having it precomputed at design time. This additional run-time computation is not free, and tooling for clocked designs is good enough that the extra slack they need is often cheaper.
Disclaimer: I'm not a material scientist, so this is probably only partly correct.
IMHO it's only a very small part of why GPU power consumption is going up. The main reason is the completely unnecessary chase for the performance crown.
From personal testing: my GPU manages to get 95% of its peak performance while being power limited to 80%. So the in order to squeeze the last 5% of performance out of the device, 20% more power is pushed through it. It stays above 99% peak performance while being power limited to ~87%.
But even just looking at the raw numbers paints a different picture. About 12 years ago, a high-end GPU (e.g. GTX 480) had a power draw of 250W at a theoretical peak FP32 performance of 1,345 GFLOPS. This year's RTX 4070 has a theoretical peak performance of 29.15 TFOPS at 200W, so we went from 5.38 GFLOPS/W to 145.75 GFLOPS/W in 12 years - a 27x improvement in efficiency and a ~22x improvement in raw performance.
Now let's compare that to the numbers from a decade ago: a GTX 580 from 2010 had a power rating of 244W at 49.41 GTexel/s. A Geforce2 Ultra from 2000 used about 10W at 2.0 GTexel/s. So we went from 0.2 GTexel/s/W to - you've guessed it - 0.2 GTexel/s/W, so same efficiency with a ~27x increase in performance over a decade, though the efficiency is only a guess, since neither GFLOPS nor official power draw figures are readily available for 2000-era hardware.
Fast forward a few years so we can get reliable power draw numbers and comparable performance in GFLOPS, we have the high end GeForce 8800 GTX at 155W for 345.6 GFLOPS in 2006. Ten years later, the comparable model would have been the GTX 1080 from 2016 with 180W at 8.873 TFLOPS. So 2.2 GFLOPS/W versus 49.3 GFLOPS/W or a 22x increase in efficiency and a ~26x increase in performance over the course of a decade.
So during the past 23 years, power efficiency steadily improved, while raw performance increase also showed no signs of slow down in the GPU space. This is given the same generous time frames, to account for the occasional generational leap.
GPU workloads will eventually run into the same scaling limits. That is we will be unable to speed up each execution unit any further, or the primary work we give the GPU will not be able to be split into more threads and accomplish useful work.
So maybe GPUs still have some room until they run into the same problem as CPUs.
Absolutely. The 2006 "A View from Berkeley" is still a great paper [1]. And we still have a long ways to go on this recommendation:
> To maximize application efficiency, programming models should support a wide range of data types and successful models of parallelism
We are still stuck in the winner-take-all mindset when it comes to software development.
[1] https://www2.eecs.berkeley.edu/Pubs/TechRpts/2006/EECS-2006-...
Not quite. NVidia's entry level price/performance has not improved much since 2016. What's skyrocketing is the price of the top of the line models.
High end is now becoming a case of throwing us much money at the problem as possible.
The difference comes from usage. CPUs are shared by processes and threads that are designed to be unaware of each other, or to be even hostile. At the same time, a lot of programs are built in such a way that they don't exploit the parallelism available to them through CPU, or, even if they do, they do it in a very clumsy way (through a bunch of wrappers with their own limitations).
To contrast this, GPU programs typically use the whole GPU at once, and are written with parallelism in mind, with little to no wrappers.
Similarly, because the basic unit of CPU usage is a process, and the model of using CPUs is that processes aren't allowed to know about each other by default, the memory use becomes more involved, inter-process communication becomes more involved, permissions, access to network etc. -- all this complicates and slows down programs which want to use CPUs.
But, if, somehow, there was an OS that could use GPU to run processes on it, use VRAM for code / data of those processes etc. -- we'd have the same problems.
I don't think this is really true anymore? I mean, on a composited desktop, pretty much any UI app is a "GPU program". If you run a video game and it's not full screen, it's sharing the GPU (and it's increasingly common for "full screen" to actually mean a borderless window, too). Video players offload decoding. And then there's all the stuff that's using GPU to accelerate generic compute.
And once you have such sharing, it's just as adversarial as processes sharing CPU, when it comes to resource allocation, security etc.
In fact if CPU core count did translate more easily to performance gains I think already with the existing CPU's we'd have a fairly signficant one-time boost.
Maybe somebody has statistical survey of how much of the existing deployed CPU core count is typically used?
I've been suggesting engineers get more cores of lower speed to gain insight on what will be performant a few years down the road since I saw my first Xeon Phi.
It's been a while since clock speeds got higher (IBM has been pushing 5GHz in their highest end for the past couple years now and it doesn't seem likely they'll cross 6 anytime soon), but we get more cores every year. We now have 4-core entry-level machines and 2-core/4-thread ones are the bottom of the barrel, with a decent one being 8-core. Ampère just announced a 192-core server beast.
And then we have another thing: performance for most users has been "good enough" for the past couple decades. I haven't gotten a new computer just because it had a faster CPU since the early 2000's - they usually turn to dust well before they become too slow to use. My wife will need to upgrade her Macbook soon-ish for regulatory reasons (when Apple EOLs and stops patching macOS 12) and her laptop is still going strong. Considering that, there is little advantage in making all but the most demanding software more parallel.
This leaves the high-end, the stuff that needs a POWER10 or a Telum to run at acceptable speeds, and the cloud vendors, who'd kill to be able to serve 1% more VMs per kilowatt because 1% of their revenue is the GDP of a small country.
If Google's only expense were electricity, and profit margin was 50%, then saving 1% on power bill would increase profits by 1%.
I skimmed the Alphabet annual report: saw $260 revenue, $80 profit, $110 operating expenses (electricity and staff). Say $10 on power, then 1% is $0.1 - improve profits a bit over 0.1%.
Anyone know how many $/year Google spends on power?
My guess is very few cores are used on average. I did some testing with Solvespace to see which build options contributed most to performance:
https://github.com/solvespace/solvespace/issues/972
Obviously using OpenMP for multi-core was the big win. But what's not shown is that in typical usage (not the test I ran) if you're dragging some geometry around it will use all cores (in my case 4 cores / 8 threads) at about 50 percent utilization. That percentage probably drops as more cores are thrown at it due to Amdahl's Law. In other words, throwing double the cores at it will give a good boost to a lot of code that is already taking less than half the time (wall clock time, not CPU time).
We added OpenMP to a number of functions for significant performance gains. And in fact, any remining single-thread operation that gets the parallel treatment is likely to have a significant impact on overall performance since that is where most of the time is spent now. At this point we're more focused on features and bugs.
Algorithmic improvements are possible and I'd like to do those in the future, but they are much harder to do than sprinkling some #pragmas around critical loops. That will improve the scalability though, where multithreading really did not.
We've known how to scale software on large silicon for a long time, but as an industry we mostly can't be bothered (or lack the skills) to do it.
In theory, performance scales logistically with the number of parallel processors (Amdahl's law).
In practice, the limit is (and has always been) memory and i/o. That's why Apple silicon kicks everyone's ass.
If we want faster computers, the biggest gains are not to be found in making processors do more work. It's in designing systems (not just CPUs) that don't let the CPU wait around to do work.
??? But it doesn't: https://browser.geekbench.com/processor-benchmarks https://browser.geekbench.com/mac-benchmarks
Absolutely not. In practise the limit is (1) how many cores are actually _used_ by programs and (2) how much work is put into making anything fast at all, ever. We're using web frontends powered by python backends over a network. The vast majority of programs use nowhere near the resoures available to them.
now that single-core GPU/CPU/TPU/whatever performance is back on the front burner i think we'll see some horsepower and compiler improvements over the next few years. luckily the i/o problem has made great strides in the meantime so network/memory/storage will be there to support it, unlike in the past. ecc ram is also plummeting in cost, so that's good.
I wouldn't exactly call Golang[1] "specialized", but it does make multiprocessing easier than most languages.
1.Or Erlang or Elixir
Over the years, Moore's Law became a household term for computer performance doubling every couple of years. Under that definition, Moore's Law died in 2005 with Dennard Scaling so for most intents and purposes, Moore's Law has been dead for a long time.
It only held under the more restrictive definition of performance for tasks that were able to be parallelized perfectly, but even that has now been broken.
You could also argue that Moore's Law died in 2005 because the term CPU used to refer to what we now know as a CPU 'core' and the term was redefined.
Ultimately, what matters is that the performance the end user experiences hasn't been doubling every 2 years since 2005.
A lot of computation-heavy workloads do in fact scale rather well with an increased core count. That's why GPUs are essentially using thousands of cores, after all: nontrivial computation virtually always implies a sizeable amount of data, and a lot of data usually implies that you can split it up into multiple sections to compute in parallel.
A lot of desktop code is still single-core, that is true. However desktop CPUs are idle >90% of the time, only waking up to do a small burst of computation every once in a while. You're not going to notice some event loop finishing in 0.8ms instead of 1ms. Even then it'll probably be running multiple tasks from multiple processes in parallel, making good use of the available cores.
CPUs have simply gotten fast enough that they aren't a bottleneck for desktop use anymore!
Only recently AMD has opened the floodgates and we got consumer systems with 16 cores, where finally it's really hard to find an excuse to leave 15 cores idle, and there's so much raw power that even suboptimal scaling can give a big performance boost.
Also Rust became a thing, and it makes much easier to write reliable multi-core software, so I'm optimistic about software catching up.
When looking at these high core count processors, the typical use case is for a server in a data centre, and these sorts of applications run 24/7 and the cost of power is a massive part of the TCO. I think you have to address power per gflop when evaluating performance for these parts, as this is the criteria they were designed against.
I think the processors are costed in consideration of the TCO of a 2U dual socket machine with a 2-3 year expected lifespan. They will be designed and costed to show year on year improvements.
Oh, and i'm not sure inflation was included as it will be relevant over the timescales involved.
Does Google Sheets provide a "inflation-adjusted dollar" function?
But anyway, I’m not sure if it makes sense to expect performance/$ to always increase anyway. I mean, I know this started out by talking about multicore, but think about single threaded performance. They’ve already grabbed all the low hanging fruit, the challenge now is finding increasingly hard to hunt down tweaks… a small improvement might require massive engineering effort.
A simple and regular ISA is what makes ARM easier to implement, which also means lower power because of fewer transistors doing thankless work like decoding instructions and reordering them.
Of course Moore's law is slowing down, but cores/$ is an extremely silly metric to use
more to the point, the comparison also represents the period of time in which AMD, the one-time "lesser" player nipping at monopolist Intel's heels by competing on price, became the technology leader and started commanding a premium for their products. Meanwhile, Intel has not been forced to cede its position in the market, their existing contracts and business model is based on premium products, not cost leadership. It will take awhile for this to sort out.
You can buy CPUs that cost a fraction of any of those listed that will absolutely demolish even the best chips from 6 years ago today. All while consuming drastically less power for that performance as well.
Even mobile chips over 2 years old are within margin of error performance distance to those entry level Naples chips listed, at a fraction of the cost and power consumption.
I would not call that stagnation.
I run a home server powered by a Ryzen 5800H, nominally a 45W part, but I've seen it maintain a power draw far higher than that for hours under sustained load.
https://github.com/intel/linux-intel-lts/issues/33
The ongoing development:
Also you are not comparing to Intel's cost per core which would show that this pricing issue is not new. I think you just didn't notice it before.
In this case, don't let the author see $/GFLOPS for the last 10 years of GPUs!
https://twitter.com/nickdothutton/status/1194978743250538496...
Maybe the IPC or GHz really will be significantly lower, but I tend to think things like cache size will be the biggest hit, and that cache size change wouldn't show up in these graphs. Essentially, same number of transistors, but more compute less cache is my guess. But perhaps the cores really are smaller & narrower & the IPC * GHz rating doesnt budge much!
The best we can hope for I think is that open source will create frameworks/foundations on top of which people can then try to build. Elixir Phoenix has been that to some extent, basically taking the rails philosophy but making it super light and fast (my Phoenix APIs run with 40MiB of memory and response times ~1ms). Maybe those sorts of advancements can save us, but I can't think of a way to address the browser that way and realistically right now the browser is a huge area of the bloat. A ton of code that runs in the browser is terribly optimized, but even the base is quite big.
This is a gross oversimplification. While there’re inefficiencies and process abuses, this doesn’t mean nobody cares about speed and resources. It might look that way to purists who’re focused only on tech part of the businesses.
The real question is what happens first: change in software to adapt to slow growing compute or change in architecture to revitalize Moore's law.
They used to be a lot faster, 22nm CPUs (Ivy Bridge) came out in 2012 - so around 2 years to go from 22 to 14.
The reason we see increasing core counts is that 2 cores consume roughly 2x the power of 1 core (well, slightly more). Doubling clock speed from 5 GHz to 10 GHz would cost way more than 2x the power. Furthermore, pumping that much heat out of a sub-100mm^2 die gets harder and harder. You can have more transistors in a small space, but you simply can't have them change state much more than the previous generation did.
http://webcache.googleusercontent.com/search?q=cache:FBaO-jB...
I know the die sizes didn't decrease much, around 10% or so. And R&D increase is surely surpassing inflation.
Maybe our great^n grandchildren will then have lives very similar to the great^n+1 grandchildren thereafter, just like the older days!
GPUs have strong potential for improvement, and moving workloads to them helps on multiple fronts: performance, cost, power consumption.
The frameworks have been helpful but at the same time -- rounding buttons is not software engineering.
So instead of spending weeks or even months on trying to squeeze the last bit of performance out of an application that's "good enough" performance-wise, developers can use that time to roll out features or fix bugs instead, i.e. generating value for their customers.
It's simply a question of economics.
First off, it's a perfect baseline when comparing AMD chips since that time. Zen 1 was similar to Intel performance-wise, winning some benchmarks and losing some.
Second, the Raven Ridge (Zen 1+) APUs were IMHO excellent performance for the price at the time - even against Intel. I have not felt the need to build a new system since the Mellori_ITX:
How is it going to account for something like the introduction of AVX512, doubling the data throughput per instruction? What about increased cache greatly reducing the number of instructions waiting for memory access, or a faster memory controller? How would you even begin to quantify a new PCIe generation doubling SSD bandwidth, or a better core interconnect reducing NUMA penalties?
IPC is basically just a marketing number with today's CPUs. If you want to get a proper comparison, use a real-world benchmark.
This year has seen a need for a decent CPU for occasions where we must single thread. Both Local AI and python development.
For so long, we never maxed out our CPU.
In that, RISC-V might get some edge, because ARM licensing is expensive, but I don't think licensing is a significant cost for the x86 crowd.
OTOH, the server ARM people are really pushing it: https://www.semianalysis.com/p/sound-the-siyrn-ampereone-192...
But Altra's ain't cheap.
Using the last ASML's technology is a big part of the situation Using on-chip memory is another
See for example https://www.cpubenchmark.net/compare/4922vs5022vs5008vs5189/...
And that is a lie, to achieve those scores it turbos and consumes 300W+ at turbo.
Today, right now, is a moment where AMD's enterprise product are outperforming Intel's in benchmarks.
Except for a very small set of specific use cases, I think anyone recommending Xeon for enterprise solutions is professionally negligent.
https://www.cpubenchmark.net/compare/5022vs5008vs5189vs5031/...
So if others want to compete, they'll always be a few years behind, since the fab capacity is reserved to Apple. So any competitor either has to magically improve the architecture dramatically - which they can't, since that would require an architectural license which Apple has, but most others don't - or find a fab that can compete with TSMC's latest tech both in terms of price and available volume.
https://www.cpubenchmark.net/compare/2966vs3238vs3598vs3862v...
I think the article is just about server CPUs getting better but the price is keeping them from being significantly better for the price. Of course, in these markets, any increase in power is justified such that the buyers are not price sensitive.