I recall a Sophie Wilson talk about how things will never get faster, past 29nm.
I recall a Sophie Wilson talk about how things will never get faster, past 29nm.
For Server is it also about core count. For the same TDP AMD offer 64 Core option. On a Dual Socket System that is 128 Core. Zen 3 Milan is also socket compatible with Zen 2 Rome.
Basically for Server Intel has been stuck on 14nm for far too long. The first 14nm Broadwell Xeon was released in 2015, and as of mid 2021 Intel barely started rolling out 10nm Xeon Part based on IceLake.
That is half a decade of stagnation. But we are now finally getting 7nm server part, 5nm with Zen 4 and and SRAM Die Stacking. I am only hoping EUV + DDR5 will bring Server ECC DRAM price down as well. In a few years time we will have affordable ( relatively speaking ) Dual Socket 256 Core Server with Terabyte of RAM. Get a few of those and be done with scaling for 90% of us. ( I mean the whole StackOverflow is Served with 9 Web Server [1] on some not very powerful hardware [2] )
[1] https://stackexchange.com/performance
[2] https://nickcraver.com/blog/2016/03/29/stack-overflow-the-ha...
overall speed increases are still nothing like the old days and spectre style mitigations have been eating away at the improvements
I remember many, many years ago my PC couldn't run the latest adobe software as I was missing SSE2 instructions on my then current CPU. I believe there was some hidden compatability mode, but the software ran extremely poorly vs. the version of software that was out before that didn't require SSE2.
The microarchitectures have improved significantly though, which does matter. For instance, Haswell-era AVX2 implementations were significantly poorer than the modern ones in, say, Tiger Lake or Zen 3. The newer ones have completely different power usage and per-core performance characteristics for AVX code; even if you could run AVX2 on older processors, it might not have actually been a good idea if the cumulative slowdowns they cause impacts the whole system (because the chips had to downclock the whole system so they wouldn't brownout). So it's not just a matter of instruction sets, but their individual performance characteristics.
And also, it is not just CPUs that have improved. If anything, the biggest improvements have been in storage devices across the stack, which now have significantly better performance and parallelism, and the bandwidth has improved too (many more PCIe lanes). I can read gigabytes a second from a single NVMe drive; millions of IOPS a second, which is vastly better than you could 7 years ago on a consumer-level budget. Modern machines do not just crunch scalar code in isolation, and neither did older ones; we could just arbitrage CPU cycles more often than we can now in an era where a lot of the performance "cliffs" have been dealt with. Isolating things to just look at how fast the CPU can retire instructions is a good metric for CPU designers, but it's a very incomplete view when viewing the system as a whole as an application developer.
There always was an "Intel advantage" to compilers for decades (admittedly Intel invested in compilers more than AMD, but they also were sneaky about trying to nerf AMD in compilers), but with AMD being such a clear leader for so many years, I would hope at least GCC has started supporting AMD flavors of compilation better.
Anyone know if this has happened with GCC and AMD silicon? Or at least is there a better body of knowledge of what GCC flags help AMD more?
But in general I think it's a bit of a red herring for the thrust of my original post; first off you always have to target the benchmark to test a hypothesis, you don't run them in isolation for zero reason. My hypothesis when I ran my own for instance was "General execution of bog standard scalar code is only up by about 50-60%" and using the exact same binary instructions was the baseline criteria for that; it was not "Does targeting a specific microarchitecture scheduling model yield specific gains." If you want to test the second one, you need to run another benchmark.
There are too many factors for any particular machine for any such post to be comprehensive, as I'm sure you're aware. I'm just speaking in loose generalities.
Note that -march will use instructions that might be unavailable on other CPUs of the target. -mtune (which is implied by -march) is the flag that sets the cost tables used by instruction selection, cache line sizes, etc.
My test was a single threaded software rasterizer with real-time vertex lighting. That I compiled in vs c++ 2008 around 2010.
But the scalar speed isn't everything, because you're often not bounded solely by retirement in isolation (the system, in aggregate, is an open one not a closed one.) Fatter caches and extraordinarily improved storage devices with lots of parallelism (even on a single core you can fire off a ton of asynchronous I/O at the device) make a huge difference here even for single-core workloads, because you can actually keep the core fed well enough to do work. So the cumulative improvement is even better in practice.
And also that is the reason why I use such processor on my own computer, it was already "outdated" when I bought it, but since one of the things I like to do is play some simulation style games that rely heavily on single-core performance, I choose the fastest single-core CPU I could find without bankrupting myself. (the i5 4690k that with some squeezing can even be pushed past 4ghz, it is a beastly CPU this one)
We're near the end of IPC scaling per core. And it was never that good in the first place. Pentium 3 IPC is only 3-4x worse than the fastest Ryzen. Most of our speed increases came from frequency.
IMO we need to get off silicon substrate so we can frequency scale again.
I wonder if the end of scaling will push everyone into faster languages like Rust. You can't sit around 2 years for your code performance to double anymore. Will this eventually kill slow languages? I think so.
The arguable shift that is more significant is not about hardware its hopefully about newer languages like Rust making this performance cost less in terms of safety and development time which is a more recent development.
It's clear that's no longer true, which gives less excuses for using say Ruby over Java.
Think about it like the tempo of a song. The entire orchestra needs to play in sync with the tempo, but how many notes you play relative to the tempo is still up to each player. You can play multiple notes per "tempo tick".
We haven't had a strict clock speed = performance ratio for over a decade, it's just one component of it now.
https://en.wikipedia.org/wiki/Cycles_per_instruction
https://en.wikipedia.org/wiki/Instructions_per_second
Rather than just bumping clock speed, they've made improvements on what can be done within the clock cycle.
I recently did a desktop rebuild.
Went from a Ryzen 7 2700X to a Ryzen 5 5600X.
On paper, this looks like a downgrade. After all, the 7 is higher than a 5, right? I have two more cores with a 2700X...
However, the 5600x has about a 28% IPC gain over the 2700x in single core performance despite running at the same base clock speed. Literally 20%+ faster in the same tasks even when taking the turbo boost out of the equation.
The 2700x was at 12/14nm process while the 5600x is at 7nm which also helps with the power consumption as the 5600x has a 40 watt lower TDP.
Since what I need is better single core performance over more cores, the 5600X is quite an upgrade despite only being two years newer. (A LOT has happened in 2 years with AMD) With two less cores but significantly higher single core performance, the 5600X outperforms the 2700X on both single core and multicore performance.
Unfortunately for Intel, they did a lot of little tweaks and cheats to gain performance at the expense of security and now the mitigation patches pulls their chips back to the Bulldozer era in terms of current performance. They were also stuck at the same node for almost a decade now (Broadwell, 2014), so they had no gains from a shrink (speed of light propagation gain from a shrink, less heat and higher clocks also).
My i7-4790k is full of jank and microstuttering now, it has become unusable. Imagine how the servers might be doing if (and they should be) they are being kept up to date on security and microcode patches.
Then there also comes with spec upgrades within the processor generations. Ryzen 2700X to 5600X also means PCI Express went from 3.0 to 4.0... Not hugely significant among desktop users, but substantial for servers that need that amount of link bandwidth for compute cards and storage.
TLDR: Magic.
There are of course architectural improvements that improve performance with space to add gates etc, but when you have 1000s of chips cost of energy is the big thing.