Instructions per cycle: AMD Zen 2 versus Intel
lemire.me
lemire.me
For instance instructions are assigned to one of a handful of ports when executed, certain instructions may only be assigned to certain ports, what ports an instruction may be assigned to differ between architectures. If an inner loop use only a few different instructions one architecture may be unlucky in that most of the instructions need the same ports, and so it can execute fewer instruction overall.
For real benchmarking use lots of different complicated jobs. It is not perfect, but it is the best way we have of comparing different processors head to head.
Picking 1 or 2 random microbenchmarks like the blog post author did is not useful to categorize overall performance across all real-world workloads. If he had picked different ones, they might have shown AMD twice faster than Intel.
> However, it is not clear whether these reports are genuinely based on measures of instruction per cycle. Rather it appears that they are measures of the amount of work done per unit of time normalized by processor frequency.
That’s precisely why nobody really uses IPC as a way to compare processors. “How much work done per unit of time” is a much better measurement and I guess for historical reasons, people conflate it with IPC.
But real textbook IPC is useless for comparison.
They behave like GPUs more and more with regards to clocks.
I think it's the other way around? CPUs had "boost" before GPUs.
In this case, there's probably only memory IO which (afaik) cannot put a process to sleep.
It's useful for comparing architectures and the implementation thereof, to gauge the potential of one line of processors over the other.
I agree that for the customer it's not the right thing to be looking for.
Even different Zen 2 CPUs have varied performance properties not just due to cores, but due to CCX count.
The exactly one use for such microbenchmark and that's optimizing the compilers.
Even if there were multiple implementations?
Also remember that x86-64 unlike x86 is not closed, and unlike POWER, RISC-V, ARM or MIPS is not actually well defined.
If AMD suddenly adds a new but useful instruction set like they did with 3DNOW in ancient times, or accelerate something reasonably common that way, say add a special SIMD conditional, where do you even start in comparison? What if Intel actually does add a useful FPGA programmable computing capability as promised or enhanced DMA?
Also, on real world benchmarks that don't fit neatly in cache, for a given chip IPC will tend to increase as you underclock it because that will cause memory latency to go down.
Say, if you used IPC only then you'd probably pick the latest Apple ARM CPU. Except it cannot go as high clock in any of the subunits as top AMD and Intel, cache is slower, and memory bandwidth abysmal in comparison.
Performance in seconds or performance per watt (unit is 1/(W*s)) in the workload you want to run is useful.
You cannot even estimate anything using microbenchmarks anymore easily since they expanded per unit local clocking in x86... (AMD in Zen+ and expanded in Zen 2, most ARM mobile CPUS, Intel since Broadwell E, expanded in Skylake.)
You get traps such as going for AVX and locally overheating the CPU where SSE2 equivalent would go faster in real life. It's all funny business.
IPC also heavily favors RISC instead of SIMD, likewise is biased against multicore. (Though not as much.) What counts as an instruction anyway?
And IPC, IMO, is a better measurement for a chip’s design than pure speed, as it removes the “but how good a process do you have access to” from the equation.
Having said that in general Intel still holds a slight edge on pure Ipc. However, considering the terrible track record of security issues and abysmal price performance ratio, a slight edge on ipc can be ignored and I would not consider Intel for most workloads at the moment. Above all, actual application benchmark trumps any ipc microbencmark.
The reason Intel had the "per core" superiority crown for years is that it had a better IPC performance due to design efficiency. Both manufacturers are pushing against the same frequency ceiling, so if you went AMD you had to significantly increase the core count to catch up, and could never match the still important single-thread performance.
We know from large scale, comprehensive benchmarks that AMD has massively picked up the pace and is neck and neck with Intel. At the same processor speed it matches the best Intel processors.
But yeah, this article is just terrible. Not just tiny, minuscule, extremely myopic benchmarks, but then a gross over-reach with conclusions. And in the way that ignorance begets ignorance, the fact that it's trending on a couple of social news sites means that now Google is surfacing it as canonical information when it's just a junk, extremely lazy analysis.
Isn't that what the Internet is for?
That is most certainly an overreach. An extraordinary overreach. Worse, it's absurdly using an AVX2 codebase, optimized for Westmere, as the baseline for "IPC" testing? The premise itself borders of gross negligence.
IPC as a generalized concept is a broad, general purpose set of instructions, not an absurdly narrow test.
Saying "Intel is faster at AVX512" is going to surprise exactly no one, and also happens to be irrelevant for the overwhelming majority of users and uses.
The microbenchmarking thing has gone on for years, and at this point anyone who has paid any attention is rightly cautious when stomping their feet and making declarations, because usually they're just pouring noise into the mix. Lazily running a couple of tiny tests is not the rigour to avoid deserved criticism.
I agree using Westmere isn't necessarily the best approach, but there is no difference in this case with either -march=native or -march=znver1.
The loop is small and simple, with only 9 instructions and compiles more or less the same regardless of march setting (I observed some basically no-op changes such as a mov and blsr swapping places). Here's the assembly (for the second test, with the bigger IPC gap):
top:
tzcnt r8,rcx
add r8d,edx
mov DWORD PTR [rdi+rax*4],r8d
mov eax,DWORD PTR [rsi]
inc eax
blsr rcx,rcx
mov DWORD PTR [rsi],eax
jne .topEven worse! Is this a defense, because it's remarkably unhelpful as one.
The blog post was clearly a cry for attention for some project -- let's just use some clickbait IPC claims to gain it -- and continually alluded to a whole project -- an extreme niche project that still wouldn't have any relevance. But instead it's a meaningless, completely misrepresentative micro-loop.
I think Daniel uses those examples because they are actual examples from projects that he is or has been working on, and he's familiar with them and actually cares about them, and because it's at least a notch more realistic than something totally synthetic.
It seems like a very roundabout thing to use as a cry for attention for SIMDjson (the project I assume you are talking about), and I don't believe that's the purpose. I see no problem in linking the project.
Picking two random benchmarks and trying to extract any kind of more general IPC claim is not on solid ground, but I'm pretty sure Daniel will say he's not doing that: he's only sharing these two specific results. That's a style that reoccurs across several entries in that blog, however, so if it triggers (as it has me on occasion) you might want to look elsewhere.
It would help if the blog post had some headings to separate the benchmarks and summary.
And to the other defense of "Well there are AMD people claiming the same in reverse, so that legitimizes this", I've seen exactly zero of those posts on here. None. They would be laughed off the site.
What we do have is that traditionally at a given frequency, per core AMD has long trailed on major benchmarks of significant, user-realistic loads. This is the the first generation in a long time where it actually doesn't, and where you don't need additional cores to make up the gap.
I am only talking specifically about the 2/3 claim at the bottom of the article, which for the avoidance of doubt, is simply a summary of the final measurement made in the article, i.e., the result of dividing 1.4 by 2.1. I know this because of its positioning in the article, because the numbers line up, because a different % IPC is given for the earlier measurement, and because an earlier version of post, with different results for the last experiment (with IPC of 2.8 and 1.4), showed a different ratio (50%).
How you are somehow interpreting the small clarification of the one line which was being discussion as wide-ranging defense of the article, I'm not sure. My broader thoughts are available here [1] and the comments on the article.
---
In this blog post, he’s paying attention to IPC because he’s typically working with inner loops where the data’s being delivered from RAM to L1 as efficiently as possible.
The main problem I have is that the claim in dispute seems to be that Zen 2 has comparable (perhaps slightly higher) IPC to Skylake, and then Daniel picks out two benchmarks and shows that Skylake has higher IPC than Zen 2... proving what exactly?
Contradicting people who said that Zen 2 had a higher IPC on every benchmark? Yes, those people were wrong, but it's easy to prove a point if you pick an argument almost no one was making it in the first place.
In the same (second) benchmark that he selected the "basic_decoder" sub-benchmark, but there is also another benchmark "bogus" which tests the empty function calling time, and this case I measure a reversed scenario: Intel at IPC 2.25 and AMD at 3.43. So should we now say that Intel IPC is "quite poor"?
I started reading through your blog last night. I’m slowly trying to learn how to go from being a programmer who doesn’t write slow code, to one who writes fast code, so absorbing a lot about vectorization and ILW, etc.
I measured 2.0, but I guess Daniel is using docker w/ a slightly different compiler version, so I think it the gap is sufficiently small that we can declare "close enough". I also measured quite different numbers for SKL (2.0) vs SKX (1.7), which is quite odd given the non-memory intensive behavior of the test: in that scenario, I'd expect SKL and SKX to perform identically.
Exactly my thoughts.
Casually hanging out sites that are bias towards AMD, even those never claimed Zen 2 has same or higher IPC than Intel.
https://www.agner.org/optimize/instruction_tables.pdf
Edit: This is wrong as BeeOnRope points out below.
The first is SIMD heavy, so Zen 2 mostly closing the gap with Intel in one of the areas where Zen 1 was very weak is a good thing.
That said, I don't agree it's a tzcnt benchmark - there are about 9 instructions only one of which is tzcnt. I'm not sure why Zen2 is worse here.
After reviewing the example again, there's no obvious reason why Zen 2 is slower, although it's likely a rare edge case. Too bad there's nothing decent like VTune on AMD platforms.
I remember one session where my choice of temporary register significantly impacted throughput while implementing an unrolled int[] hash fn on my Kaby Lake processor. I never figured out exactly why, but sharp edges do exist even on Intel chips.
Also, I could not reproduce Daniel's results: I got IPC of 1.77 (SKX) or 2.00 (SKL) compared to Daniel's reported 2.80 (SKL, I think), so Intel still better but by a smaller margin. Waiting for clarification on that one.
This is an interesting idea, but I'm not sure how you could derive meaning from comparing two vastly different architectures at such a high level.
There is more than execution ports in design of processors. Not every task can be SIMD optimized to extent of approaching theoretical IPC limits, most will be bottlenecked by memory access or even IO.
I prefer the "fake" but real-world IPC. Same clocks, same real world task, measure time to finish.
If you care enough about a particular CPU to do benchmarks, you should benchmark what YOU care about.
Lemire's job is to improve the implementation of particular algorithms to make optimal use of the hardware. Knowing the different theoretical hardware limits tells you how good an implementation is doing along different axes, and benchmarking those limits is a critical part of doing Lemire's job correctly.
You probably have a different use case for computers than Lemire, and it is therefore completely reasonable for you to care about different benchmarks.
> Instructions per cycle (IPC)
> For many people, this is the holy grail of CPU measurements in terms of how fast an architecture per core really is.
Based on his work with simdjson, professor Lemire seems to be quite aware of microbenchmarks being problematic. But general articles out here and on HN are proclaiming Intel is doomed and can never recover, due to mitigations/lack of cores/lack of chiplets. Those concerns have yet to be reflected in the stock price.
IPC MB’s, in my experience, tend to benchmark best case scenarios and that is probably the exception rather than the rule for application workloads in modern MA’s. Case in point, microbenchmarks showed significant improvements in IPC for Zen2 in lieu of Skylake yet for the application workload (CPU data bound), Skylake held up neck and neck.
The more appropriate benchmarking metric for post-Zen2 processors is CPI [0].
[1] There are some rare exceptions, such as https://travisdowns.github.io/blog/2019/03/19/random-writes-... , but it is unlikely to matter here.
It’s also false if you’re using a hypervisor that mitigates the iTLB multihit issue.
So it's something worth checking.
No hypervisor involved.
One thing itches me with the presented comparison: it is running very few benchmarks generated with the same compiler. For a thorough IPC analysis, shouldn't the tests rather being programmed in assembly to exclude any influence by the compiler choice? Also probably a wider range of algorithms should be checked, as IPC on modern processors depends less on how many cycles a certain instruction takes (you should be able to find that in the manuals), but how well multiple components of the processor can be utilized at the same time. Which extremely depends on the actual program to be run.
If you can keep it fed then ok but one cache miss to main mem, either instruction or data, will allow the instruction buffers to completely empty and stay empty for quite a long time. I don't think you can control placement to reasonably assure cache hits always for anything but the most trivial code, am I missing something?
Also if you could keep a consistent throughput like this I wonder if thermal throttling might have to kick in. I mean you're doing a lot of work...
Found this http://manpages.ubuntu.com/manpages/trusty/man2/perf_event_o... and that article doesn't instill much confidence in the reliability of these counters. Comment for CPU_CYCLES says "Be wary of what happens during CPU frequency scaling", comment for INSTRUCTIONS says "these can be affected by various issues, most notably hardware interrupt counts", BRANCH_INSTRUCTIONS says "Prior to Linux 2.6.34, this used the wrong event on AMD processors" and so on.
If I wanted to measure what OP was measuring, I would disable frequency scaling (probably doable on overclocker-targeted motherboards, also search finds some utilities which claim to do that, both windows and linux ones), measure time, then divide by frequency.
perf is all about reading actual hardware counters. It's awesome for this. There is essentially nothing made up about perf's output, except to the extent that the hardware itself reports inexact output. (For example, perf annotate may attribute events to an instruction near the instruction in question on older hardware, because older hardware has a small amount of skew when sampling.)
Sorry guys but Intel is still king of single core performance. But that's not a problem because I'm sure by 2050 most desktop applications and games will correctly make use of many cores, then AMD will reign
Assuming civilization will survive until then, given current political trends, is rash.