Performance variation in ‘identical’ processors
shape-of-code.coding-guidelines.com
shape-of-code.coding-guidelines.com
Die are sorted based on something called a bin split: die are binned immediately after wafersort based on their leakage (there are special transistors implanted near the scribe-lines that indicate tons of characteristics, as well as DFX units through out the die that are rings of 20 inverters that oscillate, also indicates tons of data on how the die behave, however testing the those buggered DFX circuits takes an enormous amount of time, and you can't slow down wafersort, so there are proxies).
The bins are designed in such a way to maximize profit and performance based on the die characteristics. Thermal throttle plays a role in this and each bin (among various vectors) is allowed some tolerance, which is exactly what OP has discovered. However, this has been going on for coming up on 30 years! So nothing really new here, I just thought I'd let you know that of course Intel is aware of this, and they never claim performance numbers outside of the tolerance allowed for thermal throttle.
Even things like the characteristics of the application of the thermal paste and the roughness of the individual fan ducts can matter, beyond the obvious bottom of rack vs. top of rack, position of processor in individual chassis, etc effects.
tl;dr-- probably most of what is measured is not silicon to silicon variation.
This kind of in situ measurement might be useful to feed into a job scheduler's weight or something but at the end of the day "environment is different" and maybe you have to derate your cluster's expected performance to account?
> We show that this variation is further magnified under a hardware-enforced power constraint, potentially due to the increase in number of cores, inconsistencies in the chip manufacturing process and their combined impact on processor’s energy management functionality
and recommendations:
> • Characterizing node performance based on averaged performance distorts the true impact of manufacturing variation on processors on the node and therefore should be avoided
But the temperature of the silicon itself has a lot to do with the performance you get at a given power level.
Voltages and currents are easy to measure relatively precisely.
Interestingly, that recommendation could be interpreted "empirical studies based on averaged node performance" give a very distorted view on "the true impact of manufacturing variation on processors".
It's an interesting set of measurements, but the assumed source of the variation is dubious, and it's not clearly what, if any, actions it really supports.
Then there is that lovely gotcha that unless all your components are from the same batch, you may well see small changes in things like ancillary chips upon motherboards due to chip supplies being from multiple vendors or even vendor variation.
Many variables that widen the margin of error and would of been nice if they at least had one system in which they tested a few CPU's, to at least eliminate those extra variables. Ideally, look at the results they got, get the 5 slowest and 5 fastest CPU's based upon the results and retest those upon one system and compare the results. That would of been great and made the difference.
Two modern miracles. Insane complex and capable tech.
Simple. Hot. Air. Basics.
Oh well, maybe laugh. I would.
[1] https://en.m.wikipedia.org/wiki/Fluorinert [2] https://en.m.wikipedia.org/wiki/Cray-2
But for real, I recently replaced the thermal paste on four 10 year old servers and performance improved by 20%!! A full core of performance by ensuring proper heat transfer
remember intel cpus will simply throttle themselves at high temps, so replacing the paste will allow heat dissipation across all the cores (as I witnessed first hand)
The station actually got techs to fly to Australia to try to diagnose the problem -- A bunch of SCSI disks stacked vertically was causing the topmost drive to overheat.
Way back, I had a 4GB SCSI hard drive that was extra tall, extra fast, extra noisy, and extra hot. I constructed a thick rubber box around it to try noise-proof it a little, but I also had to strap a CPU cooler to it and have the airflow enter the box and exit at a hole in the box after wrapping all the way round the hard drive. It worked. But just wrapping it in rubber would have been a very quick way to cook the drive. This is a drive that if just left running bare on a table would get too hot to touch.
Hard drives produce a heck of a lot less heat these days.
Plus, and I could be wrong here - the mounts are not really designed to be heat conductors (the rails are plastic on my high end Dell workstation); I suppose the heat is extracted via air flow.
Probably still an issue today, though there's a lot fewer vendors and variation.
My employer runs a few HPC grid workloads on AWS and Azure and these softwares would benchmark the instances on boot and drop the ones that were below a certain threshold, and we'd keep spawning instances until the fleet was up past our threshold. (Ultimately, we didn't buy the software).
I wondered (not too hard) why there could be such a difference in performance, but your comment has given some good insight into that.
It's very similar to the "download boosters" of old which worked until they got popular which is when download sites started throttling.
Do we really think that someone who is looking at this type of thing in enough detail to produce the graphs in the article wouldn't think of ambient temperature differential? REALLY?!
Something not being documented well enough for us does not mean that it wasn't considered... Some thing that we thought of going unmentioned does not indicate that said thing was not considered or tried.
This article doesn't mention power supply equivalency or anything about how many people were present in the datacenter during a given test. Here comes the patented HN 'hot take': "uh well actually they don't mention power supplies - are we sure these devices were even powered on?"
The paper itself is pretty clear with the methodology: the throughput of each processor was measured in an installed supercomputing cluster. There are obviously sources of variation in processor performance in a supercomputing cluster beyond the actual silicon plugged into the socket. The experiment has no ability to control for this variation, and no attempt was made.
The reason why we have research and research papers is so we can read about the methodology and think about the limitations of what is measured. It's an interesting measurement that the researchers made in situ; but it also has obvious limitations.
> Third, the median processor performance (shown by the green, blue and red lines) between processor 0 and processor 1 on Sandy Bridge show up to 1% difference. However, that difference in median processor performance increases to up to 5% on Broadwell
This socket 1 vs socket 2 variability alone explains 1/3rd of the magnitude of the difference measured in the study-- which makes it quite clear we're not just measuring silicon variability.
e.g., memory, which presumably have their own temperature characteristics
To my knowledge, and this is based on DDR3 and older; memory frequencies are essentially fixed because the transceivers on both ends need to sample in the middle of a bit cell, and to do that they need to know the clock period, which must not change once it's known. There's a delay-locked-loop (DLL) in the RAM which generates a phase-shifted local reference clock.
If the processors could all be locked to one constant frequency (i.e. all the power/performance "dynamic tuning" stuff disabled) that would help show whether there's other sources of noise. This of course also assumes the clock generators are identical.
I would say that if it make 15% difference it _was_ related to performance... It might not have been an option on a screen specifically stated as being performance related options, but that could just be a bad menu layout.
I'd be interested to know what the option was, if you have any specific memory of that?