Exception would be a branch that's near impossible to predict, like one that depends on a randomly generated value or otherwise doesn't correlate well with global history.
163 karma · joined September 16, 2022
Exception would be a branch that's near impossible to predict, like one that depends on a randomly generated value or otherwise doesn't correlate well with global history.
Not mentioned was that it doesn't inherit a lot of A72's quirks. For example, A72 has peculiar decode restrictions. It can only handle one NOP per clock, and NOPs somehow consume out of order resources beyond just a ROB entry. Kryo has none of these weird characteristics.
(also I'm the author of the article)
But using Windows, as another commentator pointed out, was done because my latency test gave very weird results from SteamOS and needed more investigation. I also used Windows to test the iGPU because Nemes's Vulkan test ran into problems under SteamOS.
Result: a pile of different architectures running different instruction sets, none of which come close to being a high performance design.
I disagree that core count should be taken to mean anything about MT performance. You always have to consider the strength of each core too. Nor does twice as many cores for the same architecture imply 2x performance, because there are always shared things like cache and memory bandwidth. And even if those aren't limiting factors, MT boost clocks are often lower than ST ones.
With regard to expectations, I don't think AMD ever said that per-thread performance would be a match for Sandy Bridge. ST performance imo was Bulldozer's biggest problem. Calling it 8 cores or 4 cores does not change that. You could make a CPU with say, eight Jaguar cores, and market it as an eight core CPU. It would get crushed by 4c/8t Zen despite the core count difference and no sharing of core resources between threads (on Jaguar).
I wouldn't say there are issues with sharing those components. More that the non-shared parts were too small, meaning Bulldozer came up quite short in single threaded performance.
It's very much a work in progress, as noted in the article. And some of the stuff that worked reasonably well on my cards, like the instruction rate test when trying to measure throughput across the entire card, went down the drain when run on Arc.
Shouldn't affect the conclusion for anything besides the PCIe copy to/from GPU tests.
But AVX512 transition latency is pretty different from clocking up from idle. The CPU is already running at full tilt, and is making a relatively small frequency/voltage change. You can see from his writeup that it's well under a millisecond. I believe newer Intel CPUs have done away with AVX-512 downclocking completely.
I probably won't be able to look into that specific issue since I don't have any AVX-512 capable CPUs.
If something is taking longer than 0.5 ms, you shouldn't be doing it in the ISR. Queue up a DPC and do your longer running processing there, or send it to user space. And yeah it might not be your fault if another driver's ISR was hogging the CPU core. That's just a case of a badly written driver screwing up the world for everyone, because they're not supposed to be doing long running stuff in an ISR in the first place.
https://docs.microsoft.com/en-us/windows-hardware/drivers/de... says an ISR shouldn't run longer than 25 microseconds. 0.5 ms is an order of magnitude off. Not something going from 800 MHz to locked 4 GHz will fix.
I agree, there are multiple factors at play. But I don't think it's basically an implementation choice. Certainly it looks like it in some cases (S821 on battery, HSW-E and SNB-E). But it doesn't seem to be the case elsewhere. For example, speed shift lowers clock transition time by taking the OS out of the control loop.
:)
> I would love to see some test done on newer CPUs
So, the site is a free time thing run by several people, all of which are either students or have other full time jobs (not full time reviewers with samples of the latest hardware). No one happened to have an ICL/TGL/ADL CPU available. I do have an ADL result from earlier today showing full boost after 0.5 ms, but that person also overclocked and may have used a static voltage. Maybe we can get an addenum out with more results but I wouldn't bet on it.
And yeah probably. In an extreme scenario you could just lock the CPU to its highest clock. Then you'll never see a transition period at all.
> logged the PSU I don't have a power meter or PSU capable of logging
And yep. You can even run a CPU at full clock all the time, meaning you will never observe a clock transition time. Cloud providers seem to do that.
They do have 3 cycle L1D latency, but again that's at low clocks, and it's a 24 KB L1D so it's 25% smaller than most L1Ds we see today. Not great considering you're eating 20+ cycle L2 latency once you miss the L1D, unlike most CPUs today that have 12 cycle latency for 256 or 512 KB L2s.
Yes, the governor can play a role. It's visible to the user, which is the point. Also, the ondemand governor is actually irrelevant to the article as the S821 and S670 used the interactive and schedutil governors respectively, and the i5-6600K was using speed shift.
I think we're disagreeing because I really don't care about how fast a CPU could pull off frequency transitions if it's never observable to users. I'm looking at how it's observable to user programs, and how fast the transition happens in practice.
> Processor still performs periodic sampling...
Same as the above, that's not the point of the article. I'm not measuring "what could theoretically happen if you ignore half the steps involved in a frequency transition even though a user realistically cannot avoid them without serious downsides" (like artificially holding the idle voltage high and drawing more idle power, as in the Piledriver example)
> In any case, a multi-millisecond delay in switching frequency isn't because the processor is waiting for the voltage to increase.
Yes, there are other factors involved besides the voltage increase. I never said it was the only factor, and did mention speed shift taking OS transition commands out of the picture (implying that requiring OS commands does introduce a delay in CPUs without such a feature).
If you want to test how fast a CPU can clock up, without influence from OS/processor polling, please do so and publish your results along with your methodology. I think it'd be interesting to see.
Since people don't usually write their own OS, I don't think it's correct to take the "maximum transition latency" reported by the CPU to mean anything, because users will never observe that transition speed. Also, processors that rely on the OS (no "Speed Shift") can transition very fast if their voltages are held high to start with, suggesting most of the latency comes from waiting for voltages to increase.
Please also read about Intel's "Speed Shift". While it's fairly new (only debuted in 2015), it means the CPU can clock up by itself without waiting for a transition command from the OS.