HNHacker News
TopNewBestAskShowJobs

clamchowder

163 karma · joined September 16, 2022

submissionscomments
clamchowder··on Characterizing gaming workloads on Zen 4
(author here) I think it's not useful to eliminate branches in hot loops. These games have giant instruction footprints and a high branch rate. A loop will probably fit within the L1 BTB and uop cache, and probably won't benefit from eliminating branches.

Exception would be a branch that's near impossible to predict, like one that depends on a randomly generated value or otherwise doesn't correlate well with global history.

clamchowder··on Kryo: Qualcomm’s Last In-House Mobile Core
Maybe it was built from Krait, but I'm not sure. I never tested Krait and don't have microbenchmark code written to deal with 32-bit arm
clamchowder··on Kryo: Qualcomm’s Last In-House Mobile Core
(author here) It is not based on A72. Kryo doesn't have anything in common with A72 besides understanding the same instructions. I covered that in detail in the article. Kryo is wider and has a very different out of order engine than A72.

Not mentioned was that it doesn't inherit a lot of A72's quirks. For example, A72 has peculiar decode restrictions. It can only handle one NOP per clock, and NOPs somehow consume out of order resources beyond just a ROB entry. Kryo has none of these weird characteristics.

clamchowder··on Kryo: Qualcomm’s Last In-House Mobile Core
(author here) yeah up to 2016. The Snapdragon 821 was released in 2016, and was the last one to have an in-house Qualcomm core
clamchowder··on ARM or x86? ISA Doesn’t Matter (2021)
All modern chips have pipelined decoders, including ARM ones. For example, the Cortex A72 has three decode stages, and it's running a 3-wide decoder at low clock speeds.
clamchowder··on ARM or x86? ISA Doesn’t Matter (2021)
Note that Golden Cove has a 6-wide decoder, and does not have a clear advantage over Zen 4 with a 4-wide decoder. Other parts of the architecture affect performance more, so it doesn't make sense to put more area and power budget into the decoder.

(also I'm the author of the article)

clamchowder··on Van Gogh, AMD’s Steam Deck APU
(author here) full disclosure, I work at Microsoft. On Azure though, not on Windows.

But using Windows, as another commentator pointed out, was done because my latency test gave very weird results from SteamOS and needed more investigation. I also used Windows to test the iGPU because Nemes's Vulkan test ran into problems under SteamOS.

clamchowder··on Van Gogh, AMD’s Steam Deck APU
(author of the article here) Agreed. I'm just commenting on the hardware architecture. The deck is a pretty competent gaming device, given its ultraportable form factor and tight power constraints. You can find that out on any number of sites that have reviewed the Steam Deck from gaming experience perspective, so I didn't think it was necessary rehash that.
clamchowder··on Loongson’s LSX and LASX Vector Extensions
They tell you how to detect support but don't describe the instructions or execution environment (registers, etc).
clamchowder··on Loongson’s LSX and LASX Vector Extensions
Feels like a shotgun approach. They have x86-64 (Zhaoxin), RISC-V (Xiangshan, Alibaba), MIPS -> Loongarch (Loongson), and aarch64 (Phytium). My guess is the central government is trying to build domestic chips by throwing piles of money at anyone who says they can make one.

Result: a pile of different architectures running different instruction sets, none of which come close to being a high performance design.

clamchowder··on Previewing China’s Loongson 3A5000 with Performance Counters
OS is a Chinese linux distro called Loongnix (Loongnix 20, DaoXiangHu). It's Debian based, and we're using binaries from their packages (not compiling from source).
clamchowder··on Bulldozer, AMD’s Crash Modernization: Front End and Execution Engine
Yeah you're right about the multiplication performance. I checked back and 64-bit integer multiplication is one per four clocks.

I disagree that core count should be taken to mean anything about MT performance. You always have to consider the strength of each core too. Nor does twice as many cores for the same architecture imply 2x performance, because there are always shared things like cache and memory bandwidth. And even if those aren't limiting factors, MT boost clocks are often lower than ST ones.

clamchowder··on Bulldozer, AMD’s Crash Modernization: Front End and Execution Engine
(author here) The FPU is not quite equivalent to a Sandy Bridge FPU, but the FPU is one of the strongest parts of the Bulldozer core. Also, iirc multiply throughput is the same on Bulldozer and K10 at 1 per cycle. K10 uses two pipes to handle high precision multiply instructions that write to two registers, possibly because each pipe only has one result bus and writing two regs requires using both pipes's write ports. But that doesn't mean two multiplies can complete in a single cycle.

With regard to expectations, I don't think AMD ever said that per-thread performance would be a match for Sandy Bridge. ST performance imo was Bulldozer's biggest problem. Calling it 8 cores or 4 cores does not change that. You could make a CPU with say, eight Jaguar cores, and market it as an eight core CPU. It would get crushed by 4c/8t Zen despite the core count difference and no sharing of core resources between threads (on Jaguar).

clamchowder··on Bulldozer, AMD’s Crash Modernization: Front End and Execution Engine
(author here) With hyperthreading/SMT, all of the above are shared, along with OoO execution buffers, integer execution units, and load/store unit.

I wouldn't say there are issues with sharing those components. More that the non-shared parts were too small, meaning Bulldozer came up quite short in single threaded performance.

clamchowder··on AMD’s Zen 4 Part 1: Front End and Execution Engine
Yeah, will have Part 2 up shortly (author here). If you try to make a Wordpress post too long, typing gets extremely laggy to the point of being unusable. So splitting it into two parts lets me go a little more in depth without waiting two seconds for a word to show up after I type it.
clamchowder··on Microbenchmarking Intel’s Arc A770
The problem is loop overhead matters on AMD, because AMD's compiler doesn't unroll the loop. Nvidia's does, so it doesn't matter for them.
clamchowder··on Microbenchmarking Intel’s Arc A770
(Author here) See https://github.com/clamchowder/Microbenchmarks/tree/master/G...

It's very much a work in progress, as noted in the article. And some of the stuff that worked reasonably well on my cards, like the instruction rate test when trying to measure throughput across the entire card, went down the drain when run on Arc.

clamchowder··on Microbenchmarking Intel’s Arc A770
Apologies for the confusion, the tests were run with ReBAR. I've updated the article to reflect that

Shouldn't affect the conclusion for anything besides the PCIe copy to/from GPU tests.

clamchowder··on How quickly do CPUs change clock speeds?
Yeah, he has some pretty interesting stuff.

But AVX512 transition latency is pretty different from clocking up from idle. The CPU is already running at full tilt, and is making a relatively small frequency/voltage change. You can see from his writeup that it's well under a millisecond. I believe newer Intel CPUs have done away with AVX-512 downclocking completely.

I probably won't be able to look into that specific issue since I don't have any AVX-512 capable CPUs.

clamchowder··on How quickly do CPUs change clock speeds?
I meant it's unrelated to how fast a CPU clocks up.

If something is taking longer than 0.5 ms, you shouldn't be doing it in the ISR. Queue up a DPC and do your longer running processing there, or send it to user space. And yeah it might not be your fault if another driver's ISR was hogging the CPU core. That's just a case of a badly written driver screwing up the world for everyone, because they're not supposed to be doing long running stuff in an ISR in the first place.

https://docs.microsoft.com/en-us/windows-hardware/drivers/de... says an ISR shouldn't run longer than 25 microseconds. 0.5 ms is an order of magnitude off. Not something going from 800 MHz to locked 4 GHz will fix.

clamchowder··on How quickly do CPUs change clock speeds?
Thanks :) I guess I can't reply to a 7th level comment, so hopefully this one shows up in the right place.

I agree, there are multiple factors at play. But I don't think it's basically an implementation choice. Certainly it looks like it in some cases (S821 on battery, HSW-E and SNB-E). But it doesn't seem to be the case elsewhere. For example, speed shift lowers clock transition time by taking the OS out of the control loop.

clamchowder··on How quickly do CPUs change clock speeds?
> I think it's great the author made an account today on HN, and is replying to questions from his post. This is one if the best parts of the community here.

:)

> I would love to see some test done on newer CPUs

So, the site is a free time thing run by several people, all of which are either students or have other full time jobs (not full time reviewers with samples of the latest hardware). No one happened to have an ICL/TGL/ADL CPU available. I do have an ADL result from earlier today showing full boost after 0.5 ms, but that person also overclocked and may have used a static voltage. Maybe we can get an addenum out with more results but I wouldn't bet on it.

And yeah probably. In an extreme scenario you could just lock the CPU to its highest clock. Then you'll never see a transition period at all.

> logged the PSU I don't have a power meter or PSU capable of logging

clamchowder··on How quickly do CPUs change clock speeds?
Yeah, didn't want to start an article with a five paragraph essay especially when wordpress pagination doesn't work, so I can't get an Anandtech style multi-page article up.

And yep. You can even run a CPU at full clock all the time, meaning you will never observe a clock transition time. Cloud providers seem to do that.

clamchowder··on How quickly do CPUs change clock speeds?
That's a different and unrelated topic. If you're concerned about how fast device driver code can respond, well you can get a lot done in 0.5 ms even with the CPU running at 800 MHz or whatever the idle clock is.
clamchowder··on How quickly do CPUs change clock speeds?
Big cores each have a 512 KB L2 cache with a latency of around 24 cycles, while the little ones each have a 256 KB L2 with ~23 cycle latency. It's really terrible considering the low 2.34/1.6 GHz clock speeds. Then there's no multi-megabyte last level cache, so it's going out to memory a lot.

They do have 3 cycle L1D latency, but again that's at low clocks, and it's a 24 KB L1D so it's 25% smaller than most L1Ds we see today. Not great considering you're eating 20+ cycle L2 latency once you miss the L1D, unlike most CPUs today that have 12 cycle latency for 256 or 512 KB L2s.

clamchowder··on How quickly do CPUs change clock speeds?
> The Linux CPU frequency governor literally uses it as part of the algorithm for calculating its sampling rate

Yes, the governor can play a role. It's visible to the user, which is the point. Also, the ondemand governor is actually irrelevant to the article as the S821 and S670 used the interactive and schedutil governors respectively, and the i5-6600K was using speed shift.

I think we're disagreeing because I really don't care about how fast a CPU could pull off frequency transitions if it's never observable to users. I'm looking at how it's observable to user programs, and how fast the transition happens in practice.

> Processor still performs periodic sampling...

Same as the above, that's not the point of the article. I'm not measuring "what could theoretically happen if you ignore half the steps involved in a frequency transition even though a user realistically cannot avoid them without serious downsides" (like artificially holding the idle voltage high and drawing more idle power, as in the Piledriver example)

> In any case, a multi-millisecond delay in switching frequency isn't because the processor is waiting for the voltage to increase.

Yes, there are other factors involved besides the voltage increase. I never said it was the only factor, and did mention speed shift taking OS transition commands out of the picture (implying that requiring OS commands does introduce a delay in CPUs without such a feature).

If you want to test how fast a CPU can clock up, without influence from OS/processor polling, please do so and publish your results along with your methodology. I think it'd be interesting to see.

clamchowder··on How quickly do CPUs change clock speeds?
Hey, author here. The Snapdragon 670, Snapdragon 821, and i5-6600K tests were done on Linux, and the rest were done on Windows. If Windows is delaying boost, it doesn't seem to be any different from Linux. And lack of intermediate states between lowest clock and max clock on all three of those (not considering the little cores) indicates the OS is not stepping through frequencies.

Since people don't usually write their own OS, I don't think it's correct to take the "maximum transition latency" reported by the CPU to mean anything, because users will never observe that transition speed. Also, processors that rely on the OS (no "Speed Shift") can transition very fast if their voltages are held high to start with, suggesting most of the latency comes from waiting for voltages to increase.

Please also read about Intel's "Speed Shift". While it's fairly new (only debuted in 2015), it means the CPU can clock up by itself without waiting for a transition command from the OS.

← PreviousPage 2 of 2