A Look at the AMD Zen 2 Core
fuse.wikichip.org
fuse.wikichip.org
How do processors that split ops into uops implement precise interrupts? I sort of understand how the ROB is used to implement precise interrupts even with pipelining and OOO, but I don't quite see how processors map uops back to the original instruction sequence.
And I don't know. I guess is if each u-op is tagged with the instruction address within the process, that would do, but that's carrying around at least 32 bits, which is quite a large tag.
Alternatively tag indirectly, which is more likely (you can have maybe 256 instructions 'hot' at any time so an 8-bit tag on each u-op pointing to a 32 or 64-bit table entry (edit: holding the actual address of the macro-op). And the window for the ROB and the other thing that does instruction issue, is ~200 instructions, so that sounds more plausible).
All speculation on my part though!
I was mostly assuming that the issue of the operation reserved a spot in the outstanding buffer which would ensure sequencing of the write or commit after execution which would walk through the buffer in sequence.
But you're right that there are still more questions to how some of that data is tracked through the pipeline.
But that's also all speculation :)
How appropriate!
But, your idea I think will not work in general where there's branching.
Which is a) amazing and b) OMG the frigging complexity of something that has to run at sub-ns speeds. It's like sausages, the closer you look the less there is to enjoy.
Thanks!
> Like if one instruction gets split into two uops and the first one completes but the second one raises an exception.
That's not a problem. It's just one of n exception types that instruction can raise. Suppose a macro (say x64 instruction, if something like this exists) division instruction where one operand could be fetched from memory, you could have
r2 <- r3 / ^r4
where ^r4 fetches the contents of memory at address held in r4.suppose it's split up into u-ops
tr6 <- ^r4 ;; tr6 is temporary register 6, invisible to programmer
r2 <- r3 / tr6
you could have a division by zero at u-op 2, or an invalid address exception for u-op 1. Either of those are valid exceptions for the original single macro-op.Extrapolating from what Tuna-Fish said, the ROB is list of macro instructions, each instruction I assume will be tagged with its actual macro-op address, and each u-op must link back to the originating macro-op so macro-op retirement can take place, so we have a small (8 bit? Because ROB queue is small) pointer from each u-op back into the macro-op in the ROB.
Follow the 8-bit u-op ptr to the ROB, get the originating macro-op address, raise exception at that address.
Assuming I'm right, and assuming I understood you question correctly. I'll have to read his answer more carefully again.
edit: swapped ^ for asterisk as deref operation, as stars interpreted as formatting. Edit 2: slightly clearer.
If a uop throws an exception (let's say a fused add+ld), each uop can have a tag that helps you backtrace to its PC (instruction address) of let's say the start of the sequence, so you know what to inform the Privileged Architecture as to what "instruction" excepted. For many reasons, you need to store a list of PCs of the inflight instructions somewhere (although it is heavily compressed), so having a small ID tag to help reconstruct a given uop's PC isn't too onerous.
Ideally, multiple instructions may map to a single uop, but either none (or up to one) can throw an exception. The hard one here is something like load-pair uops; since each load can throw an exception. Some machines, if a fault is encountered, will refetch the pair and re-execute as independent/unfused loads. Other designs will just pay the pain of tracking which of the pair excepted and do some simple arithmetic off of that.
Heh, random note, I just looked up shift instructions on x86, there are 6 different ones, not RISC. But today there's a lot over 1000 instructions so a few shift variants are peanuts.
So in say 'rol mem_addr, shift', your inner ISA would be cracked to something like.
ld reg_temp0, mem_addr
rol reg_temp0, shift
st reg_temp0, mem_addr
This is all hearsay though; I could have certainly misheard/misremembered.Few games will suffer greatly from it, but there are several titles with RAM bottlenecks, like PUBG and FarCry.
Anyway, AMD has a much better price/performance offer than Intel. For general puprose Intel is totaly anihilated, but for the games they are still more than competitive.
It seems implausible that this user benchmark is a good indicator. The Zen 1 architecture exhibited nothing of the sort[0] -- it would be an order of magnitude performance regression.
I expect we'll start to see more accurate tests once the processors are actually released into the wild.
[0]https://www.tomshardware.com/reviews/amd-ryzen-7-2700x-revie...
Of course, the effects of ~20ns more latency on system memory accesses may not be as easy to observe in practice, especially if throughput is excellent. But we'll see, I guess; 'gaming' benchmarks should be a good test. (Meanwhile, I'm just going to be happy with Zen 2 if it provides a large speedup in compile times.)
However, I think it's still true that the actual release of the Zen 2 processors will make userbenchmark scores much more reliable as the scores regress to the mean.
>69 ns at 3600C16
I'd go for that.
That said, the X570 chipset has PCIe 4.0 with improved power management to match, even at the ITX end. But PCIe 4 will likely go through implementation improvements and right now is only useful for storage in really fast Raid configurations, 15GB/sec performance from 4 drives has been demonstrated at a relatively cheap cost, comparable to high end commodity NVMe PCI 3 storage.
I wish it were obvious which part was “best” but different implementations lead to different optimizations/costs/benefits...
For memory latency to matter, you need the worst case scenario to play out, which is a miss on all caches.
For cache to be made entirely irrelevant in performance, you'd need a situation where every access is a cache miss.
Yet even a pattern of completely random memory accesses will hit cache once in a while. And even then, a completely random access pattern is absolutely in disconnect with real world applications.
Thus, performance is not determined by memory latency. It is only a factor. And latency has actually improved, like other metrics, when compared to previous generations.
In short, wait for benchmarks. Just some three hours left for NDA lift.
I was kind of curious what the latencies might be for other contemporary processors/builds, and I'm not sure 70ns is actually really outside the normal margins.
Here's an i7-8700k build that is already pretty close to 70ns: https://www.userbenchmark.com/UserRun/18173216 - Also acknowledging that, this is not the 'best case' performance. But seems to be not so unusual either.
(edit: Removed section about latency ladder as I just realized it was caching and not system memory latency that we were most likely seeing initially, and therefore not terribly relevant.)
As far as I can tell, Zen 1 has similar latency characteristics[0][1]... so I guess I won't notice any degradation when upgrading to Zen 2.
But my point is more about advanced user builds and press benchmarks.
[0] - 2400 cl16. It's much worse than 3600 cl16, used in referenced Ryzen 3000 test. It would be interesting to see how Ryzen would perform with it. May be zen2 comparison with 2400 vs 3400 (3600 is very rarely achievable there as IMC trait of the platform) would be a good illustration.
[1] - 3200 C16. But it performs way bellow expectations. Modern Intel chip should have ~55ns with it. With 3600 cl16 intel will have ~40ns.
I'm not too concerned though, since I already use Zen 1 and it looks to be around the same. If I had to take a shot in the dark, I'm guessing it's just a consequence of the chiplet designs of Zen and Zen 2 that enabled them to scale so well. If it is such a tradeoff and there is not future gains to be made on latency for Zen platforms, I can accept that.
Though it does make me curious if Intel will ever stumble upon the same problems.
It affects general tasks too, but with much less magnitude than gaming, because games are concerned with frame times and overall latency the most.
What I'm asking is, is there a good way to test the effects of just memory latency?
Most increase is between 3200cl14 vs 3200cl12. 12%. Difference between this two is almost purely a Latency.
Then compare 3200cl12 and 3600cl14 - 3%, marginally no increase. Almost no difference in latency, only throughput and IF.
Past 3200 RAM throughput and inter-core-communication (IF) has very little influence for Zen1 gaming. For Zen2 this scenario would differ in some ways but not too much.
[1]: https://www.reddit.com/r/Amd/comments/c9x8v7/2700x_memory_sc...
If you develop on something unix-ish valgrind's cachegrind will tell you about your L1 performance. On recent Linux you can get this straight from the kernel with `perf stat` https://perf.wiki.kernel.org/index.php/ (cache-misses are total misses in all levels)
The most basic question is: are you randomly accessing more then your processor's cache worth of memory?
https://www.anandtech.com/show/14605/the-and-ryzen-3700x-390...
It also seems to ignore the comment on userbenchmark itself that indicates that latency is less of a bottleneck than bandwidth.
I'll say coming to the conclusion that Zen 2 is less than competitive in comparison to Intel offering a day ahead of launch is more than brave.
if you're playing a AAA singleplayer game, the GPU will almost certainly be the bottleneck. any i5/i7 tier CPU from the last eight years will be powerful enough not to starve the GPU most of the time. when I play this sort of game, I mostly just care about the average framerate. I don't care too much if every 10-15 minutes a complex scene causes a brief stutter. if this is the only kind of game I played, I would just get a cheap midrange processor (unless I had requirements due to unrelated workloads).
on the other hand, when I play a fast-paced multiplayer game (especially fps), I care a lot more about worst case framerates than average. in a game like counterstrike, framerate drops tend to happen when there are smokes, flashes, and/or multiple players onscreen simultaneously. in other words, they happen in the most important moments of the game! while I'm happy to average around 45-60 fps in the witcher, I want the minimum in csgo to be no less than 120. if you have a goal like this, even an old source engine game becomes pretty demanding.
there are also games like factorio where the simulation itself is difficult or impossible to split into multiple threads. singlethreaded performance and memory bandwidth set an upper bound to how big your base can be while still having a playable game.
Also games that are sensitive to memory latency are very poorly written games. This is not a new or contraversial opinion of PubG, which is know for being terribly written (and that is supposedly a reason for fortnite taking its audience so easily). It is also written using Unity and C#, which don't require hopping around in memory, but do make that the most obvious path, which is a trap for people who don't know better.
Farcry is not a game I would expect to have these problems. Ubisoft seems generally technically sound, despite releasing wildly unstable games, but it will take real numbers to get a clearer picture.
This inside relationship Epic had was the source for some of the erm friction between the two Battle Royale games.
1 - https://en.m.wikipedia.org/wiki/PlayerUnknown's_Battleground...
Didn't work. Intel smoked them, especially on server workloads that were cache optimized. I wonder if the same things is going to repeat itself?
I haven't build a desktop in almost a decade, so I was thinking of it when the top shelf ryzen chips hit the market. I don't play any games, but would use it for server dev and basic ML work.
AMDs smaller L1 was as definite negative at the time. This was back when hyperthreading could be a net negative because of the reduced L1 cache per thread so we would turn that off to.
The memory bottlenecks you encounter with the games you mentioned revolve around bandwidth and timing, not latency.
For heavily dynamic applications like editing tools where the user is free to shape the data as they please and the worst case is much worse, latency becomes a much bigger issue.
Is this a gimmicky marketing term, or is it logical\descriptive\rational?
Looks like an already studied technical description going back to 2005:
https://github.com/ChampSim/ChampSim/blob/master/branch/hash...
https://www.jilp.org/cbp2014/paper/DanielJimenez.pdf
> Introduced by Tarjan and Skadron 2005
> Basic idea:
> - Hash segments of branch history into different tables
> - Sum weights selected by hash functions, apply threshold to predict
> - Update the weights using perceptron learning
Since you don't have branch predictors for each address, you share the same prediction data for multiple addresses (this is one often underlooked cost to heavily branchy code).
There's a very approachable explanation in the following (around the 1:01:40 mark): https://youtu.be/8I_1TSs695I?t=1h1m40s -- part of Design of Digital Circuits - Lecture 18: Branch Prediction II (ETH Zürich, Spring 2019). Related readings: https://safari.ethz.ch/digitaltechnik/spring2019/doku.php?id....
These lectures are pretty great, by the way, highly recommended to anyone interested in computer architecture: http://people.inf.ethz.ch/omutlu/lecture-videos.html
Incidentally, this year's High-Performance Computer Architecture Test of Time Award has been given to "Dynamic Branch Prediction with Perceptrons" referenced in the lecture (from 2001, https://www.cs.utexas.edu/~lin/papers/hpca01.pdf): https://engineering.tamu.edu/news/2019/02/jimenez-receives-h....
Even Microsoft has made VM-sandboxing a consumer feature. It's time for AMD to join the bandwagon and help along the way, not keep the virtualization/other security features exclusive to server customers.
Available on UDOO Bolt, https://www.kickstarter.com/projects/udoo/udoo-bolt-raising-...
At least AMD made DRTM (SKINIT) available on consumer CPUs, where Intel still keeps DRTM (TXT) limited to business vPro chipsets. ECC is available on some CPU/motherboard/bios combinations for Ryzen, although ECC and low-power for Ryzen APUs is segmented to "Pro" models only available to OEMs.
It's still near impossible to buy devices with Ryzen Embedded, despite several system announcements from ASRock, but SEV is available today on Supermicro Epyc 3000 mITX motherboards.
Are there any articles like this about Apple or Intel chips? I like hearing about actual implementations, not just educational processors. (I am looking through the other articles on this site.)
I remember Stokes had a recommended book and quite a few articles on Ars Technica back in the day. Covering older hardware, but typically a prereq in understanding newer compatible designs.
Also a great book though maybe dated by now is Modern Processor Design by Shen.
>"Perceptrons are the simplest form of machine learning and lend themselves to somewhat easier hardware implementations compared to some of the other machine learning algorithms."
Can someone explain what is it about perceptrons that make them easier to implement in hardware?
These lectures are pretty great, by the way, highly recommended to anyone interested in computer architecture: http://people.inf.ethz.ch/omutlu/lecture-videos.html
Incidentally, this year's High-Performance Computer Architecture Test of Time Award has been given to "Dynamic Branch Prediction with Perceptrons" referenced in the lecture (from 2001, https://www.cs.utexas.edu/~lin/papers/hpca01.pdf): https://engineering.tamu.edu/news/2019/02/jimenez-receives-h....
Expect large price cuts from Intel soon.