What's new in CPUs since the 80s? (2015)
danluu.com
danluu.com
> Also, first-level caches are usually limited by the page size times the associativity of the cache. If the cache is smaller than that, the bits used to index into the cache are the same regardless if whether you're looking at the virtual address or the physical address, so you don't have to do a virtual to physical translation before indexing into the cache. If the cache is larger than that, you have to first do a TLB lookup to index into the cache (which will cost at least one extra cycle), or build a virtually indexed cache (which is possible, but adds complexity and coupling to software). You can see this limit in modern chips. Haswell has an 8-way associative cache and 4kB pages. Its l1 data cache is 8 * 4kB = 32kB.
Having helped build the virtually indexed cache of the arm A55 I can confirm it's a complete nightmare and I can see why Intel and AMD have kept to the L1 data cache limit required to avoid it.
Interestingly Apple may have gone down the virtually indexed route (or possibly some other cunning design corner) for the M1 with their 128 kB data cache. However I believe they standardized on 16k pages which would allow still allow physical indexing with an 8 way associative cache. So what do they do when they're running x86 code with 4k pages? Does they drop 75% of their L1 cache to maintain physical indexing? Do they aggressively try and merge the x86 4k pages into 16k pages with some slow back-up when they can't do that? Maybe they've gone with some special purpose hardware support for emulating x86 4k pages on their 16k page architecture. Have they just indeed implemented a virtually indexed cache?
This does not seem feasible because the 16k pages on ARM are not "huge" pages; it's a completely different arrangement of the virtual address space and page tables. The two are not interoperable.
Bigger than that. (https://github.com/apple/darwin-xnu/blob/2ff845c2e033bd0ff64...)
/* I-Cache, 256KB for Firestorm, 128KB for Icestorm, 6-way. /
/ D-Cache, 160KB for Firestorm, 8-way. 64KB for Icestorm, 6-way. */
Apple also uses a base 16KB page size on macOS, with 4KB being present for backwards compatibility (x86_64 apps) and the wider ecosystem (running Windows for example).
> Maybe they've gone with some special purpose hardware support for emulating x86 4k pages on their 16k page architecture
Nah, it's that when you run an x86 process, one of the TTBRs is pointed to a page size corresponding to 4KB pages.
https://www.anandtech.com/show/16226/apple-silicon-m1-a14-de...
Apple definitely takes advantage of VI=PI for their caches.
160kB is not a good size for a cache, surely must be a typo.
i$ has more options for handling aliases. You don't actually even need to eliminate them at all because it's a read-only cache. So it's not unusual to see these caches exceed the VI=PI geometry.
ECC? Needs hugepages? Or hardware erratum? I don’t know for sure.
Also, at least from the VM perspective icache is reported as PIPT, which is not too surprising but still worth noting.
Ridiculous. You're just making things up. It clearly doesn't have a 160kB 6-way L1 dcache. Nor are the small cores 64kB 6-way -- those don't even divide that's ridiculous. Why do you say it's not a typo?
They "explicitly" say it is 160kB in that backwater code comment. They also explicitly say it's 128kB in this presentation https://www.youtube.com/watch?v=5AwdkGKmZ0I&t=534s so what makes your source the authoritative one?? Pretty obviously it's a quick cut and paste job with a typo (two in one line actually) in it.
> ECC? Needs hugepages? Or hardware erratum? I don’t know for sure.
ECC? Erratum? Come on. It's clearly because it's 128kB.
And yeah, 192KB/128KB is what the HW behaves as by all accounts by measurements, making this even more odd…
- cache entries are assumed to be looked up by the virtual address (not physical), to save a TLB lookup
- within a single page, the page line index is the same in both a physical and a virtual address
- if the cache has enough sets, it doesn't just use the line bits, but also some bits of the page address
Here's the crucial bit: a single physical page address may correspond to several different virtual addresses, so if the cache kept using virtual addresses, it might have had many copies of the same line. Kaboom.
That's why having more sets in a cache than lines in a page is troublesome.
One thing I would have added to "since the 80s" is the x86 timestamp counter. It really changed the way we get timing information.
It does look like Hardware Lock Elision/Transactional Memory is something that seems like it will be consigned to the dustbins of history (again).
I believe a lot has happened around mobile and power as well. Apple boasts their progress every year, and at least some of them are real. But they are too secretive to talk about that. I hope some competitors have written some related papers. For example, the OP talks about dark silicon. What's going on around it these days?
> I hope some competitors have written some related papers.
You should read this:) http://people.engr.tamu.edu/djimenez/pdfs/exynos_isca2020.pd...
The new AMD Ryzen 5800x3d has 96MB of L3 cache. This is so monstrous that the 2048x entry TLB with 4kB pages only can access 8MB.
That's right, you run out of TLB-entries before you run out of L3 cache these days. (Or you start using hugepages damn it)
----------
I think Intel's PEXT and PDEP was introduced around 2015-era. But AMD chips now execute PEXT / PDEP quickly, so its now feasible to use it on most people's modern systems (assuming Zen3 or a 2015+ era Intel CPU). Obviously those instructions don't exist in ARM / POWER9 world, but they're really fun to experiment with.
PEXT / PDEP are effectively bitwise-gather and bitwise-scatter instructions, and can be used to perform extremely fast and arbitrary bit-permutations. I played around with them to implement some relational-database operations (join, select, etc. etc.) over bit-relations for the 4-coloring theorem. (Just a toy to amuse myself with. A 16-bit bitset of "0001_1111_0000_0000" means "(Var1 == Color4 and Var2==Color1) or (Var2==Color2)".
There's probably some kind of tight relational algebra / automatic logic proving / binary decision diagram / stuffs that you can do with PEXT/PDEP. It really seems like an unexplored field.
----
EDIT: Oh, another big one. ARMv8 and POWER9 standardized upon the C++11 memory model of acquire-release. This was inevitable because Java and C++ standardized upon the memory model in the 00s / early 10s, so chips inevitably would be tailored for that model.
This is more reasonable than it sounds. A TLB miss can in many cases be faster than a L3 cache hit
EDIT: Well, unless you use 2MB hugepages of course.
I know that modern cores have dedicated page walking units these days, but I admit that I've never tested the speed of them.
Each 64byte cache line could feasibly come from a different page in the worst case.
I think modern processors actually pull 128 bytes from RAM at the L3 level, if each 128 L3 cache line is from a different page, that's 768k pages in the 96MB L3 cache.
That being said, huge pages won't help much in this degenerate case. So your assumption might be valid for this argument actually.
So maybe it's not that much of an error.
And yes, they can be extremely useful for efficient join operations in some contexts, that would be challenging to implement without those instructions. Also selection for some codes. Not everyone needs them, but when you need them you really need them. And those use cases are frequently worth it. I use them to implement a general algebra, much like you suggest.
There are some more details in there but that's the main gist. The power control unit is its own operating system!
I've heard the joke that RISC won because modern CISC processors are just complex interfaces on top of RISC cores, but there really is some truth to that, isn't there?
There was an Arm version of AMD's Zen, it was called K12, but it never made it to the market since AMD had to choose their bets very carefully back then.
So RISC was a revolution over CISC, but CISC arguably always had 'a RISC core' inside, even before RISC was a thing.
Having coded all three (vertical ucode, horizontal ucode, and RISC decode stage HDL) the biggest difference I've found is that RISC tends to be three address, and CISC vertical microcode to be two address. I think that comes from the same place as assuming the presence of an I$. The gate count niche that lets you have a large SRAM block for cache also let's lets you have a large SRAM block for your register file. You're therefore not as dependent on a single accumulator or two tightly coupled with your ALU like earlier CISC designs had to encourage.
But as CPUs got bigger and suddenly had enough hardware to do "everything" at once, they had the problem of circuit depth. Sure, you can execute the whole instruction in one cycle but it's a really LONG cycle.
You fix that problem with pipelining ([R]eally [I]nvented by [S]eymore [C]ray, of course).
But you can't pipeline a complicated microcoded instruction set. Everything that happens has to fit in the same pipeline stages. So, the instruction set naturally becomes "reduced".
Basically: RISC is the natural choice once VLIW gets rolling. It's not about simplification at all, it's about exploiting all the transistors on much more "complicated" chips.
> But you can't pipeline a complicated microcoded instruction set. Everything that happens has to fit in the same pipeline stages. So, the instruction set naturally becomes "reduced".
That's what they said in the early 80s, but then the 486 came out. AFAIK the longest pipelined general purpose systems were also fairly heavily microcoded (Netburst).
Basically, if you were handed the transistor budget of an 80486 and told to design an ISA, you'd never have picked a microcoded architecture with a bunch of addressing modes. People who did RISC in those chip sizes (MIPS R4000, say) were beating Intel by 2-3x on routine benchmarks.
Again: it was the budget that informed the choice. Chips were bigger and people had to figure out how to make 2M transistors work in tandem. And obviously when chips got MUCH bigger it stopped being a problem because dynamic ISA conversion becomes an obvious choice when you have 200M transistors.
It's a sort of RISC core, but count the number of micro-ops some things translate to, it appears that some of them are pretty chunky.
Multi-level caches, TLBs, out-of-order execution, NUMA memory, SIMD, multi-threading - were all normal features of big machines (Crays, Cybers, etc) and are now normal on single chip machines. We have various optimizations and special case handling (eg branch prediction) but these are not normally directly visible to programmers.
GPUs (very large specialized coprocessors) were not a feature of older computers; and their capacity routinely dwarves the "main" CPU.
At a lower level, the routine use of high speed serial comms is new-ish, or at least a reinvention. It was used in very early computers, but use of serial comms for PCIe devices and SD cards is now routine. Perhaps it's just an optimization, but it was not usual a few decades ago.
One thing that I don't recall seeing on older computers is the multitude of fine-grained performance counters now available - useful for fine-tuning, and for peeking at what's going on outside your own process!!
oh and chips come with light up fans and crap now but theres no open standard on how to control the light color so everyone just leaves it in Liberace disco mode so its like a wonderful little rainbow coloured pride parade is marching through my case.
For a while it seemed that senator-level freedom would be available to the great unwashed, but it was but a mirage.
Carry on.
It used to be that you could do almost anything with impunity on the internet, trade books and mp3's, bitch on forum posts anonymously and anything less than the CIA getting a federal warrant (which they never did because even if they could see what you were doing it was obvious you were just being an edgy teen) would mean that your bullshit would be hidden and lost to the annals of time.
Now if you post something edgy on Reddit and someone takes offense to it then they can track down your twitter and insta and fb and swat your house and get you fired and turn you into a social pariah.
Hopefully this is obvious hyperbole, just saying that I get where they're coming from. From before the Matrix made the internet seem like it was cool and full of vigilante Jesus bullet-dodging ninjas with infinite guns and everyone wore black leather and hacked into the mainframe. There was a time when the internet was mostly text and pretty much sucked but that was still amazing and beautiful and now it's an all out slugfest between 5 billion people all bickering and posing for points, fame and fortune from the other 5 billion humans on the same 5 websites while corporations spur the up and comers into greater feats of dickishness in order to garner attention that can then be used to possibly sell some random cream made by child labor in Morocco to fight off a butt wrinkle or something stupid like that that doesn't matter.
The internet is an amazing thing and computers are amazing but this ship did not end up where its original captains were sailing it and no one knows if we could ever turn back.
This is why I continue to pay a premium for workstation/server grade hardware even when I'm assembling the system myself.
Most of the other lighting software sets it off because for whatever reason, they appear to do something with running processes or windows registry or something that causes EAC to panic and crash my computer. The default ASUS motherboard lighting controls set it off because apparently they aggressively reach into the running game binary in order to set the lighting appropriately.
A) threats to freedom and privacy of computing
B) that it is difficult to control the lights of cpu coolers
And there will be very helpful responses fixing B).
agreed on all points too
The compiler alone, but also the code, can create branches where the binary checks if certain instructions are available and if they are not, use a less optimal operation.
Backwards compatibility for modern binaries basically. But not forward ability to see the future instructions that haven’t been invented yet.
Not all binaries are fully backwards compatible. If you’re missing AVX, a surprising number of games won’t run. Sometimes only because the launcher won’t run, even though the game plays without AVX.
But for a pure pre-compiled example there is Apple Bitcode which is meant to be compiled to the destination architecture before run (as opposed to JIT code which is compiled when run). It's mandatory for Apple watchOS apps and when they released watch with a 64bit CPU they just recompiled all the apps.
If they come out with new watches, then they can re-compile all of the code for the new watches with no developer involvement. It is really the best solution for all involved.
The upside is your program may get updated, accelerated routines if it is dynamically linked to a library that you update. The downside is the calls are always indirect via the PLT, which isn't very efficient. It is suitable for things like block encryption and compression where the function entry latency is not very large compared to how long the function runs. It is not very suitable for calls that may be extremely short, like memcmp.
p { maxwidth: 1000px, text-align: center, margin-left: auto, margin-right: auto }
body { text-align: center }
MUCH easier (code snippets = toast)
This is surprising to me. Running any workload that's not your own will trash the CPU caches and will make your workload slower.
Consider for example your performance sensitive code has nothing to do for the next 500 microseconds. If the core runs some other best effort work, it will trash the CPU caches, so that after that 500 microseconds, even when that other work is immediately preempted by the kernel, your performance sensitive code is now dealing with a cold cache.
caches, pipelining, branch prediction, memory protections, SIMD, floating point at all, hyper threading, multi-core, needing cooling fins or let alone active cooling
I wonder how much I've forgotten
Hyper threading is the idea where the multiplexing of multiple threads happens dynamically based on data dependencies / stalls.
HT as an implementation tries to give long sequential execution windows to a hardware thread, it's designed to hide occasional long latencies (IO typically), but pays by having a higher switching cost.
A barrel processor trades off individual instruction latency for a higher degree of parallelism and improved latency hiding.
Prior to 1994 the CPU-memory speed delta wasn't so bad that you needed to cover for stalled execution units constantly. Looking at the core clock vs FSB of 1994 Intel chips is a great throwback! [1] Then CPU speed exploded relative to memory, as was probably anticipated by forward looking CPU architects in 1994.
With slow memory there are a few obvious changes you make to the degree you need to cover for load stalls: 1) OoO execution 2) data prefetching 3) find other computation (that likely has its own memory stalls) to interleave. On the thread level is a pretty obvious grain to interleave work, if deeply non-trivial to actually implement.
Performance oriented programmers have always had to think about memory access patterns. Not new since the 80s to need to be friendly to your architecture there.
Segmented memory was a 16-bit processor trying to access a 32-bit space (good grief, I forgot the details).
You'd have "segment registers", such as "ds" or "es". Reading/writing to memory was by default to the "ds" segment (a 64kB region). If you wanted to read/write to a different segment, you needed to make sure that es, or fs, or one of the other segment registers were available.
If you look at old Win32 code, you might see "far pointer" and "near pointers". "far pointers" were 32-bit pointers that would load the appropriate segment register, while "near pointers" were 16-bit pointers that only made sense within certain contexts.
It was a nightmare. "Flat Mode" eventually became the norm when processors had enough space to just store full sized 32-bit pointers (and later, 64-bit pointers). "Flat mode" made everything easier and now we can forget that those dark ages ever existed...
If you had to have just 16 address bits, having code (CS), stack (SS), data (DS), extra (ES), etc. segments was actually pretty nice. Memory copying and scanning operations were natural without needing to swap bank in the innermost loop.
Of course if you could afford 32-bit addressing, there's no comparison. Flat memory space is the best option, but I don't think it came for free.
Though probably few of those things are technically new since the 1980s.
I remember seeing an add for Power which advertised memory speed at half the processor speed. I think only consumer CPUs have slow memories.
Ok... it also has things like a built-in FPU, multiple cores, cache, pipelining, branch prediction, more address space, more registers, and manufactured with a significantly better (and smaller) process.