How many registers does an x86-64 CPU have?
blog.yossarian.net
blog.yossarian.net
* 16 general-purpose registers.
* 16 or 32 vector registers.
* 7 vector mask registers
* 8 x87 floating-point registers, with 8 MMX registers aliased. (To separate or not separate x87/mmx is definitely a challenging question)
* 3 normal status registers: RIP, RFLAGS, MXCSR
* 6 x87 status registers: FSW, FCW, FTP, FDP, FIP, FOP
* 6 segment registers
* 6 debug registers
* Relevant MPX registers (I don't know this ISA extension very well, so I can't count these registers accurately)
These are the registers that I would expect to be able to poke at in a debugger or inspect/modify via something like ucontext_t, and they're going to be found in whatever kernel abstraction you use to save not-currently-running thread information.
Contrast this with a hypothetical cpu that has only one register "base" and allows to address 32 words after the address at that base register. i.e. things like arithmetic instructions would have 5 bits to address operands which will be interpreters as base+8*n. To make things even more interesting this architecture's instruction pointer lives at base+0.
Such an architecture would have one register under your metric (as only one register needs to be saved/restored to context switch an entire "register file").
However, implementations (microarchitecture) could actually shadow that memory range into hardware registers, and page in/out the whole register bank upon writes of the base register (effectively performing a hardware assisted context switch; hello TSS).
However, since each instruction in this hypothetical ISA must have enough space in the encoding to address these operands, for all intents and purposes this architecture would have 32 registers.
Deciding instructions, addressing operands, dealing with consequences of code density (icache misses), ... are all way more frequent events than context switches.
Hence I do agree with TFA that operand encoding should be the default metric to count registers. And this also includes sub/overlapping registers, if they are independently addressed.
The broader point, though, is in deciding whether or not to include registers like CR0 and DR0. The principle I'm using here is that registers that are not expected to be saved/restored on task switches should be excluded. Registers that are per-process (i.e., page tables in general, or segment descriptors on x86) or per-CPU (most MSRs) are thus excluded by this criterion.
FSBASE/GSBASE are extremely borderline--I wouldn't complain if they were or if they weren't excluded from a list of registers. These act as a mixture of user-visible registers (even if accessibly only via syscalls until very recently) and segment descriptor information. They're not in Linux's userspace-visible mcontext_t struct, but they are in the kernel's equivalent to mcontext_t.
This, er, wasn't really hypothetical. The TMS9900, the CPU used in the TI-99/4A, had three hardware registers: a program counter, a status word, and what was called a Workspace Pointer (WP). General purpose "registers" lived in RAM, and were referenced by an offset off the value in the WP. Subroutine calls were initiated by saving the PC and changing the WP to a fresh new register context before branching.
I really think this article would have been more interesting if he had discussed the microarchitectural registers. Those are important to understand for optimization even if they're not directly visible.
Also discussing MSRs but not special information tables like the page table and VMCS is a somewhat odd distinction. While they're probably stored differently, they are somewhat similar in how they are used.
Also, isn't the TLB like a kind of set of registers? The TLB entries are very frequently accessed. How about store buffers and the like?
Are there any x86 profiling tools which give any metrics about the real utilisation of the register file?
The reason it is less relevant for integer computations is that integer ops have normally lower latency and tend to have shorter loop carried dependency chains.
Like a mode switch bit in a CR register. So MSR-r are just the interface. And the MSR register access can be "slow", so no synchronization or optimization required.
But, the idea of more register makes better architecture, is a total bad assumption. See the dead body of Itanium (128 general-purpose 64 bit integer registers, 128 floating point registers etc. )
With multitasking, one have to switch between context, and larger context (register file size) takes more time.
There are cases when you are better using just the GPRs, rather than the SIMD registers. (Linux kernel does not use FPU or SIMD registers)
Also SIMD usage may slow down the clock, like AVX in x86_64. So you may trust your compiler for vectorization, but it may make more harm than good.
Sparc chips got around that by having sliding windows of registers: instead of having to push all the registers to the stack you just moved the window.
So, instead of a weird but very fast CPU, it ended up being not very fast both in x86 and native modes, while still being weird. (The makers of the Cell CPU did not compromise, went full weird, and had a winner of sorts.)
Maybe count how many bits of registers there are. Then count RAX as 64 bits of registers
Not saying that it's not interesting to know how much actual storage the register file offers; just highlighting that TFA focuses on the instruction encoding angle of the question, which is also important.
CPU architectures are masterpieces of tradeoffs.
Put too many registers and your instructions steam is not dense enough and you cannot keep your cpu busy due to stalls in the fetch phase. Also context switches become expensive (there are solutions to that though).
Put to few registers and you have to spill registers to memory too often, and thus also consume precious instruction stream space.
I disagree strongly with that characterisation. Just no.
* Both x86 and other ISAs have registers that can't be stored to at all, like `k0` for the constant opmask and a whole bunch of read-only MSRs. But they can be read from wholly independently and as discrete registers.
* There are lots of cases where registers can be programmed to clobber other registers, particularly in the performance counter and PAT MSRs.
Sub-registers are definitely a stretch from the above, since they explicitly share bits in the x86 model. But then again, even the x86 model exposed to assembly programmers is a lie: the underlying microcode dynamically renames a large arena of anonymous registers at runtime, and subregisters like AL and AH have been separated in the microcode (to avoid some cases of partial register stalls) for over a decade.
That's not what most people care about when they talk about how many registers a processor has.
I think most people do consider the instruction pointer and status word to be registers, despite also violating the constraints you specified.
Most people probably don't think about MSRs at all, and so maybe just aren't interested in a count of them. But I'm interested in counting the different pieces of on-CPU state that would be necessary to faithfully model an entire x86 core, and both Intel and AMD refer to those bits of state as "registers."
* How many bytes of register files does the OS have to save on task switches? (With this question, subregisters shouldn't be counted.)
* How many registers does an emulator (such as Rosetta 2) have to implement and test? (Subregisters should be counted.)
Even these one might argue aren't directly useful; when considering context switching, one could dig down further into how much of the context switching time is attributable to saving the registers, validate that with experiments across architectures, etc.
it still shouldn't because they can't be distinct in the emulator as writing on a subregister affects the larger register and viceversa.
There is precedent for all this too.
Take the 6809, having 8 bit A and B accumulators. These can be addressed as D, a 16 bit register.
D is not generally counted as it's own register because the D addressing does not point to anything new.
A and B are counted as two registers.
If somehow D brought new bits, say it was 24 bits long, then it would need to be counted.
However if I were doing emulation, I would have to count it because it is a register case that has to be addressed. Just like the D register has to be dealt with on a 6809.
My own personal way to resolve this has always been to determine whether or not a given register specification, that can be addressed somehow, brings new information not contain in any other register specification, to the table.
Fact is, CPU designers do all kinds of crazy things with registers. They overlap, they may be indirect, like not directly addressable, but still there as consideration for the programmer.
Think about something like the REP instruction found on some CPUs. There's a little circuit it keeps track of account, and some rules, and account may be a register that may or may not be directly addressable any other way.
My general take on this article is, "wow, that's a lot of registers!"
We can all quibble about what the quantity of a lot is, and it's all good fun, and I don't think it means anything really.
"Rosetta is a translation process"
https://developer.apple.com/documentation/apple_silicon/abou...
This is not the case for Rosetta 2, which does a translation pass before running the generated ARM binary. No x86 code is run.
This is something that has been bothering me for some time now - actually since the mid-80's: why not implement multiple contexts as an index into a large register file? This way, a context switch would take the time it takes to write to the `task-id` register. It will impact latencies, but would the impact of having, say, 8 contexts not be smaller than having to hit L1 or L2 for the same data?
I would assume that, instead, a modern CPU tags decoded instructions in the reorder buffer with the virtual core number (and register set) it should be applied to. This way, the parallelism would be much easier to exploit.
http://archive.gamedev.net/archive/reference/articles/articl...
I plan to do an update to the post this afternoon.
And while counting architectural registers is something, in reality modern out-of-order processors do something called "register renaming" so that more registers can be used on-the-fly as it dynamically creates a data flow dependency graph. Yes, inside each processor.
Also, if it’s in the silicon and there’s a bug in the algorithm, a microcode update can fix it for everyone. If the bug was in the compiler, you’d need to recompile everything to fix it.
It's only noteworthy "feature" was that it was able to ship ARMv8 support before ARM had a proper ARMv8 CPU design. Time to market for a new ISA was fast, but that's about it.
There exists both 16 bit protected mode (available since the 80286) and 32 bit protected mode (available since the 80386).
Once you remove a single opcode, it's not really x86 any more, and you need emulation at OS level. But then, once you've done that, why not remove more instructions? Why not remove all of them and start again on a much more power-efficient platform? Why not remove the memory model?
One of the brilliant ideas in the A1 is a flag for whether the current process insists on the slower but more comprehensible x86 memory ordering.
You can e.g. remove all MMX instructions and signal this fact through CPUID. Very few apps require MMX (as in can't work without it).
Another option is to remove the instructions from hardware and emulate them in microcode.
That said, the MMX instructions in particular are so problematic to use (and SSE ubiquitous and strictly better) that I suspect you could introduce a processor that lacks MMX support and break almost nobody, certainly far fewer people than removing x87. I don't know if there is an implicit or explicit actual dependency on MMX anywhere.
I had similar issues with PPC software that just blindly assumed I had Altivec on my G3. It wasn't fun.
Removing the 16-bit and 32-bit modes don't actually remove any instructions from the platform (save the binary-coded decimal instructions)--you're largely saving only a few bits of decoder table entries at best. Furthermore, processors reset into 16-bit mode on startup for compatibility reasons, so killing 16-bit and 32-bit mode would introduce major compatibility headaches.
ISA extensions can be more easily removed since there's already a CPUID bit that tells operating systems and applications whether or not they are used. The MPX extension for bounds checking is now regarded as a mistake, and Intel has already confirmed that they are removing it from future processor generations. The TSX extension for transactional memory is apparently on the hit list because of Spectre, and was removed from some processor generations.
The only significant processor execution unit space that is truly obsolete I can think is the x87 floating-point execution unit logic, with the concomitant MMX execution unit logic--SSE is just strictly better for everything here, except if you're trying to actually get the 80-bit precision. But the existence of 80-bit floating point in the 64-bit ABI (i.e., long double) means you'd have a hard ABI break that would potentially break software even written today, and the pain of breaking that ABI is probably not worth whatever savings you get out of it.
Didn't the switch to UEFI effectively reduce the scope of this problem to only apply to motherboard firmware? Operating systems no longer need 16-bit code to boot.
https://en.wikipedia.org/wiki/TOP500
> As of November 2020, all supercomputers on TOP500 are 64-bit, mostly based on CPUs using the x86-64 instruction set architecture (of which 459 are Intel EMT64-based and 22 are AMD AMD64-based. The few exceptions are all based on RISC architectures). Thirteen supercomputers, including the no 2. and no. 3 are based on the Power ISA used by IBM POWER microprocessors, three on Fujitsu-designed SPARC64 chips. One computer uses another non-US design, the Japanese PEZY-SC (based on the British ARM[8]) as an accelerator paired with Intel's Xeon.
There are non-x86 architectures in the TOP500, including ones which have less cruft than x86, but the x86 chips keep on being used in some of the fastest machines on the planet. My hypothesis is that x86 cruft doesn't really matter, and you'd need to go to a cruft level that was orders-of-magnitude worse for the ISA choice to dominate performance.
Intel's Pentium processors (the original ones) were doing pretty badly because they made the pipeline too deep at the expense of other things.
You are probably thinking of the Pentium 4 which was designed as a speed demon with a very deep pipeline and failed to reach its target frequency.
Word and Autocad also run on the Mac!
It's always at the expense of something. Transistor budget is fixed.
Also, take into account the CPUs are not always the more expensive part of the compute node - GPUs, HBM, lots of DDR4, and fast networking gear are also pretty expensive and will be more or less constant as you change CPU architectures.
The vast majority of CPUs shipped are not x86 compatible.
Statista claims 23.5 billion microcontrollers are shipped annually.
I know microchip (the PIC people) made press releases roughly annually as they shipped another billion flash microcontrollers. Google found the one from 2011 when they shipped their tenth billion PIC chip.
I find it difficult to get x86 sales figures. Intel gross revenue is high because they have their fingers in everything. AMD financial statements claim about $2B/quarter total revenue, so if you figure the average shipped price of a AMD cpu is $200 and they made all their revenue off CPUs, that would be 40 million CPUs shipped per year, which seems both ridiculously high AND about a 25th the quantity of microchip PICs shipped.
One way to look at the number of ARM CPUs shipped is the licensing / holding company has made enough licensing fees to pay for about "a hundred and fifty billion" ARM chips in its lifetime.
From programmer’s perspective, the rest of them are either different ways to access these 32-34, or very exotic and rarely used.
P.S. Modern compilers don’t normally emit x87 nor MMX instructions, because they are often slower compared to SSE (all AMD64 processors are required to support at least SSE1 and SSE2). For instance, FSQRT on Zen2 has 22 cycles both latency and throughput, VSQRTPD has 20 cycles latency and 8.5 cycles throughput (lower is better), despite taking 4 square roots in one shot. I think it’s safe to assume x87 and MMX instructions are only left for backward compatibility with old 32-bit binaries, when writing new code they can be ignored.