How many x86 instructions are there? (2016)
fgiesen.wordpress.com
fgiesen.wordpress.com
[1] https://github.com/Battelle/sandsifter
Hopefully this will have either saved you a click or validated your time in reading the article.
> Does a non-instruction that is non-defined and unofficially guaranteed to non-execute exactly as if it had never been in the instruction set to begin with count as an x86 instruction? For that matter, does UD2 itself, the defined undefined instruction, count as an instruction?
While having instructions for everything that are slow in early models but can be significantly improved in silicon over time is one way to look at CISC, I genuinely wonder how much silicon is spent on instructions that are so rarely used they'd be better in software.
Or to ask another way: how many instructions are in billions of x86 cores that rarely if ever get used? Hmmm...
Of course, that’s not an interesting observation.
It gets (somewhat) interesting when you realize that one may inadvertently include in-line assembly. Do you call including system headers that happen to contain inline assembly cheating? Including headers for a standard C library? Compiling with old-fashioned compiler flags (for example to generate FPU instructions for the x87)?
$ objdump -w -j .text --no-show-raw-insn -d /usr/bin/emacs-gtk | egrep '^ *[0-9]+:' | awk '{print $2}' | sort | uniq | wc -l
130
$ objdump -w -j .text --no-show-raw-insn -d /bin/ls | egrep '^ *[0-9]+:' | awk '{print $2}' | sort | uniq | wc -l
97
$ objdump -w -j .text --no-show-raw-insn -d firefox-bin | egrep '^ *[0-9]+:' | awk '{print $2}' | sort | uniq | wc -l
136I wonder if the numbers would change significantly for gentoo or anything compiled manually that had more knowledge of the CPU specifics?
$ objdump -w -j .text --no-show-raw-insn -d /usr/lib64/firefox/firefox-bin | egrep '^ \*[0-9]+:' | awk '{print $2}' | sort | uniq | wc -l
134
So actually, less!(granted, that’s only 32/64 bits, but still…)
I have a blog post about coping strategies for working around the absence of PMOVMSKB on NEON:
https://branchfree.org/2019/04/01/fitting-my-head-through-th...
We used these techniques in simdjson (which I presume still uses them; the code has changed considerably since I built this): https://github.com/simdjson/simdjson
The best techniques for mitigating the absence of PMOVMSKB require that you use LD4, which results in interleaved inputs. This can sometimes make things easier, sometimes harder for your underlying lexing algorithm - sadly, it's not a 1:1 transformation of the original x86 code.
I'm somewhat curmudgeonly w.r.t. SVE, insisting that while the sole system in existence is a HPC machine from Fujitsu, that for practical purposes it doesn't really exist and isn't worth learning. I will likely revise this opinion when ARM vendors decide to ship something (likely soon, by most roadmaps). There's only so much space in my brain.
AVX-512's masks are OK. They're quite cheap. There are some infelicities. I was irate to discover that you can't do logic ops on 8b/16b lanes with masking; as usual the 32b/64b mafia strike again. This may be a symptom of AVX-512's origin with Knights*.
It would be nice if the explicit mask operations were cheaper. Unfortunately, they crowd out SIMD operations. I suppose this is inevitable given that they need to have physical proximity to their units - so explicit mask ops are on the same ports as the SIMD ops.
I also wish that there were 512b compares that produced zmm registers like the old compares used to; sometimes that's the behavior you want. However, you can reconstruct that in another cheap operation iirc.
Fair enough. I have high hopes for SVE, though. The first-faulting memory ops and predicate bisection features look like a vectorization godsend.
> There's only so much space in my brain.
I'm still going to attempt a nerd-sniping with the published architecture manual. Fujitsu includes a detailed pipeline description including instruction latencies. Granted its just one part, and its an HPC-focused part at that. But its not every day that this level of detail gets published in the ARM world.
https://github.com/fujitsu/A64FX/tree/master/doc
> I was irate to discover that you can't do logic ops on 8b/16b lanes with masking; as usual the 32b/64b mafia strike again.
SVE is blessedly uniform in this regard.
> It would be nice if the explicit mask operations were cheaper. Unfortunately, they crowd out SIMD operations.
This goes both ways, though. A64FX has two vector execution pipelines and one dedicated predicate execution pipeline. Since the vector pipelines cannot execute predicate ops, I expect it is not difficult to construct cases where code gets starved for predicate execution resources.
I salute your dedication to nerd-sniping. I need my creature comforts these days too much to spend days out there in the nerd-ghillie-suit waiting for that one perfect nerd-shot. That may be stretching the analogy hopelessly, but working with just architecture manuals and simulators is tough.
I am more aiming for nerd-artillery ("flatten the entire battlefield") these days: as my powers wane, I'm hoping that my superoptimizer picks up the slack. Despite my skepticism about SVE, I will retarget the superoptimizer to generate SVE/SVE2.
FMA4 https://en.wikipedia.org/wiki/FMA_instruction_set#FMA4_instr...
XOP https://en.wikipedia.org/wiki/XOP_instruction_set
TBM https://en.wikipedia.org/wiki/Bit_manipulation_instruction_s...
Indeed, those who think x86 is complex should also take a detailed look at the ARM64 instruction set, particularly its instruction encoding. If you thought making sense of x86 instruction encoding was hard, and that a RISC might seem simpler, AArch64 will puzzle you even more.
To use the MOV example, the closest ARM equivalent might be the 40 variants of LD, which the reference manual (5000+ pages) enumerates as: LDAR, LDARB, LDARH, LDAXP, LDAXR, LDAXRB, LDAXRH, LDNP, LDP, LDPSW, LDR (immediate), LDR (literal), LDR (register), LDRB (immediate), LDRB (register), LDRH (immediate), LDRH (register), LDRSB (immediate), LDRSB (register), LDRSH (immediate), LDRSH (register), LDRSW (immediate), LDRSW (literal), LDRSW (register), LDTR, LDTRB, LDTRH, LDTRSB, LDTRSH, LDTRSW, LDUR, LDURB, LDURH, LDURSB, LDURSH, LDURSW, LDXP, LDXR, LDXRB, LDXRH. Some, like LDP, are then further split into different encodings depending on the addressing mode.
My suspicion is that to achieve acceptable code density with a fixed-length instruction encoding, they just made the individual instructions more complex. For example, the add instruction can also do a shift on one of its operands, which would require a second instruction on x86.
I think this is why assembler can be faster many times. Not because I'm better than a compiler. But because the structure of the language nudges you into faster approaches.
It's been a problem (optimizing) for some time though. I remember it being some work to beat the compiler on the i960CA. OTOH, I seem to remember the i860 being not-so-great and for sure the TI C80 C compiler was downright awful (per usual for DSPs).
And it changes so often, instructions that are fast on one CPU are not so fast on the next one, and vice versa.
Not to mention branch prediction and out-of-order execution makes it very difficult to meaningfully benchmark. Is my code really faster, or just seems like it because some address got better aligned or similar.
I've gotten significant speed gains in certain projects by simply replacing certain hand-optimized assembly in libraries (ie not my code) with the plain C code equivalent. The assembly was probably faster 10-15 years ago, but not anymore...
That's an interesting point, plus there's the portability issue.
My own breadcrumbs of legacy code for this kind of innerloopish stuff has been to write a straightforward 'C' implementation (and time it), an optimized 'C' version (which itself can depend on the processor used), and a handtuned assembly version where really needed.
It allows you to back out of the tricky stuff plus acts as a form of documentation.
Also, the programmer can "cheat" by doing things the compiler would consider invalid but are known to be ok given the larger context of the application.
The restrict keyword in c gives an example of how one can optimize code by knowing the larger context. https://cellperformance.beyond3d.com/articles/2006/05/demyst...
The problem is the ROI is usually pretty bad as these assumptions rarely hold as the code evolves, in my experience, and the optimization usually only lasts for finite (sometimes shockingly short) amount of time. i.e. OS changes, hardware changes, memory changes, etc. etc. etc.
[1] https://store.steampowered.com/app/370360/TIS100/?curator_cl...
[2] https://store.steampowered.com/app/504210/SHENZHEN_IO/?curat...
Whenever bored, I read or watch chibiakumas tutorials. There's no such thing as knowing too many assemblers.
ARM has also been willing to drop older optional legacy stuff like Java oriented instructions that almost nobody used and Thumb. X86 supports nearly all legacy opcodes, even things like MMX and other obsolete vector operations that modern programs never use.
I mean ... that's the theory.
Which seems reasonable.. but you just burned 3/4 of the single opcode instruction space, which may not be worth it for most general purpose loads.
For thumb, 32-bit instructions can be allocated up to 3/32 of the potential 32-bit space, and 16-bit instructions can use 29/32 of the 16-bit space (3 of the potential 5-bit opcodes denote a 32-bit instruction.) Which is probably a better ratio than 1/2 or 1/4 of each, for instance. Though I'm not sure how much of that encoding space is actually allocated or still reserved.
Related, I believe ARM has allocated about half of the 32-bit encoding space for current A64 instructions.
Are there examples of ISA or chips that handle variable length instruction encoding efficiently?
Contrast that to something like utf8 where that isn't possible.
I think this old Intel patent talks about one of their decoder implementations:
FWIW, this is exactly what RISC-V does.
There aren't any standard extensions with instructions >32b yet, but the extensibility is there in the base spec.
This complexity is pushed down to operating systems, compilers, assemblers, debuggers. It ends up causing brutal human time overhead throughout the chain, and its cost effectively prevents security and high assurance.
This more than justifies moving away from x86 into RISC architectures, such as the rising open and royalty-free RISC-V.
Of course, these days the Z80 instruction set looks trivial :-)
Is that really true though? In my experience, the amount of complexity in a layered system is roughly constant, and if you make one layer simpler, you have to make another more complex.
I would expect that compiler for a RISC architecture, for example, needs to do a lot of work figuring out which set of 'simple' instructions encode an intent most efficiently. It also needs to deal with instruction scheduling and other issues that a RISC compiler does not care about.
How Many X86-64 Instructions Are There Anyway? - https://news.ycombinator.com/item?id=14233296 - April 2017 (133 comments)
How many x86 instructions are there? - https://news.ycombinator.com/item?id=12358050 - Aug 2016 (39 comments)
Does a compiler use all x86 instructions? (2010) - https://news.ycombinator.com/item?id=12352959 - Aug 2016 (189 comments)
How Many X86-64 Instructions Are There Anyway? - https://news.ycombinator.com/item?id=11535178 - April 2016 (1 comment)
$ ndisasm -b 64 /dev/urandom
I always think of the AMD 29000.
In the modern day it mostly just means that you can implement fancy out of order schenanigans with a bit fewer engineer-years than non-RISC ISAs.
https://dl.acm.org/doi/10.1109/12.4607
"An interrupt is precise if the saved process state corresponds to a sequential model of program execution in which one instruction completes before the next begins. In a pipelined processor, precise interrupts are difficult to implement because an instruction may be initiated before its predecessors have completed."
There hasn't been any new CISC architecture worth mentioning in decades.