X86 Register Encoding
eklitzke.org
eklitzke.org
Partially as a consequence of this, the REX prefixes take up a lot of space in most x86-64 instruction streams. In fact, the average size of each instruction is almost exactly 4 bytes, exactly the same as in classic 32-bit RISC architectures. (This is why I dislike it when people link to that old Linus post about how x86 is better than RISC architectures because of code size; it may have been true then, but not now.)
Lea too if your arithmetic fits in it... but then that has issue restrictions... Three operand and the various conditional stuff in arm64 is worth more overall.
Edit: Ah well, beaten by your edit. :)
The greedy register allocator uses the CostPerUse hints to prefer the lower registers, but it has to balance many optimization goals at once, so it is not always obvious what is going on. For example, it will:
- Defer using a callee-saved register for the first time until it is worth the cost of spilling it in the prologue. It won't create a prolog spill just to use a better register. - Try to reuse function argument and return value registers to minimize register copies around function calls.
IIRC, the CostPerUse attribute gave a code size savings of around 1% on x86-64 and Thumb2. On both architectures, the effect is reduced by the fact that all register operands must be from the low register set before you can use the smaller encoding.
I was curious about this claim so I wrote a quick and dirty perl script to test it out:
vmlinuz-linux = 2.71957329365681
bash = 3.95324321387071
firefox = 3.56640365053712
geany = 3.44776119402985
gzip = 4.17925462998247
perl = 4.10894941634241
python = 3.82481751824818
tar = 3.93348845041101
thunar = 3.82471016115786
radeon_drv.so = 4.02697782644402
Seems spot on, except for the kernel. I'm not sure why that has shorter instructions on average.EDIT: I ran it on git too; average was 3.96747336803081. It seems Linus' x86 wizardry does not extend to userspace in this case.
Those instructions tend to be larger, so it may partly explain the difference in average instruction size. The same explanation would apply to why the Radeon driver has a higher average.
It calls "objdump -d" on the binary, which is binutil's disassembler.
objdump prints each instruction out in the following format:
5dc: 67 80 7d 00 00 cmpb $0x0,0x0(%ebp)
The first part is the instruction offset; the second is the bytes that comprise the instruction; the third part is the instruction in AT&T assembly syntax.The script uses regexs to cut off the front and back, then counts the number of bytes (in hex format) in the second part.
Then it calculates the average for all instructions.
Maybe I'm missing something obvious, but why is code size a big concern for Facebook's server code?
http://arstechnica.com/business/2012/04/exclusive-a-behind-t...
Which is not the case at all. Those REX prefix bytes used to be perfectly good 32-bit x86 instructions that now simply don't work in 64-bit mode with their original encodings. So the "compatibility" between 32-bit and 64-bit modes is mythical -- the Opteron could have had a nice shiny new 64-bit programming model that was far less confusing than the dog's breakfast that is x86-64, but just didn't.
Kind of strange, nowadays, an embedded armv8 cpu will come with multiple decoders (ARM32, thumb2, AARCH64).
That is the one advantage that i think ARM has over x86/amd64. That ARM encoding is much shorter.
Presumably because it doesn't matter that much, so it isn't worth the investment for Intel.
It feels to me like people have trouble accepting that both of these are true: (1) RISC vs. x86 doesn't matter that much in practice; (2) technically speaking, RISC is a superior design to x86.
Intel invested in a new instruction set (IA64). AMD did not.
The market chose AMD's offering.
The lesson I take from history is that there is no need to maintain ISA compatibility as long as the old mode is accessible on a process level.
The secret is that Apple controls their hardware/appstore and android apps are ISA independent.
Didn't the AMD64 architecture spank Itanium in the marketplace long before the iPhone became a big deal?
It wasn't Intel but AMD that came up with x86-64. Intel (and HP) indeed wanted to give you a clean 64-bit design - Itanium. And yes, it was meant as a barrier to entry among other things. But AMD won because Itanic didn't run existing x86 software and generally turned out to be slower than Intel hyped.
As for the internal RISC, you won't get it because it changes from generation to generation (also between vendors) and these changes are part of why x86 keeps getting faster. Furthermore, AFAIK these internal ops aren't as dense as x86 to make them simpler to decode and fetching them from memory would make bandwidth more of a bottleneck.
"The C calling convention on x86 systems specifies that callees need to save certain registers."
Practically speaking is this the prologue that that C run time - crt0.o provides automatically/implicitly?
crt0.o is the glue between how the kernel loads a program into memory and how main() expects things to work. it does basic setup tasks that can vary between platforms, but is generally things like collecting command line arguments and setting up the stack. it will also invoke exit() if main() returns, since that's how the kernel expects the process to be destroyed
When you say "how the kernel loads a program into memory" I assume you are referring to ld-linux.so.2? Is that correct?
I imagine then that ld-linux.so.2 calls __start in crt0.o and crt0.o jumps to main(). Is this correct?
maps the executable to memory, and if an elf interpreter (/lib/ld-linux.so) is specified (which it normally is) also the elf interpreter.
It then jumps to the entry point of the raw binary or of the elf interpreter.
I think your "normally" means "in case of dynamically linked executables".
Its easy to sometimes conceptually think or talk about "a loader" as if its some standalone entity. However this can be misleading(at least to me anyway.) Because in reality its linked into every dynamically-linked binary, along with crt0.o which as the other poster mentioned takes care of some ABI requirements and the setting up of the stack.
For anyone else who might be interested this also an illuminating source code file to read, libc-start. Which would be part of crt0.o
http://repo.or.cz/glibc.git/blob/HEAD:/csu/libc-start.c#l105
I don't have to think this low level on any kind of regular basis but its a great thought exercise to do so from time to time.