The Art of Picking Intel Registers (2003)
swansontec.com
swansontec.com
This usually results in degradation performance-wise, but has its uses.
Each register has its own personality and certain things it just does best. It is a little sad to see that RISC has "won" and modern CPU's have dozens or hundreds of registers which are just numbers.
The classic example of where register-poor ISAs hurt isn't about "subtle differences" at all. It's the fact that there are still only 8 (for i386) named registers, and so anything that needs to deal with a working set beyond that needs to do spill/fill to memory, and you can't "rename" memory accesses (though you can sort of cheat, as with the store forward optimization -- but that doesn't work nearly so well as renaming does).
Everything else is general memory optimizations that apply for everything like the aforementioned store forwarding. It's still expensive if the CPU can't use them (mismatched load/store size, incorrect speculation, etc.)
The only weird thing there is the compiler lucks out in generating a version that allows OoOE to save the programmer from themself.
In effect, the compiler was asked to compile
void foo(float &resVal, float &val1, float &val2, float &val3, float &val4, float &val5, float &val6, float &val7) {
float tmp = val1+val2+val3+val4+val5+val6+val7;
resVal += tmp;
}
whereas the author wrote asm doing void foo(float resVal, float val1, float val2, float val3, float val4, float val5, float val6, float val7) {
resVal += val1;
resVal += val2;
resVal += val3;
resVal += val4;
resVal += val5;
resVal += val6;
resVal += val7;
}
Which isn't the same thing. And if the author unrolled it with a hundred registers, it would still be every bit as slow.And it should be obvious that it's slow not because it's messing up OoOE, it's slow because it's one incredibly long dependency chain with a stall between each add, due to it depending on the result of the previous add. OoOE happens to be able to fix the programmer's mistake for the first case, which is luck because the programmer obviously didn't intend for it to do so.
But now I'm repeating myself...
Well, I'm assuming the programmer wrote it like that, because otherwise a smart compiler would have calculated val1+...+val7 before the loop and just added that precalculated value repeatedly.
Also it doesn't matter if it did keep resVal (and val1..val7) in a register, it would be every bit as fast as it is now.
Similarly, the compiler has no leeway about the order of floating point additions. If the C++ was written to have identical output as the asm, the compiler could not be faster than the asm without violating the C++ standard.
Again, OoOE plays absolutely positively no part whatsoever in making the asm slow. Neither does which values are in registers or memory.
It's slow because a chain of instructions that use the output of the previous as input cannot execute in fewer cycles than (latency)*(num instructions). Period.
It's the sort of thing that should have been written by the programmer as something like
res += (val1+val2)+(val3+val4);
res += (val5+val6)+(val7);
then we wouldn't be having this discussion...Having participated in the writing of an x86 code generator for a compiler, I wish to thank the author for a much-needed laugh from the above sentence.
Then you can copy-paste the assembly language from the article (I've changed the spacing for readability and added dummy definitions for the names so it will compile):
;demo1.asm
source_address equ 0x100
destination_address equ 0x200
loop_count equ 0x10
mov esi, source_address
mov edi, destination_address
mov ecx, loop_count
my_loop:
lodsd
;Do some calculations with eax here.
stosd
loop my_loop
And assemble it, in your favorite shell type: nasm -l demo1.lst demo1.asm
Then here's my alternative implementation that uses different registers. You can no longer use LODSD, STOSD or LOOP instructions since these instructions only work if you chose the same registers as demo1. ;demo2.asm
source_address equ 0x100
destination_address equ 0x200
loop_count equ 0x10
mov ebx, source_address
mov edx, destination_address
mov esi, loop_count
my_loop:
mov eax,[ebx] ; these two instructions instead of stosd
add ebx,4
;Do some calculations with eax here.
mov [edx],eax ; these two instructions instead of lodsd
add edx,4
dec esi ; these two instructions instead of loop
jnz my_loop
I get a demo1 of 24 bytes and a demo2 of 38 bytes. The demo1.lst and demo2.lst files produced show how many bytes are taken up by each instruction. (And if you get addresses in a crash dump, they can be used to track down the corresponding source code line.)If you want to actually run these programs, nasm's default output (raw machine language instructions) cannot be used by most OS's. (In DOS, you can -- just rename to .COM. But a DOS target needs to tell NASM 'bits 16', to have it emit the proper prefixes for those new-fangled 32-bit instructions, and will crash without the DOS exit syscall, INT 0x20.) The magic incantations for standalone Linux assembly language programs are here:
http://blog.markloiseau.com/2012/04/hello-world-nasm-linux/
The .o file produced by an intermediate step of the instructions at the above link can be linked with C code (unless you use a Microsoft toolchain, in which case you have to instruct nasm to output .obj instead). Then you can call assembly language functions from C and vice versa. (Figuring out how to retrieve your function's arguments in assembly language is very interesting and will enlighten you about the implementation of high-level languages.) Most "real-world" assembly code does this: Most of the program is written in C, and only the functions that need the particular advantages of assembly language are written in it.
http://www.posix.nl/linuxassembly/nasmdochtml/nasmdoca.html
From this you can see that many arithmetic and logic instructions have special encodings that use AL/AX/EAX as the destination. (You can also use the more general encoding to produce exactly the same instruction, but in more bytes.)
If you're working with AL/AX/EAX, then if you arrange for ESI and EDI to point into your source and dest buffers respectively, you can use LODS and STOS to fetch and store bytes. If your processing is working in a streaming-like fashion (and this isn't a bad idea, if you can arrange it) then you can fetch or store a bytes or dword, and bump the pointer, using one instruction that's 1 byte. Doing this with a MOV followed by an INC would be 2 instrutions and 3 bytes.
There is no short form for the [EBP] addressing mode, nor [EBP+reg], because you virtually always use fixed offsets with a base pointer. I think both of these end up as [EBP+0] or [EBP+reg+0], so you'll alwas waste an extra byte for the fixed displacement if you try to use EBP for something more general.
The rotate instructions have special encodings that take the rotate count in CL, since C is for count after all.
There's probably also a few more bits and pieces like the above, which was just what stuck in my mind from when I wrote code to generate x86 machine code directly. I just skimmed through the NASM doc to double check my memory.
I couldn't say whether using the shorter forms makes any realistic difference on real code these days, aside from the general 'fewer bytes = probably quicker' rule. But If you're trying to make EXEs that are 4K or less, as this guy seems to, I bet the little savings would add up.
At any rate, hopefully, one day soon. We will move away from the x86 and its legacy. While some people believe that architecture at this level doesn't matter... and they are mostly right... the x86 is still ugly.