A fundamental introduction to x86 assembly programming
nayuki.io
nayuki.io
http://savannah.nongnu.org/projects/pgubook/ (free pdf!)
This is a practical book and teaches assembly programming on Linux.
Author Jonathan Bartlett wrote this book because he was frustrated to no end with the existing books. At the end of them he could still ask, "How does the computer really work?" and not have a good answer. Jonathan's goal is to take you from knowing nothing about programming to understanding how to think, write, and learn like a programmer. You won't know everything, but you will have a background for how everything fits together.
Fun story: I remember how I went through this book in 2004, a day before a job interview, and I exactly got asked a question about how C functions get compiled to assembly, how the stack and memory management works. I got that job.
This book takes a bit of a different angle on explaining computers. I really enjoy the history.
My only complaint would be that I think he has a Microsoft bias (but then I guess I am biased myself).
And similarly, I have a minor annoyance with the OP's mention of Linux, and only Linux, when he touches on calling conventions. This bias toward one system, and ignorance of others, is typical of many websites and documentation. To be fair, the OP mostly avoids it.
To be clear, great explanations of computers to me are ones that either:
1) take great care to stay completely neutral and only discuss universally shared traits across systems,
2) go to great lengths to try to be as comprehensive as possible, including many systems and all their commonalities and idiosyncracies, or
3) focus only on one system and go into great detail how it works.
The more the author strays from 1, 2 or 3, the less likely I am to read their work.Petzold pays ample attention to Morse code and similar succinct ways of communicating information. In my opinion this type of focus is the mark of a skilled coder. When I look at the entries to IOCC, it is no surprise to me that Morse code is (or at least was) a frequent focus of the entrants.
For a quick idea of what ASM can look like if you build up the foundations step-by-step, and understand what you're working with:
.include "record-def.s"
.include "linux.s"
#PURPOSE: This function reads a record from the file descriptor
#
#INPUT: The file descriptor and a buffer
#
#OUTPUT: This function writes the data to the buffer
# and returns a status code.
#
#STACK LOCAL VARIABLES
.equ ST_READ_BUFFER, 8
.equ ST_FILEDES, 12
.section .text
.globl read_record
.type read_record, @function
read_record:
pushl %ebp
movl %esp, %ebp
pushl %ebx
movl ST_FILEDES(%ebp), %ebx
movl ST_READ_BUFFER(%ebp), %ecx
movl $RECORD_SIZE, %edx
movl $SYS_READ, %eax
int $LINUX_SYSCALL
#NOTE - %eax has the return value, which we will give back to our calling program
popl %ebx
movl %ebp, %esp
popl %ebp
ret
With the definitions in place, its like a whole different language.For reference, the definitions are simply a text file with contents similar to:
#System Call Numbers
.equ SYS_EXIT, 1
.equ SYS_READ, 3
.equ SYS_WRITE, 4
.equ SYS_OPEN, 5
.equ SYS_CLOSE, 6
.equ SYS_BRK, 45
Or (record-defs.s in the example) clearly describing the data with: .equ RECORD_FIRSTNAME, 0
.equ RECORD_LASTNAME, 40
.equ RECORD_ADDRESS, 80
.equ RECORD_AGE, 320
.equ RECORD_SIZE, 324
This book opened my eyes to the concrete, data driven nature of the problems I'm trying to solve, at an atomic level.
It somehow dispelled all the magic behind programming, while exciting the mechanical side of my brain, leaving me with that "its just a machine, I can solve any problem if I just trace things patiently until I understand the parts and how they interact".The function should read in a loop, to fulfill its contract (read a record).
#include <unistd.h>
ssize_t read_record(int fd, void* buf, size_t record_size, size_t* written) {
for (*written = 0; *written < record_size; ) {
ssize_t n = read(fd, (char*)buf + *written, record_size - *written);
if (n < 1)
return n; /* error or EOF */
*written += n;
}
return *written;
}Get a TIS-100 game and play it (that got me interested again). After that, I compiled simple(r) C programs with gcc and looked at their .S output. After that (along with that) grab a tool like ollydbg, x64dbg or (if you can!) IDA Pro and open up your favorite programs and modify them. Whenever you stumble upon an unknown (to you) instruction, look it up in the intel manual and google for it to see idioms people use. This process has worked really well for me, for now, albeit it feels like I'm cracking software or something like it (it's fun though). Along with that you can start writing asm blocks in your programming language of choice and/or full asm with any of the assemblers (flat, yasm, nasm, whatever).
Only thing you need to know beforehand are the basics of C and data/memory manipulation.
A simple way to circumvent this problem (I don't claim it is the best) is
sub rsp, 4
mov eax, [rsp]
where eax of course contains the value to push.That said, as long as you keep RSP aligned you can do whatever you want. Consider this code:
extern void value(int* a, int* b, int* c);
int main() {
int a, b, c;
value(&a, &b, &c);
return a+b+c;
}
This is how LLVM compiles it: subq $24, %rsp
leaq 20(%rsp), %rdi
leaq 16(%rsp), %rsi
leaq 12(%rsp), %rdx
callq value
movl 16(%rsp), %eax
addl 20(%rsp), %eax
addl 12(%rsp), %eax
addq $24, %rsp
retq
Note that the int values are allocated at 4-byte alignment, but rsp is aligned to 8-bytes. If you add an additional parameter, 'd', you'll see that the compiler still allocates 24-bytes of stack, and stores the additional parameter at 8(%rsp) (which is unused in the code above).If you want to call C code conforming to the x86-64 SYSV ABI, RSP needs to be aligned to 16 bytes when you execute the call. If the code you generate never calls alien code, 8 byte alignment is enough.
Since 8 bytes are occupied by return address pushed by the call which started your function, you need to decrease RSP by further 8, 24, 40, 56, 72, ... bytes before calling code generated by others.
Reason: having stack 16 byte aligned makes it easier to allocate aligned 16 byte stack variables and this is useful because x86 has 16 byte registers (SSE) which are most efficiently loaded/stored to aligned addresses.
However, it isn't only performance that you lose by neglecting alignment. I learned the hard way that some code generated by gcc crashes if you call it with unaligned stack.
That's why in this example LLVM allocates 24 bytes, even though 16 would be enough for 3 ints.
Another example (gcc):
extern void bar();
void foo() {
bar();
}
0000000000000000 <foo>:
0: 48 83 ec 08 sub $0x8,%rsp
4: b8 00 00 00 00 mov $0x0,%eax
9: e8 00 00 00 00 callq e <foo+0xe>
e: 48 83 c4 08 add $0x8,%rsp
12: c3 retq
To anyone writing x86-64 compilers, I recommend finding the x86-64 SYSV ABI spec and reading it. Saves debugging time.x87 allows to do calculations in extended precision, i.e. word width is 80 bits. SSE is limited to double precision (64-bit words). Also FPU has some advanced math instructions, like sin, cos, tan, exp, etc., and these operations will be available in AVX512F.
People store doubles, and people get upset when compiler optimizations (like when to spill from registers to memory and whether to use one or two instructions for multiply-and-add) change not just the performance but also the result of computations.
But you can, if you are after precision as the OP apparently was.
I'm not sure what this c-word is doing in your post, I thought we were talking assembly, but as far as c-things go, at least gcc and clang represent long double as 80b extended precision on x86.
So, for example, compiling this beauty (which is too large for x87 stack):
long double a[64], b[64], c[64], d[64];
// load a,b,c,d from somewhere
long double x =
((((((((a[0]+a[1])+(a[2]+a[3]))+((a[4]+a[5])+(a[6]+a[7])))
+(((a[8]+a[9])+(a[10]+a[11]))+((a[12]+a[13])+(a[14]+a[15]))))
+((((a[16]+a[17])+(a[18]+a[19]))+((a[20]+a[21])+(a[22]+a[23])))
+(((a[24]+a[25])+(a[26]+a[27]))+((a[28]+a[29])+(a[30]+a[31])))))
+(((((a[32]+a[33])+(a[34]+a[35]))+((a[36]+a[37])+(a[38]+a[39])))
+(((a[40]+a[41])+(a[42]+a[43]))+((a[44]+a[45])+(a[46]+a[47]))))
+((((a[48]+a[49])+(a[50]+a[51]))+((a[52]+a[53])+(a[54]+a[55])))
+(((a[56]+a[57])+(a[58]+a[59]))+((a[60]+a[61])+(a[62]+a[63]))))))
+((((((b[0]+b[1])+(b[2]+b[3]))+((b[4]+b[5])+(b[6]+b[7])))
+(((b[8]+b[9])+(b[10]+b[11]))+((b[12]+b[13])+(b[14]+b[15]))))
+((((b[16]+b[17])+(b[18]+b[19]))+((b[20]+b[21])+(b[22]+b[23])))
+(((b[24]+b[25])+(b[26]+b[27]))+((b[28]+b[29])+(b[30]+b[31])))))
+(((((b[32]+b[33])+(b[34]+b[35]))+((b[36]+b[37])+(b[38]+b[39])))
+(((b[40]+b[41])+(b[42]+b[43]))+((b[44]+b[45])+(b[46]+b[47]))))
+((((b[48]+b[49])+(b[50]+b[51]))+((b[52]+b[53])+(b[54]+b[55])))
+(((b[56]+b[57])+(b[58]+b[59]))+((b[60]+b[61])+(b[62]+b[63])))))))
+(((((((c[0]+c[1])+(c[2]+c[3]))+((c[4]+c[5])+(c[6]+c[7])))
+(((c[8]+c[9])+(c[10]+c[11]))+((c[12]+c[13])+(c[14]+c[15]))))
+((((c[16]+c[17])+(c[18]+c[19]))+((c[20]+c[21])+(c[22]+c[23])))
+(((c[24]+c[25])+(c[26]+c[27]))+((c[28]+c[29])+(c[30]+c[31])))))
+(((((c[32]+c[33])+(c[34]+c[35]))+((c[36]+c[37])+(c[38]+c[39])))
+(((c[40]+c[41])+(c[42]+c[43]))+((c[44]+c[45])+(c[46]+c[47]))))
+((((c[48]+c[49])+(c[50]+c[51]))+((c[52]+c[53])+(c[54]+c[55])))
+(((c[56]+c[57])+(c[58]+c[59]))+((c[60]+c[61])+(c[62]+c[63]))))))
+((((((d[0]+d[1])+(d[2]+d[3]))+((d[4]+d[5])+(d[6]+d[7])))
+(((d[8]+d[9])+(d[10]+d[11]))+((d[12]+d[13])+(d[14]+d[15]))))
+((((d[16]+d[17])+(d[18]+d[19]))+((d[20]+d[21])+(d[22]+d[23])))
+(((d[24]+d[25])+(d[26]+d[27]))+((d[28]+d[29])+(d[30]+d[31])))))
+(((((d[32]+d[33])+(d[34]+d[35]))+((d[36]+d[37])+(d[38]+d[39])))
+(((d[40]+d[41])+(d[42]+d[43]))+((d[44]+d[45])+(d[46]+d[47]))))
+((((d[48]+d[49])+(d[50]+d[51]))+((d[52]+d[53])+(d[54]+d[55])))
+(((d[56]+d[57])+(d[58]+d[59]))+((d[60]+d[61])+(d[62]+d[63]))))))))
;
produces only 80b spills (fstpt in GNU syntax): $ objdump -d fpmonster |grep fst
400457: db 7c 1c 10 fstpt 0x10(%rsp,%rbx,1)
400462: db bc 1c 10 04 00 00 fstpt 0x410(%rsp,%rbx,1)
400470: db bc 1c 10 08 00 00 fstpt 0x810(%rsp,%rbx,1)
40047e: db bc 1c 10 0c 00 00 fstpt 0xc10(%rsp,%rbx,1)
400b47: db 7c 24 10 fstpt 0x10(%rsp)
400d91: db 3c 24 fstpt (%rsp)
Clearly, extended precision can be done right both in C and raw assembly.> People store doubles, and people get upset
That's their fault :) and another story altogether. For reproducible low precision, indeed SSE is the way to go.
The nice thing about this book is that it guides the reader at understanding how the machine works first, and only then to assembly programming.
The sad thing about this book is that it references 32 bit intel-compatible processors.
My guess is that the original author has grown old and is not interested in producing a fourth edition of such book.
On this matter, I would like to ask: is it worth learning assembly for the x86/32-bit instructions, now that pretty much every computer is built on the amd64 architecture ?
"Well, that isn’t completely my decision. When the publisher wants a new edition, they contact me, and then I begin writing. I don’t see a new edition on the horizon for a couple of years yet. Worse, I need additional pages to cover 64 bit issues, and the number of pages I have in the book is limited. Unless I can persuade the publisher to go beyond the 600-page mark, I’m going to have to eliminate other material to cover 64-bit assembly."
[1] http://www.contrapositivediary.com/?page_id=1808#ASMSBS3E
The Cogwheel Brain: Charles Babbage and the Quest to Build the First Computer (Doron Swade) is also an interesting read, especially as Babbage's ideas were eventually vindicated when the Science Museum built it in the 1990s and it worked.
http://www.pentesteracademy.com/course?id=3 http://www.pentesteracademy.com/course?id=7
cmpl $9, -4(%rbp)
jle .L3
under GCC ('if i <= 9, loop') and: cmpl $10, -8(%rbp)
jge .LBB0_4
under clang ('if i >= 10, end the loop'). Neither GCC or clang does an exact conversion, and they both produce different assembly instructions (optimisations disabled in both cases).If you build a cross-compiler, you can also output assembly for architectures other than your local machine, though this can be quite fiddly (see crosstool-ng for a project which has done most of the work for you).
Curiously, it doesn't actually appear all that ineffecient. Go uses it AFAIK. I wonder whether anyone has studied it. I also wonder whether gccgo uses that convention or defaults to SysV x64.
I don't think of it in terms of like-vs-dislike. My observation is that it's a difficult thing to get right without a compiler, and thus avoided for introductory material.
As far as I am aware, storing values in a register requires fewer instructions. However, I have never personally confirmed the performance difference of this calling convention.
Handling of all calling conventions is one of the many things that I am personally much happier leaving up to GCC and LLVM in practice.
intel CPU is just crap from a design standpoint, and since they had to remain backward compatible, it's gotten a lot faster with lots and lots of tricks, but it still sucks in 32-bit mode. No amount of tricks will change that. It has to be run in 64-bit mode to gain a performance boost and simplify the code, whereas processors with fixed 32-bit instruction encoding run faster in 32-bit mode and code simplicity is a constant.
> ... far more elegant than a 32-bit four register intel CPU.
I count eight: EAX, EBX, ECX, EDX, ESI, EDI, EBP, ESP.
x86 has also advantage of supporting arbitrary immediate values, so you don't need to allocate registers just for constants.
When I originally wrote four general purpose registers, I had eax, ebx, ecx and edx in mind, but after fully listing them above, I revise my earlier statement: the x86 assembler has two general purpose registers. Crappy processor architecture with lots of specialized registers, but too few general purpose ones.
Compared to MOS 6502, Motorola MC680##, or SPARC, only the eax and edx are really general purpose registers -- even mentioning ecx, the counter register, would be iffy.
68k is pretty nice, d0-d7 registers are indeed interchangeable. Of course a0-a7 are just for addressing, I think a7 was usually stack pointer.
SPARC I've never programmed, so no comments about it.
I've written x86 code in the past (20 years ago) using all 8 registers for general purpose task -- yes, even ESP. It was faster that way to implement a texture mapper. Ugly but fast.
Those 8 x86 registers are mostly general purpose, apart from some exceptions.
Multiplication was the only annoying one, getting result in EAX:EDX.
I always succeeded making x86 do whatever I wanted, despite some limitations with register use.
I agree that stosb/rep would probably be confined to memory management, and saving/restoring register around such ops isn't the end of the world. Not sure about movsb -- I suppose if you're copying enough data, store/restore is going to be negligible overhead in terms of speed, but if you're actually trying to write clear code, it would certainly be easier to not have to worry about the book keeping?
Currently the fastest way to memset large chunks of memory is probably to use SSE or AVX. I'd guess this is what gets generated if compiler target arch allows.
With SSE/AVX you also have an option to use non-temporal moves to avoid polluting caches. This might have a negative impact on any memset micro-benchmark [1], but significantly help any concurrently executing memory bound CPU cores.
Properly aligned (cache line 64-byte boundary) you might be able to avoid read-for-ownership as well, further reducing memory bus traffic.
So most use of rep-prefix might be pointless, unless you can accept the performance hit.
[1]: Just like micro-benchmarking any other resource constrained operation. Micro-benchmarks can give you very wrong idea of what is best for the system as a whole.
The CPUs have lots of duplicated logic to process many instructions in parallel and, on "friendly" code, can sustain average throughput of 2 or more instructions per clock cycle, provided that the instructions are simple enough.
The end result is that a loop made with normal adds, cmps and jnes outperforms those dedicated looping instructions.
They are only used by compilers when optimizing for code size and maybe by people who want concise hand written assembly, though I'm not sure why wouldn't they just use C in such case.
See "Software Optimization Guides" released by AMD/Intel for more info.
In the sense of instruction encoding they are. Additionally, the x86-64 call convention typically does not use ebp as frame pointer (except when you use something like alloca). Instead functions typically allocate their nessesary stack amount at the begin. So ebp is generally used as a general purpose register by most compilers in x86-64.
- The ARM has a standard MMU (well, two major revs). MIPs has quite a few variants, and they involve TLB invalidation / reload. They're both worth working with.
- ARM has made some interesting architectural choices over the years, and it's worth studying what they've done. MIPs is more a static platform, not as much market pressure or will to innovate.
ARM has a few ecosystem problems that I'd rather not deal with. From what I understand there is a lack of hardware discovery. ARM is more embedded then 'user' computer. Nothing is swapable.
http://www.hugi.scene.org/compo/compoold.htm
The 256B and below categories in the demoscene are also sources of interesting x86 Asm programs:
For mips there is an excellent book called 'see mips run' by dominic-sweetman. you might find it to be quite instructive...
I've done some SPARC assembler programming for fun, and for someone with 6502 / MC680x0 assembler background, SPARC has a really exotic assembler (register windows and synthetic instructions, most notably). It also illustrates just how complex and powerful the Scalable Processor ARCitecture is; for a RISC CPU, it has loads and loads of advanced features all designed for high performance, which leads me to think that the existing compilers generating SPARC code must be crap, since some of the SPARC processors (notably UltraSPARC II, UltraSPARC III, and UltraSPARC T1) are generally known to be slow when it comes to non-parallelized, number crunching performance. The design of the SPARC processors and their assembler is in direct opposition to that.
http://www.amazon.com/SPARC-Architecture-Manual-Version9-Int...
http://sparc.org/technical-documents/specifications/
are generally known to be slow when it comes to non-parallelized, number crunching performance
I think that's due to a similar set of design decisions that lead to the Itanium also being an amazing benchmark performer, but dismal in "general purpose" code. IMHO SPARC and Itanium are architectures that have optimised heavily for the sort of massively parallel, predictable/few branches, predictable-memory-access-pattern that "high performance computing" benchmarks show off, but at the expense of the small, branchy, less predictable "serial byte/bit manipulation" that tends to be common in other, more general-purpose code. To use a car analogy, it's like a dragster vs. a rally car.
http://studentnet.cs.manchester.ac.uk/ugt/COMP15111/syllabus... http://studentnet.cs.manchester.ac.uk/ugt/COMP22712/syllabus...
Despite its age, ARM System-on-Chip Architecture (Steve Furber) is still a good introduction to the processor and assembly (I'm re-reading it at the moment). ARM Assembly Language - an Introduction (J.R. Gibson) is worth reading if you want to learn ARM assembly.
The materials page for COMP22712 also has some interesting resources, including the lab manuals and a small ARM assembler written in C (source is also on GitHub: https://github.com/uomcs/aasm). Unfortunately COMP15111 is now hosted on Blackboard so you can only get at the materials if you're a current student enrolled on the course.
(I've been an undergrad, postgrad and staff in CS at Manchester, so I'm familiar with the courses - other universities may have similar resources with fewer access restrictions).