Int 80h (2001)
int80h.org
int80h.org
ARM became dominant in mobile computing, which was barely a glimmer back in 2001. x86-64 became commonplace. Just targeting Intel and ARM processors now necessitates writing code for four architectures: x86, x86-64, AArch32 and AArch64 for a full range of compatibility.
While MS-DOS syscall numbers were documented and often _required_ assembly to use, the opposite is true on modern Windows - Windows system call numbers are undocumented and change from release to release, requiring programmers to link against Kernel32.dll or equivalent to stably call into the OS.
Compiler technology has advanced significantly, and processor manufacturers now often optimize towards patterns employed by compilers rather than by humans (case-in-point: the x86 "loop" instruction, while extremely convenient for handwritten code, is significantly less performant than a cmp-jmp loop).
Code has gotten more and more complex. Whether this is a good thing or bad, the reality is that hundreds of millions of lines of assembly would be required to replicate complex modern programs like web browsers - projects of that scale will always require powerful HLLs to manage abstraction, something that assembly does not and cannot provide on its own.
View this page as what it is - a historical artifact. Although assembly programming definitely still has its uses, the arguments this page makes in favor of it are largely no longer relevant.
UNIX kernels could run on a variety of processors, including ARM; UNIX was a relative latecomer to x86.
As for Windows kernels, both then and now, UNIX kernels are open source and do not deliberately obscure system calls.
It's not unreasonable for someone to argue the benefits in writing programs for UNIX.
"View this page as what it is..."
It's one of the few pages written about UNIX assembly. That's how I view it. Years later it was put into The FreeBSD Handbook.
"... aiming to provide performant tools ..." (1995) - https://books.google.com/books?id=uZjYdk4bQgkC&pg=PA1&dq=per...
definitions for different aspects of 'peformant' (1998) - https://books.google.se/books?id=FYkmwYKBoCwC&pg=PA235&dq=%2...
"The network and hierarchical data models, introduced in the 1960's, provide a very performant data organization for well-defined, update-intensive transaction processing of large volumes of data at the cost of conceptual complexity" (1993) - https://books.google.com/books?id=Ro3zjjTNE9UC&pg=PA327&dq=p...
Here's the n-grams trend. https://books.google.com/ngrams/graph?content=performant&yea... .
There are many domain-specific terms which have not yet made it to general dictionaries. For example, from your comments, I see you used the word 'multicast', This is not in Merriam-Webster, at http://www.merriam-webster.com/dictionary/suggestions/multic... . What makes it a word but not performant? Is "zero-copy" a word? "smarthost"? "type-punning"? "upstream"?
Software can be big or small, fast or slow, efficient or inefficient (for combinations of size and speed.) I don't think we need any more adjectives beyond those.
Instead of "significantly less performant", you could just say "much slower". (Note that LOOP was much slower on older Intel CPUs, but was the same as the longer sequence on AMD's. It would not surprise me if newer Intel ones now decode LOOP into the same sequence of uops. In any case, unless the loop is a tiny one, what happens in the body is likely to make a much more significant difference, especially if cache misses are involved.)
Some interesting reading on why LOOP was slower on Intel CPUs --- because Windows 95's timing loop used it (observe the lack of "performant" in any of those posts): https://groups.google.com/d/topic/alt.windows98/liuMAIctHRg
Also, I think performant is often used as a synonym for "fast enough", which is slightly different than "fast."
In looking around just now, economics seems to prefer the second definition, which is why I added it. All of the uses I found are compatible with this definition.
What's the definition of "object oriented" again?
Languages evolve. Whether for better or worse is up to the reader. If it weren't for evolving languages, we'd all still be writing like Shakespeare.
[0]: https://en.wikipedia.org/wiki/Phonological_history_of_Englis...
For example:
- observing => observant =?= observing well (yes, exists in this meaning)
- informing => informant =?= informing well (does not exist as adjective, but could in the future?)
- thinking => thinkant =?= thinking well (does not exist at all, but could in the future?)
Scientist, doctors, lawyers, and other engineering disciplines use precise words. Software engineers should too.
As more and more people use "performant" fewer and fewer people will understand what good performance actually means. I wouldn't be surprised if this encourages a culture of wasteful premature optimization.
On the other hand, when using Asm you can often leverage techniques not possible in HLLs to avoid complexity and abstraction. This is a fact which seems not commonly known, even in CS courses involving Asm, as those tend to be more about inspecting compiler output and the few "write in Asm" exercises are mimicing code that a compiler would generate, which IMHO misses the whole point of writing in Asm.
Here is an operating system containing many nontrivial applications, all written in 100% Asm:
Among other things, it contains a minimal yet functional web browser (around the same level of minimalism as Dillo or NetSurf, i.e. no JS or fancy CSS), and the whole package fits on a single floppy disk.
An entire OS with kernel, drivers, and applications in total can be smaller than a "Hello World" application in a "modern" programming language (which depends on countless other external libraries too, including an OS several GB or more.) It really makes you think.
Here's a small snippet I found from draw.asm:
add r9 , [copyxe]
add r9 , [copyxe]
add r9 , [copyxe]
sub r9 , [copyx]
sub r9 , [copyx]
sub r9 , [copyx]
add r10 , [sizex]
add r10 , [sizex]
add r10 , [sizex]
add r14 , 1
72 bytes.Here's what GCC generates with -Os for the same code:
imul rax,qword ptr [copyxe],3
inc r14
add r9,rax
imul rax,qword ptr [sizex],3
add r10,rax
imul rax,qword ptr [copyx],-3
add r9,rax
36 bytes.As a bonus, GCC with -O2:
mov rax,[copyxe]
add r14,1
lea rax,[rax+rax*2]
add r9,rax
mov rax,[sizex]
lea rax,[rax+rax*2]
add r10,rax
mov rax,[copyx]
mov rcx,rax
neg rcx
lea rax,[rax+rcx*4] ; clever!
add r9,rax
52 bytes.So, not only does GCC hugely beat this snippet of handwritten Menuet assembly in code size, it can even do so handily while folding all the repeated additions into LEA and eliminating most of the loads.
This is an excellent example of why you shouldn't write in assembler.
However assembly programmers will also instinctively avoid things that result in massive amounts of code because they have to manually write every line of it. Compared to C++ where templates are effectively a code generation meta-language. C makes you work a lot harder to generate huge amounts of machine code - and its binaries tend to be smaller because of it.
I don't think that's correct; assuming RIP-relative for all the global variables, the first 9 instructions are 7 bytes each and the last one 4, giving a total of 67 bytes.
You can cherry-pick examples all you want, but I think this is a great example of why you should use Asm: It's not even trying to be optimised, and it still somehow manages to be smaller overall!
Try to optimise it, however, and you can definitely beat GCC...
lea rcx, [copyxe]
imul rax, [rcx], 3
add r9, rax
imul rax, [rcx+copyx-copyxe], -3
add r9, rax
imul rax, [rcx+sizex-copyxe], 3
add r10, rax
inc r14
...with 32 bytes. This human gives, for "-O2", the following 41-byte snippet: lea rcx, [copyxe]
mov rax, [rcx]
lea rax, [rax+rax*2]
add r9, rax
mov rax, [rcx+copyx-copyxe]
lea rax, [rax+rax*2]
sub r9, rax
mov rax, [rcx+sizex-copyxe]
lea rax, [rax+rax*2]
add r10, rax
inc r14
GCC may be clever to turn (-3) * x into x - 4 * x, but not clever enough to realise that turning 3 * x into x + 2 * x like it did for the other two "statements" and a subtraction would be shorter and faster because it means one less operation on the critical path.http://blog.stuffedcow.net/2013/05/measuring-rob-capacity/
The repeated use of rax is also no problem because of register renaming. Despite two dependency chains specifying rax, the CPU can detect that they're independent and assign different physical registers for rax, allowing them to execute in parallel.
I don't really understand this argument. If a few lines of c can generate dozens of lines of assembly, how can it not be faster to have the compiler write assembly for you?
Your comment focuses on how fast the compiler generates lines of assembly, after the programmer has constructed a solution.
But unless I have misunderstood the author was focusing on how fast the programmer can construct a solution.
The final step of depressing keys to input the solution into the computer can be sped up by using libraries and perhaps macros.
It's also possible to write programs ("code generators") that generate lines of assembly based on a "template". And that "template" need not be written in C or any other popular computer language.
It could be some "DSL", a shorthand for assembly, that the programmer creates for himself.
djb's qhasm might be considered an example.
While it may be generally true that C requires less key punching than assembly (certainly it requires fewer "lines"), this does not mean constructing a solution in C necessarily comes faster than constructing one in assembly... if the programmer is familiar and efficient working with assembly.
edit: idk how to put a asterisk up there in the example above
I have no idea where people get the idea that one line of C equals 10 lines of asm.
PS Assembly is not hard at all, and is good to understand/read assembly when coding C. The only thing is that it is a bit rough to start coding in it (no good tutorials is one of the two-three difficulties).
Because coding in asm takes more than just the raw instructions.
x86 doesn't have many registers so you're shifting around stuff across registers and memory all the time, not to mention stack setup/teardown and the error checking (e.g. div0, overflow) you are supposed to do after operations.
I also disagree with "shifting around stuff across registers and memory all the time". An "average" function doesn't really have all that many variables and x86 has 8 GP registers (16 on amd64). You don't even have to declare variables, you just take a register (and/or a location in memory) and use it as you wish.
I'm now sitting here thinking what can be done in C that would end up being a lot of instructions, and i can't think of anything. Going over an array does not, neither does recursion (recursion is just syntactic sugar in the end). What else.. i can't think of anything. Macros ? But macros are cheating in this comparison (not to mention that the flat assembler has actual macros, that are extremely powerful)
This is one of the things where (x86) Asm is definitely easier: divide overflows naturally cause an exception which will get handled automatically, and add/subtract/multiply overflows can be checked with one extra (JO or JC) instruction. In C, this is all undefined behaviour.
I'm now sitting here thinking what can be done in C that would end up being a lot of instructions, and i can't think of anything.
Crypto algorithms come to mind, but like you said, macros can take care of that easily... and on this page is an AES-256 CFB implementation in Asm that is around the same size of source as others in C (but much smaller binary):
See e.g. [1] for a more comprehensive overview concerning the handling of system calls on Linux.
[1] http://blog.packagecloud.io/eng/2016/04/05/the-definitive-gu...
http://web.archive.org/web/20120615223202/http://www.x86-64....
What is even worse is that I think Intel made the mistake of checking for canonical addresses and GP#ing on the SYSRET instruction itself.
I'm guessing this article is somewhat out of date. The systems where performance matters most - those most likely to be constrained in some way, where coding in assembly would yield the most benefit - are mobile platforms, which invariably means ARM not x86.
I think the emphasis on the 0x80 interrupt was to draw parallels to the old DOS calls - int 21h - and make a point that nothing technically stops you from doing the equivalent on Unixy system. (A few years after this the sysenter thing came along and so now int 80h is a legacy path..)
Sadly a lost art, all of this. Talking about this is a great way to get blank stares from many people.