x86-64 Assembly Language Programming with Ubuntu (2022)
egr.unlv.edu
egr.unlv.edu
this book also looks like it could be a very solid base for a course that did cover one or more of those additional topics; the course would also include some supplementary material
it probably teaches more than everything i know about amd64 assembly, except the most important thing, which is that arm assembly is much better
(while being an order of magnitude simpler)
The main pitfall while writing assembly is the abuse of a preprocessor. Being dependent on the grotesquely and absurdely complex and massive compilers out there is one thing, but moving that dependency to a complex preprocessor is not that much better. So caution and care about that issue must be kept in mind.
I wish all RISC-V CPU vendors to pay the silicium real estate price for 64bits because writting once a 64bits RISC-V code and to be able to "run-ish" it everywhere, from "embeded" to servers passing thru workstation, wow.
I am writting currently x86_64 assembly, namely the manual "register-ization" of some code paths is done, ready for an easy RISC-V port. Not to mention RISC-V has double the amount of registers, and some code paths will really benefits from this additional register space (intel plans to follow that route).
Ofc, this will be mostly micro-arch agnostic assembly code, as it is the case already for x86_64, with maybe simple or very generic static "optimizations" (cache line, alignment, registerization, leveraging some instruction fusions, etc). Worst case scenario, adapted (not rewritten from scratch) assembly written code paths to fit better on a micro-arch with a runtime switch/installation... if really needed. I guess correct code will become very important, maybe more than fast code ("should" become true for hardware design too, if the performance penalty is not too high).
RISC-V will need compiler support though, that for legacy support, and sometimes, humanization of the assembly output of some compiled programs may be usefull.
All that depends on the success or not of RISC-V, which for that will need ultra-performant implementations all over the board, micro-archs and best silicium process.
RISC-V is not perfect, but a more than good enough modern ISA, and the real risk of fragmentation is 32bits/64bits code paths even though some care was provided to make 32bits<->64bits code adapation easy.
Mistakes will be made (micro-arch with critical bugs), so it won't happen overnight.
the preprocessor thing depends on what it is that bothers you about the grotesque compilers. a macro preprocessor can be pretty small; gpm (basically unix m6) was reputedly 250 machine instructions
historically 'worse is better' and 'the innovator's dilemma' suggest that adoption of an innovation like risc-v depends more on conquering the low end of the market long before the high end. but then there's the tesla roadster...
shame about the t bit tho
Developing a compiler from scratch is the other significant use for writing it. Of course it is quite common to need to read it.
[1]: https://gcc.gnu.org/onlinedocs/gcc/x86-Built-in-Functions.ht...
Edit: also useful are https://gcc.gnu.org/onlinedocs/gcc/Extended-Asm.html and https://gcc.gnu.org/onlinedocs/gcc/Local-Register-Variables....
This is a notable differentiation - Writing assembly is a different skill to reading it from a disassembly. Reverse engineering, malware analysis etc. does not inherently require you to be able to write asm, although it certainly would help.
The pinning local variables to registers thing didn't work in llvm a couple of years ago (which seems consistent with the gcc docs) but does work at the boundaries of inline asm and that's generally enough. I like a pin-register intrinsic, something like `u64 pin(u64, enum reg)` where the compile time constant enum names the register and the semantics are a no-op other than constraining the register allocator, but that doesn't seem to be readily available in gcc/clang.
I don't have a good answer to constraining instruction scheduling.
On reflection it's all somewhat more horrible than it needs to be, perhaps inline compiler IR is a better idea.
My point is that conditional moves are one of usecases badly supported by compilers, and that (may) require dropping to assembly.
dan bernstein makes the argument that, as computers get faster, we use them on bigger problems, which means that computer performance is increasingly dominated by small inner loops, which is precisely the situation where it becomes more rational to put effort into hand-optimizing your small inner loops than to hack on the compiler to hopefully speed up all parts of the program, just as it was in the 01960s for different reasons
a different way to attack that problem in many cases is to write a domain-specific compiler from a domain-specific language to machine code, as thompson's regexp engine did, and as verilog compilers do. but i'm not sure how you speed up a media codec that way
bernstein has also written a fair bit of assembly to eliminate timing side-channel leaks from cryptographic code
Example: https://github.com/dddrrreee/cs140e-23win/blob/85b9ae3bd46c7...
Conservative garbage collectors. Scanning the native stack for pointers can be done in C but isn't quite enough since there might be pointers in registers. So I wrote assembly code to spill all the registers onto the stack prior to scanning.
I would agree that direct asm is very rare these days in game engines outside of third-party libraries. There can be significant gains with tuned asm but some combination of intrinsics and ISPC is usually good enough. But it is far more useful to be able to _read_ assembly, for debugging in an optimized build or analyzing release crashes.
I would have preferred to emit something like LLVM IR instead, but couldn't because of several constraints.
See the .s files in: https://cs.opensource.google/go/go/+/refs/tags/go1.21.5:src/...
I occasionally see it in compression as well.
(my own ASM experience is limited to using TASM to write some games in Z80 on my TI-83 graphing calculator)
The Intel syntax, common in the PC world since the MS-DOS days, used across Windows and OS/2 as well.
In macro Assemblers, inline Assembly in high level languages, and naturally Intel and other x86 manufacturers CPU manuals.
It uses the format,
op dest, source
Then you have the AT&T syntax, for whatever reason when support was added to UNIX for x86, they chose the format used by other architectures.It is only found on GNU/Linux and BSDs, and follows
op source, dest
Where op is written with the data size prefix, and addressing modes are somehow more complex.Intel:
lea eax, [eax + eax * 4]
AT&T: lea (%eax, %eax, 4), %eaxSince this book targets Ubuntu I'm assuming Linux supports Intel syntax now, so I guess that this also means that it will be a more "portable skill" across different systems.
Android is probably the only Linux based system where Intel's syntax is favoured via yasm's inclusion on the NDK.
Generally AT&T syntax is seen as a bit weird in the x86 realm, and there are some Unix assemblers like NASM that use Intel syntax.
https://developer.apple.com/documentation/xcode/writing-arm6...