x86 assembly doesn’t have to be scary (2018)
blog.benjojo.co.uk
blog.benjojo.co.uk
I think it's fairly common to recommend beginners to write a compiler but I think it's less common to actually recommend trying to emulate parts of x86. I think it's a particularly easy way to get started just because it's the architecture you already know or because its the architecture that all compiler tutorials use (if they don't use LLVM). And if you are a programmer you probably incidentally have gcc, gdb, and objdump on your system ready to help you out.
Doing both compilers _and_ emulators really helped my understanding of x86 and C (even if I wasn't writing C).
My background is in web development and my reason for doing these projects/writing is purely educational.
[0] https://notes.eatonphil.com/compiler-basics-lisp-to-assembly...
[1] https://notes.eatonphil.com/compiler-basics-an-x86-upgrade.h...
[2] https://notes.eatonphil.com/emulating-amd64-starting-with-el...
[3] https://notes.eatonphil.com/emulator-basics-a-stack-and-regi...
For example, the codegen for some C++ that was discussed a few days ago [1]
This website also has another article which explains how a C++ code can be understood in terms of C code. https://www.avabodh.com/cxxin/cxx.html
I also wrote another article which explain how C++ code can be understood in terms of C code: https://www.avabodh.com/cxxin/cxx.html
The x86 instruction set was a cobbled-together mess from day one, while 68k assembly coding was pure joy because of its elegant and consistent instruction set.
Back then it only had the one register width (16-bit, e.g. "ax"), whereas now it has the 32-bit series (e.g. "eax") series and the 64-bit series (e.g. "rax"). It also now has SIMD, SSE/AVX, virtualization support, and other technologies. Back then it just had one operating mode (real mode), whereas now it has protected mode, long mode, system management mode, and a few other intermediate modes (e.g. "unreal mode").
So a lot of the complexity that x86 has now was introduced after that decision was made. It's definitely conceivable that the 68000 line would have developed similarly had it been chosen instead of x86 for the PC.
but I started in my first job dissassembling 68000, very easy to work with
Yeah, X64 is only used by windows, that's like ~80% of the desktop computing market, plus nearly every server and cloud instance out there, so not much at all. /s
Why do some users assume that the whole world revolves around Apple's iOS/M1 Mac ecosystem as if it exists in a vacuum?
Because for those users, it does, and Apple coddles and encourages that mindset.
Desktop computers are a declining and relatively small market. Your home, car and office are full of ARM devices. Potentially hundreds.
x86 lives in the data centre (Linux, not Windows, so contrary to grandparent post, but whatever) and in some desktop systems. Not the majority of systems.
Has nothing to do with an Apple fetish.
Is this true? I thought Linux dominated the server space.
LUI a0, 0x7FFFF
ADDI a0, a0, 0xFFF
is superior to MOV eax, 0x7FFFFFFF
Except, of course, that RISC-V snippet doesn't actually load 0x7FFFFFFF into a0 because Reasons, it has to be LUI a0, 0x80000
ADDI a0, a0, 0xFFF
"But the assembler has the LI pseudoinstruction so it will properly do this calculation for you!". Right, so much for "nice assembly language": you need an actual smart macroassembler to write it.The funny part is, there is now the "C" extension to RISC which introduces 16-bit instructions which are allowed to freely mix with 32-bit instructions — so now those 32-bit instructions can be 16-bit aligned and even be split between two physical memory pages which kinda kills the whole "but at least fixed-length encoding prevents Spectre-like exploits" argument.
Assembly isn't supposed to be convenient and expressive to write. for that we have high level languages. But some consistency makes for less error prone and easier analysis.
But RISC-V specification is explicitly structured around describing several basic cores and the extensions to those so it looks like it's all very unified and consistent: and indeed, it mostly is since it was developed mostly in one continuous effort with consistency in mind. Bute there are still some inconsistencies between how things are done in different extensions as well, for example, the "C" extension uses zero-extended immediates in half of its instructions unlike the rest of ISA and the other half of this very extension because of pragmatics: nobody would like to have negative offsets in those shortened instructions, so those are unsigned.
In x86 the mnemonic "MOV" can be translated into instructions with a different opcode according to the addressing mode, immediate value size, or target register. Most of the x86 instructions have a similar issue, while RISC-V macroassembler are pretty simple. Therefore, the x86 assembler must actually contain much more intelligence than the RISC-V assembler to make it look "simple" and "nice".
Then I read some of the Intel stuff. OMG.
Bare metal embedded coding is a different topic of course, but regular application development with assembly works just fine.
'Regular applications' that use inline asm should really be raising eyebrows.
https://en.wikipedia.org/wiki/RollerCoaster_Tycoon_(video_ga...
> The game was developed in a small village near Dunblane over the course of two years.[2][5] Sawyer wrote 99% of the code for RollerCoaster Tycoon in x86 assembly language, with the remaining one percent written in C.[3]
That's maybe a little spun. The separated address/data registers (there are two kinds of registers, and memory operations need to use one from each to combine to a final address) played hell with optimizer strategies for years.
In fact 68k compiler output was significantly sub-par pretty much throughout its lifetime. The "cobbled-together mess" had (post-386 anyway) a significantly more orthogonal instruction set and was just plain easier to optimize, even for humans.
On the 68020 you get even more flexibility.
That's the kind of complexity that really hurts software. Compare vs. the commonly-cited x86 nonsense (the REP prefix, say), which complicates silicon implementations but generally makes software easier to write (c.f. decades of optimized inline memcpy implementations).
The point being the 68k was a dead end in a different direction. It was a "clean" archiecture from the perspective of a 1970's assembly programmer, but not a late 80's compiler writer.
When you have 15 registers to work with that's not going to be more of a problem than only having 7 or 8 registers to work with.
How hard would it really be to custom build a chip that was really simple, but modern. I'm thinking like the 6502 in the Commodore, but much faster. Or is the complexity in x86 an inherent property of modern performance? I guess what I'm getting at, is could you build something that is actually pretty darn fast if you don't need to run a modern OS like windows or Linux on it, but keep it drastically simple?
I've been thinking it might be neat to have a blazing fast NUC sized computer that just boots into some barebones forth sitting on top of a few assembly words. Maybe with just enough peripherals to do some actual work (like load from SD card).
A lot of the performance gains over the last couple decades haven't come from machine code changes, but rather from various forms of pipelining and superscaling. Rather than run one instruction at a time, CPUs run hundreds of instructions at a time. The complexity of doing that is that many instructions often depend on the result of instructions immediately before them, and so the CPU needs a lot of shortcut paths internally to keep from stalling out.
You could get to Ghz speeds with a custom ASIC, and you could use base RISC-V as a modern, not crufty ISA. It still won't be anywhere near as fast as a high end x86 or ARM CPU, though, unless it's similarly pipelined and superscaled.
You could make a CPU optimized specifically for forth or whatever, and likely achieve better performance for the number of transistors than otherwise. For example, if virtual memory isn't helpful, than you can omit that whole subsystem.
Underneath the covers, however, the underlying hardware complexity is an inherent property of modern performance. For example, branch prediction.
It's confusing as a beginner to encounter the "di" and "dx" registers, and to understand why the 8-bit version of "r13" is "r13b" but the 8-bit version of "rdx" is either "dh" or "dl".
I like x64 a lot more because IMO it allows you a few important memory related ones in x86 making life a lot simpler.
As a newer developer this is a problem I've had with computing in general - a lot of newer languages, web frameworks, etc. are solving problems that I only understand through seeing what came before me. For example - why should I care about memory ownership if I've never tried to write a large-scale C program?
That said, the C standard is also an impenetrable mess, yet people can start using it without resorting to it.
Those are highly obscure.
It would be certainly a smaller tutorial, but historical perspective might be valuable for understanding cruft layers of design decisions that made sense decades ago and addressed problems that have ceased to exist.
https://github.com/torvalds/linux/blob/master/arch/x86/entry... (subject to change, see your specific kernel source)
I think more fun is doing assembly on a µC without OS with some interesting peripherals and many MCUs have an instruction set inspired by x86.
Maybe this is just a case of having learned and written 32-bit x86 assembler for years in my youth, but I strongly prefer it to x86_64.
Changing the signature of functions requires rearranging which registers are being populated, and if there's enough parameters some go in the stack, and if you've rearranged the order, now you're moving some from stack back to registers and visa versa. At least when everything was always going through the stack, you just rearranged their position in the stack. x86_64 will always be more annoying in this regard, it's not an ABI decision made with humans in mind at all. The assumption is (rightfully) that compilers are doing this work, and the perf win is significant.
If a function return value could fit into %r0 on the PDP-11 (%ax/%eax/%rax on x86), it would be returned in there. %r1-%r4 would be used to pass function parameters in, if they could fit, and/or spill over into the stack.
Heck, even UNIX system call conventions on x86 can be traced back to PDP-11, i.e. the syscall number is passed in %r0 (%eax) followed by a TRAP (INT on x86) instruction (can't remember which TRAP number, though).
I love articles like this that break it down. The biggest challenges so far have been getting the assembler to output the correct format (ie, 16 bit real mode) and learning inline assembler in C (and getting GCC to output the correct format).
It's basically "Writing an Operating System from Scratch for Dummies".
I actually wrote my own graphical x86 OS starting from the code in there.
Do you think it's a waste of my time to be learning this in 2021?
How smart are compilers these days ? Say, to optimize small function, for example computing a scalar product or applying a 3D matrix transformation to a set of points.
Generally, they do lots of inlining, and then once you inline you can get some more optimizations in, rinse and repeat. Ends up pretty optimal. (This is C++, can't speak for other languages.)
We work on very perf-sensitive code and we never drop down to assembly. For hot loops, we usually inspect the generated assembly and if it's not great, it's fairly easy to "nudge" the compiler towards the better-performing solutions by tweaking the source code. Also some manual unrolling might be needed to better saturate the vector processing cores of modern CPUs.
And when you're working with signed integers, you still have to do stuff like a >> 1 instead of a / 2 :)
But if it's now up to some nudging, it's vastly better to me.
I provided math stuff and vectorisable stuff on purpose :-) Happy to see that vectorisation is somewhat automatic. I remember the MMX days and there were not that funny :-)
Here's an example that looks like magic at first glance where the compiler converts manual bit twiddling code to a popcnt instruction (with the right compiler setting), but do the bit counting any other way, and the whole thing falls apart:
So while 68000 ASM may look pretty and be easy to remember, its obvious why x86 won - you could get the critical jobs done.
At the same time, I used a clever 12-byte sequence of those opcodes to swap software interrupt vectors on process switch. So each process could have its own floating-point exception handlers etc. We called it the 'soft vector chain' and it was among the least known parts of our kernel.
A colleague (John McGinty) suggested in our old age we could consult by scratching our beards and saying "Ah! It must be the soft vector chain!" for every problem.
It is a bit less obvious, in fact.
Motorola never looked at 68k CPU's as a serious business, more like toying around with the CPU's all the time or looking at them as a less important spin-off of the main business. Their main sources of income were defence contracts (e.g. specialised or hardened microchips), microcontrollers, DSP's (which were pretty cool, by the way, – all implementing the Harvard architecture), memory chips (I think), radios and the field radio equipment and later mobiles.
They carried largely the same attitude to 88k RISC and PowerPC CPU lines (albeit trying to compete more seriously for a while), but ultimately failing to catch up and leading to the eventual PowerPC demise. After that failure, they spun off anything CPU, DSP and microcontroller related into Freescale, and the rest is now history.
...just ugly.
There was a recent thread [1] optimizing some very primitive trig functions.
I recently watched a djb interview where he talked about the importance of fully utilizing available hardware [2 at 5:15]. That can be a good starting point, although at work its usually easier to just consume more resources than to use what you've got more efficiently.
; To build:
; nasm -fobj hello.asm
; alink -oPE -subsys con hello -entry main
bits 32
import imp_CreateFileW kernel32.dll CreateFileW
import imp_WriteConsoleW kernel32.dll WriteConsoleW
import imp_CloseHandle kernel32.dll CloseHandle
import imp_ExitProcess kernel32.dll ExitProcess
import imp_GetLastError kernel32.dll GetLastError
extern imp_CreateFileW
extern imp_WriteConsoleW
extern imp_CloseHandle
extern imp_ExitProcess
extern imp_GetLastError
%define CreateFile [imp_CreateFileW]
%define WriteConsole [imp_WriteConsoleW]
%define CloseHandle [imp_CloseHandle]
%define ExitProcess [imp_ExitProcess]
%define GetLastError [imp_GetLastError]
global main
section .text use32
main:
push dword 0 ; hTemplateFile = 0
push dword 0 ; dwFlagsAndAttributes = 0
push dword 3 ; dwCreationDisposition = OPEN_EXISTING
push dword 0 ; lpSecurityAttributes = NULL
push dword 2 ; dwShareMode = FILE_SHARE_WRITE
push dword 0x40000000 ; dwDesiredAccess = GENERIC_WRITE
push dword filename
call CreateFile
mov [handle], eax
push dword 0 ; lpReserved = NULL
push dword nChars ; lpNumberOfCharsWritten
push dword 16; nNumberOfCharsToWrite
push dword hello ; lpBuffer
push eax ; hFile
call WriteConsole
mov eax, [handle]
push eax
call CloseHandle
push dword 0
call ExitProcess
err:
call GetLastError
push eax
call ExitProcess
section .data use32
filename: db __utf16__('CONOUT$'), 0, 0
hello: db __utf16__('Hello Console!'), 13, 0, 10, 0, 0, 0
handle: resd 1
nChars: resd 1What browser are you using?
It's fine when I initially open the page, but once I click 'run' I can no longer use PgUp, PgDn, F12, etc.
I'm guessing the emulator captures some input. But even when clicking on other parts of the page, it doesn't 'uncapture'.
FWIW, it doesn't really bother me.
Before hitting the emulator (and/or editor, not sure) all keys work fine. After, well, only scroll on the mouse worked for moving through the page.