“Hello World” in assembly language on Linux
jyotirmoy.net
jyotirmoy.net
# as hello.S -o hello.o
# ld hello -o hello
# ./hello
.data # .data section starts
msg: .ascii "Hello, World!\n" # msg is ASCII chars (.ascii or .byte)
len: .int 14 # msg length is 14 chars, an integer
# .int 32 bits, .word 16 bits, .byte 8 bits
.text # text (instruction code) section starts
.global _start # _start is like main(), .global means public
# public symbols are for linker to link into runtime
_start: # _start starts here
movl len, %edx # value 14 copied to CPU register edx
movl $msg, %ecx # memory addr of msg copied to CPU register ecx
movl $1, %ebx # file descriptor 1 is computer display
movl $4, %eax # system call 4 is sys_write (output)
int $128 # interrupt 128 is entry to OS services
movl $1, %eax # system call 1 is sys_exit (prog exits)
movl $77, %ebx # status return (shell command: "echo $?" to see it)
int $128 # call OS to do it via interrupt #128Things like system calls and program startup are more advanced topics you'd introduce later, even if they make for shorter code.
And to nit even further: your example is "lower level" but still relies on a metric ton of magic being executed by the linker. If you're allowed to "just invoke ld" the OP's "just call and return from main" trick doesn't sound so bad.
Check this out http://www.muppetlabs.com/~breadbox/software/tiny/teensy.htm...
Who thinks in terms of CPU registers these days, for instance?
The people for whom it makes sense, same as the people who think about RTCs or deep sleep modes or programming single-board computers to read temperature sensors.
It isn't like the low level has gone away or become completely inaccessible. It's that now we have more options to write working code, and we can optimize for things beyond cycle-level and byte-level efficiency. Optimizing for readability, for example, wasn't really an option when all of the readable algorithms were unacceptably slow due to constant terms.
Breaking every rule about data-hiding to get somewhat better constant terms in the big-O analysis isn't virtuous, it's just what was forced on us by insufficient hardware.
That's pretty much the only way I think these days, but I guess assembly does not qualify as "modern language" even though it can and is used to control the lion's share of the world's "modern" computers. Matters little to me; I'm hooked.
I do not write in Lua, and I am sure I will be quickly corrected by someone who does, but isn't it considered both (virtual) "register-based" and "modern"? My sincerest apologies to the Lua experts if I am wrong.
.data
msg: .ascii "Hello, World!\n"
len = . - msg
.text
.global _start
_start:
mov r0, #1 /* fd 1 = stdout */
ldr r1, =msg /* message */
ldr r2, =len /* length of message */
mov r7, $4 /* write */
swi #0 /* syscall */
mov r0, $0 /* status */
mov r7, $1 /* exit */
swi #0 /* syscall */
I couldn't get the # style comments to work, so had to resort to /* */ instead. It's very similar because the code is quite simple, but could be a good starting point if you are interested in ARM assembly. You can use the same commands as above to build and run.Having tried the perfection that is MASM, I'm always disappointed by what NASM lacks in comparison and I'm not at all comfortable with GAS.
Choosing a good assembler, an assembler that you like, that you're comfortable with, is critical.
Heavy lifting is done by an external routine being called.
That does not make you learn assembly language.
ASM is about interrupt (bottom half), calling other piece of assembly, saving and restoring the registers, and above all mastering the mov, and the art of deciphering how registers are used.
This is only just assembly quiche programming.
for one it uses more code than necessary to do things it doesn't need to do (checking arguments, storing the string in two pieces)
for another its just not in the spirit of hello world, which is supposed to show something that just displays hello world.
its making assembly look harder than it is.
Is this a typo? The return value is stored in %rax or %eax, not %ebx.
AT&T syntax:
movl mem_location(%ebx,%ecx,4), %eax
Intel syntax: mov eax, [ebx + ecx*4 + mem_location]
The other differences are minor compared to that, IMHO. mov eax, [ ptr ]
is like eax = *ptr; // or, eax = ptr[ 0 ];
The offset/multiplier memory addressing format for AT&T syntax was always more troubling for me. Coming from a TASM/MASM/NASM/PASCAL/x86 background first, it felt "icky" to put offsets outside of the "brackets" (or parenthesis, as it were) [0][1].[0] https://github.com/lpsantil/rt0/blob/master/src/lib/00_start...
[1] https://github.com/lpsantil/rt0/blob/master/src/lib/00_start...
68k, Alpha, PDP-11, SPARC, and VAX have the destination on the far right.
The order is probably as contentious as the great endianness debate, but I think one of the most awkward parts of having src, dst order is that subtraction looks backwards. I prefer dst, src because it corresponds closely with the direction of assignment in higher-level languages:
op a, b, c ; a = b op c
op a, b ; a op= b[0] http://wiki.osdev.org/SYSENTER#Compatibility_across_Intel_an...
[0] http://en.wikibooks.org/wiki/X86_Assembly/Interfacing_with_L...
Here you go:
; Hello World, linux x86_64, nasm syntax
section .rodata ; Begin read only data section
hello: db "Hello, World",0x0a ; String, 0x0a is \n
hello_len equ $-hello ; $ is current address, length is address after string - address of start of string
section .text ; begin code section
global _start ; export _start so the linker can see it
_start: ; program entry point
mov rax, 1 ; write(2) syscall number
mov rdi, 1 ; stdout
mov rsi, hello ; string address
mov rdx, hello_len ; string length
syscall ; execute the write syscall
mov rax, 60 ; exit(2) syscall number
mov rdi, 0 ; exit status
syscall ; execute the exit syscallIt's not ideal, and should really fall upon designers to fix, but a workaround (in Chrome or Firefox) is to edit the page's source (either with Chrome webtools or Firebug) and add a line-height attribute to <p> or whatever tag or class is most applicable, and set it to something like 125% or 135%. Usually results in a much more readable page.