AT&T syntax is broken (2021)
outerproduct.net
outerproduct.net
AT&T Syntax versus Intel Syntax - https://news.ycombinator.com/item?id=33585154 - Nov 2022 (104 comments)
Also, it's worth mentioning that what people are normally referring to when they say "Intel syntax", beyond the syntax in the oficial docs, is more like "TASM IDEAL" than MASM or Intel's little-known "official" assembler (named ASM386.EXE, seemingly rare and hard to find these days); in the former, a label is consistently an address constant and all memory operands are enclosed in [ ], while the latter distinguishes between labels and variables.
mov.b (a3), #42
or whatever.For example: suppose I want to be able to unambiguously use the names of C globals in my assembler - I might insist that they all get extended with an initial '_' or I might insist that registers get a '%' prepended
Or: many of my potential customers use IBM 029 card punches, I can't use [] to indicate indirection because they can't punch those characters, so I'll use () instead
Or: people in the UK don't have a $ key, I'll use a # to indicate literals instead
These are historical issues I've certainly had to deal with when designing long lived things
To be proficient in reading or writing assembly (versus higher level languages) means to deal with a stream of instructions. Once you are fluent with that concept parsing individual instructions is a lookup. Operand ordering is just a small part of that. It can be x86, arm, or tis-100.
Switching between x86-intel and x86-at&t is no different than switching between x86 and arm.
mov eax, ebx
means moving ebx into eax.When you move something, you move it to its destination, not from its destination.
mov A, B
should mean "Move A into B".The mnemonic operation names are in English though. mov = move, add = add, jmp = jump etc. It's not like it's just hex codes, APL or K. So I think the English argument kind of makes sense.
> It’s somewhat telling that assembly has no equivalent to the word ‘to’, and few enough production rules to count on both hands.
Assembly was designed to be fairly concise. Blaming a 1970s assembly syntax for not having "to" instead of a comma is kind of silly. It used commas and will keep using them probably. If the syntax was more advanced, it wouldn't be assembly, it would be C or some other higher level language.
TFA> The association between this numeric form and its meaning under the ISA is completely arbitrary
P> It's not like it's just hex codes
The numeric form of machine code, appropriately in hex or octal, is not arBITrary at all, the bits mean things that help group and decode the instruction with minimal logic. In the first example in the article, it's easy to see the registers he's talking about, and it would be easy after that to decode the addressing modes, alu functions, etc.
They're not arbitrary, was my point.
MOV EAX, 5
Vs
MOV $5, %EAX
- 6502 uses "load"/"store" (LDA/STA, LDX/STX, LDY/STY), but also "transfer" (TXA, TAX, TYA, TAY, TSX, TXS).
- Z80 has "load" (LD, LDD, LDIR) but NO "store" (stores are LD src=reg,dest=ram I think)
- Some 4-bit Sharp CPUs use "t" meaning transfer as a JMP.
- Intel 4004 has mnemonics with Fetch, Read, Write, Load
- MIPS has both "load"/"store" (lb/sb,lbu,lh/sh,lhu,lw/sw,lui,la,li) and a couple "move" mnemonics (mfhi/mthi, mflo/mtlo).
- 68000 has MOVE (MOVE, MOVEA, MOVEM, MOVEQ), but then you do have that pesky LEA (Load Effective Address) instruction.
- Signetics 2650 is consistent, you have LODA/STRA, LODI, LODR/STRR, LODZ/STRZ for load/stores of registers to/from RAM and LDPL/STPL, LPSL/SPSL, LPSU/SPSU for load/stores of the PSW to/fom RAM.
There are a whole bunch of instructions with no operands at all that nonetheless have sizes. The most useful that comes to mind is SYSRET. sysretl and sysretq are both useful but are not interchangeable.
But x86_64 made 32-bit operations zero the high 32 bits of the destination register, so, logically, 0x90 would clear the high bits of RAX. Oops.
As I’ve heard the story, this was noticed a bit on the late side, and an 11th hour change was made to special case 0x90 and make it still be a nop.
(I personally find cmp to be really annoying in AT&T syntax —- the AT&T syntax cmp; jge is backwards, whereas the exact same sequence in Intel syntax makes sense. (jge is “jump if greater or equal”. But which values are compared in which order? [0]). AVX-512 in AT&T syntax is extra bizarre, too, if for no other reason than that it doesn’t match the documentation.
[0] For extra mental gymnastics, CMP with two register arguments can be encoded in two different ways: the way where the first operand could have been, but wasn’t, in memory, and the way where the second operand could have been, but wasn’t, in memory. Unlike SUB, CMP doesn’t really have a destination, but AT&T still reverses it for, ahem, consistency.
A=A&C A; ?A=0 A; GOYES bit_is_not_set
Then MOV style syntax felt weird..
https://www.scs.stanford.edu/05au-cs240c/lab/i386/s17_02.htm
https://xem.github.io/minix86/manual/intel-x86-and-64-manual...
So some of the world’s most popular compilers, got it.
If GCC had been sold at Solaris C compiler prices it wouldn't be as popular.