It would be delightful if GCC were to adopt something similar after this stabilizes in Rust.
It would be delightful if GCC were to adopt something similar after this stabilizes in Rust.
Thank you for helping to validate that decision; this is the kind of feedback we needed.
(And note that you can still choose to use AT&T syntax, it just isn't the default.)
That said, this is a good decision because C compilers seem to be the major holdouts at this point—binutils has had .intel_syntax for a long time now, it’s just not supported inline.
Microsoft's assembler uses Intel syntax, and their inline assembly also uses Intel syntax.
Their inline assembly is much less explicit about inputs, outputs, and clobbers, and "just" has you mix C symbols and labels with assembly instructions, which I imagine is a pain to get right.
However Microsoft has de-emphasized inline assembly and doesn't allow it on amd64 or ARM. For things that need specific instructions you need to either use intrinsics or put assembly in a different object file.
That's just sad :(
Examples: An abort call on an embedded system that needs to disable interrupts to prevent task switching after the abort. One instruction.
An implementation of _start for embedded that can do everything in C except setting a couple processor flags. Two instructions.
Running (with a lock) event callbacks in an embedded system on their own stack, so that not all tasks need to have their stack big enough to handle the stuff the callbacks might do; 5 instructions that are far better inlined than having another call and indirection. (In addition to keeping doc/code together).
But the way they had it, they presented a seamless blend of assembly instructions and any C identifier, in or out, and I guess the compiler would need to parse out any side effects and cope with them. My guess is they looked at porting all that to ia64 [which they supported until Server 2008 R2], amd64 and ARM and balked.
No idea if they ever had it on some of the old architectures they supported in NT4 days or on CE (alpha, mips, ppc).
add.s.ne
ldm.ia.cc
(in both orders, for those cases with two suffixes)
asm(".intel_syntax noprefix;"
"xor eax, eax;"
"xor edx, edx;"
"1:;"
"mov r8, [rdi];"
"mov r9, [rsi];"
"sub ecx, 64;"
"jl 2f;"
"cmp r8, r9;"
"jnz 3f;"
"lea rdi, [rdi - 8];"
"lea rsi, [rsi - 8];"
"jmp 1b;"
"2:;"
"not ecx;"
"shr r8, 1;"
"shr r9, 1;"
"shr r8, cl;"
"shr r9, cl;"
"cmp r8, r9;"
"3:\n"
"seta al;"
"setb dl;"
"sub eax, edx;"
".att_syntax prefix;"
: "=&D" (d0), "=&S" (d1), "=&d" (d2), "=&c" (d3), "=&a" (cmp)
: "0" (l), "1" (r), "3" (nr_key_bits)
: "r8", "r9", "cc", "memory");Most Rust code is quite portable, targeting ARM, x86, MIPS, PPC, WASM, Sparc, s390x, riscv, ... That means, that for many snippets of inline assembly, you might encounter ~8 of them, one for each architecture, all using different syntaxes.
Intel syntax is quite similar to that of other popular architectures `op dst, args...`.
Adding another second syntax for x86 just doesn't add that much value IMO, and adds quite a bit of cost: now everybody dealing with x86 assembly needs to learn 2 syntaxes... and everybody dealing with portable code now needs to be at least able to read 2 syntaxes for x86... Without talking about the cost of implementing a second syntax in the compiler, etc.
If you prefer AT&T, you can always write a proc macro that translates it to Intel, and use that in your projects.
If I ever need to deal with such code, I'd just expand the macro to read the actual Intel syntax, modify that, and either fork the project, or submit a patch with a fix using Intel syntax.
If you prefer AT&T (or you have a large body of existing AT&T code you don't want to have to translate all at once), use asm!("...", options(att_syntax)) and it'll Just Work.
Nowadays GCC and LLVM support both styles and archs pick whatever they prefer and nobody cares, really.
> either fork the project, or submit a patch with a fix using Intel syntax.
That sounds a bit extreme? Reading/writing in both styles is not an issue for anyone that has dealt with x86 professionally.
Is this actually true? Admittedly I've done mostly x86 and ARM for the past several years (almost entirely Cortex-M, so v6-M and v7-M profiles, using ARM's GCC builds for embedded) and the only toolchain that prefers AT&T syntax is x86 GCC and those explicitly trying to be compatible with it. All the ARM inline assembly I've written, targetting GCC backends, has been ARM syntax, and likewise for all the disassembly output.
The DSPs and DSP-likes are... always weird. So I try to stay away from them and make them someone else's problem. But I don't think they use AT&T syntax either. It doesn't work so well for truly strange processors anyway.
I'm one of those guys who tends to have the makefiles output the disassembly, and have it open on the other monitor while I'm working, so I'd notice if it were different....
1. It’s the syntax in the manual. The last thing I want to do when reading or writing asm is to mentally translate from the manual to AT&T syntax.
2. Addressing like (%rax) is tolerably. But the AT&T scale * index + offset syntax is inexcusable. Give me the verbose Intel addressing syntax any day, please.
(As a kernel programmer, I’m more familiar with AT&T. I still hate it. I’m morbidly curious how Intel syntax ought to handle things like SGDT. Maybe SGDT SIXBYTE PTR [address]? The fact that four bytes is called a DWORD isn’t great.)
Right! I worked on C++ compilers for years and I don't even know where is the canonical book of AT&T mnemonics. At the rate that Intel is adding new instructions, using anything other than their official docs (and thus their official names) seems nuts.
However, I recognize that it's objectively worse for a human coder and contains some syntactic footguns which shoot even a experienced coder regularly:
- `number` (memory displacement/pointer) used instead of literal `$number`. It's an easy automatic mistake even for a person who knows this very well.
- the SIB clauses for x86 (memory addressing with constant Displacement and Scale and Index, Base in registers) look like `D(B, I, S)`: it's possible to remember this, but reading/writing it is not as obvious as `[D + B + I * S]`.
- Intel syntax in general is more similar to high-level languages, even though it's more verbose.
- AT&T syntax has syntactic redundancies like '%' before each register which make code much noisier than needed.
In practice, though, naming a symbol with a register name is error-prone and confusing. I'd rather have a way to escape a symbol when it's really needed than to pay the price of noise for an admittedly bad idea.
This allows for unusual languages like LISP and FORTH, without mangling the symbols. Symbols could have commas and spaces.
Keeping tricks like that stable and clearly documented is, uh, not for the faint of heart.
That's not to say that the Rust team should have tried to find (or create) an alternative; it's outside their core mission, and choosing between popular existing alternatives is the right framing.
Why there is always a guy who doesn't like the default and unaware that it's just one of the options?
Modula-2, Pascal, C, C++, Basic based compilers with inline Assembly parsers or intrisics that interact with host language type system instead of manually dealing with strings UNIX style.
The asm macro does know about what group of registers (or what specific register) a input/output is stored/loaded from and compilation will fail if the input isn't compatible with the register.
Then it's delegating to a assembler for the rest.
Sure it could also have a list of all asm commands the given target does support and parameters/registers it can be used with but this means keeping track of all supported targets and all commands for all potential featuresets of all targets and all way they can be called. This is a lot of work. Even more such support can be added later one, too. For now it's more important to support asm on stable.
Also with proc-macros you can implement this as a library on top of `asm`!
Also with a more "knowledgeable" asm any support for a new platform would need to manually add support for all asm of that language so having a standard stringly interface is a good think anyway.
Saying that fits with the type system is a stretch... it is like any other FFI.
The important part, type system integration, is already there. Sizes must match, move and borrow checking still happens, etc. This avoids the typical failure modes of unix-like stringly-typed systems. In fact if you read the announcement and RFC, you'll see that this was a primary goal of the new design!
I think they were absolutely right not to go down that route, because then you have to basically specify this syntax for every possible architecture targeted by Rust. Integrating ASM in the language syntax is fine if you only care about x86, but it's going to be a mess if you want to support MIPS, ARM32, AARCH64, Sparc and whatnot because the ASM dialects for each of these architectures have special bespoke syntax to deal with various quirks of the underlying ISA. MIPS has special syntax to load the top and bottom half of an address as well as dealing with delay slots (.noreorder), ARM has special sugar for PC-relative addressing etc...
Importing that baggage into Rust would be a fool's errand IMO. It would also make it harder both to port Rust to new architectures and to transfer existing assembly code into Rust since it would require adjusting the syntax.
I'm very happy with Rust's current approach, it's a good middle ground IMO.
function MultiplyBy2Reg(aInput: int64): int64 assembler
register;
asm
SHL RAX, 1;
end;
function MultiplyBy2(aInput: int64): int64 assembler;
asm
mov RAX, aInput;
SHL RAX, 1;
end;This is a very nice feature to have if you know your environment is overwhelmingly tied to a certain architecture and you want to make it easy to target this architecture in particular.
This might actually be a somewhat reasonable, if short-sighted approach if we were in 2005 and you could make the point that anything that's not x86 is effectively legacy or niche but nowadays you need to at the very least support x86-64, ARM32 and AARCH64. Baking all these flavors of assembly into a language would be very heavy handed and a maintenance nightmare.
Beyond that we're not in the 90's anymore, ASM is not routinely used outside of very low level code these days. Even for performance people often opt for intrinsics instead. Compilers got massively better at optimizing code while humans got worse due to ever more complex architectures.
Being able to seamlessly blend ASM with normal code isn't that much of a killer feature anymore IMO. It would be a costly gimmick to implement.
Also that same compiler still does it for x86.