Developing a compiler from scratch is the other significant use for writing it. Of course it is quite common to need to read it.
[1]: https://gcc.gnu.org/onlinedocs/gcc/x86-Built-in-Functions.ht...
Edit: also useful are https://gcc.gnu.org/onlinedocs/gcc/Extended-Asm.html and https://gcc.gnu.org/onlinedocs/gcc/Local-Register-Variables....
This is a notable differentiation - Writing assembly is a different skill to reading it from a disassembly. Reverse engineering, malware analysis etc. does not inherently require you to be able to write asm, although it certainly would help.
The pinning local variables to registers thing didn't work in llvm a couple of years ago (which seems consistent with the gcc docs) but does work at the boundaries of inline asm and that's generally enough. I like a pin-register intrinsic, something like `u64 pin(u64, enum reg)` where the compile time constant enum names the register and the semantics are a no-op other than constraining the register allocator, but that doesn't seem to be readily available in gcc/clang.
I don't have a good answer to constraining instruction scheduling.
On reflection it's all somewhat more horrible than it needs to be, perhaps inline compiler IR is a better idea.
My point is that conditional moves are one of usecases badly supported by compilers, and that (may) require dropping to assembly.
dan bernstein makes the argument that, as computers get faster, we use them on bigger problems, which means that computer performance is increasingly dominated by small inner loops, which is precisely the situation where it becomes more rational to put effort into hand-optimizing your small inner loops than to hack on the compiler to hopefully speed up all parts of the program, just as it was in the 01960s for different reasons
a different way to attack that problem in many cases is to write a domain-specific compiler from a domain-specific language to machine code, as thompson's regexp engine did, and as verilog compilers do. but i'm not sure how you speed up a media codec that way
bernstein has also written a fair bit of assembly to eliminate timing side-channel leaks from cryptographic code
Example: https://github.com/dddrrreee/cs140e-23win/blob/85b9ae3bd46c7...
Conservative garbage collectors. Scanning the native stack for pointers can be done in C but isn't quite enough since there might be pointers in registers. So I wrote assembly code to spill all the registers onto the stack prior to scanning.
I would agree that direct asm is very rare these days in game engines outside of third-party libraries. There can be significant gains with tuned asm but some combination of intrinsics and ISPC is usually good enough. But it is far more useful to be able to _read_ assembly, for debugging in an optimized build or analyzing release crashes.
I would have preferred to emit something like LLVM IR instead, but couldn't because of several constraints.
See the .s files in: https://cs.opensource.google/go/go/+/refs/tags/go1.21.5:src/...
I occasionally see it in compression as well.