Asmttpd: Web server for Linux written in amd64 assembly (2017)
github.com
github.com
Yes, it's assembly, but it's basically a thin script over a set of string library functions which are themselves just naive versions of what's already in libc. Most of this is just strcpy/strcat/strchr/strstr; the only reason it's not one giant stack overflow is that it doesn't actually implement most of HTTP (for instance: nothing reads Content-length, and everything works off a static buffer that is sized to be as large as the maximum request it will read). What's the point of writing assembly if you're going to repeatedly scan to the end of strings as you try to build them?
I'm not taking away anything from the project as an exercise. It's neat! But like: it's not something you'd ever want to use, or a productive path to go down to make a good webserver.
I'd have to benchmark to make any claims about speed, but from my past, quite extensive experience with using Asm, it's not uncommon for even a simple/naive/asymptotically less efficient algorithm written in Asm to outperform a more complex/clever/asymptotically efficient one in an HLL at the problem sizes encountered in practice, simply by virtue of the constant factor being much smaller. For the same reason a bubble/insertion sort in Asm can easily beat the standard library sort if your inputs aren't that big.
Of course if you combine Asm with efficient algorithms (which can sometimes be easier to write than in a HLL), then you can get much closer to how much the CPU can actually do.
nginx and Apache are both written in C, but nginx can outperform/outscale Apache. That's because of high-level architectural differences, not from tweaking assembly code. An HTTP server isn't doing much in data-transformation/algorithmic terms, it just has to scale efficiently.
If your code is synchronous and non-parallel, it's going to lose every time, even if it's highly tuned assembly.
You're conflating two different things: constant factors based on algorithm choice and the merits of hand-written assembler. It's true that sometimes asymptotically worse algorithms beat asymptotically better ones in practice. It's a lot less clear that humans can regularly beat compilers for typical scalar code.
I wonder how efficient it would be to spawn a few hundreds.
I proposed to debian (on an IRC channel for that) to make a package for rwasa but I've been told that "nobody does asm in 2017". I wonder if it's still true in 2018.
Picolisp 64bit
Luajit (mostly)
All sources are really well outlined, and worth the read through. But you might start with the 64bit's README. [0]
Bash is undeniably the language with the highest possible stability, you can literally feel the stability and security when a file is served and the hard drives go crazy.
%macro stackpush 0
push rdi
push rsi
push rdx
push r10
push r8
push r9
push rbx
push rcx
%endmacro
%macro stackpop 0
pop rcx
pop rbx
pop r9
pop r8
pop r10
pop rdx
pop rsi
pop rdi
%endmacro
I guess that's one way of saving registers. Not particularly efficient, I guess, but it works…The same goes for the longer sequence of individual pushes or pops --- they all depend on the stack pointer. In fact, the single instruction needs to only adjust it once by the total number of registers pushed/popped (since it is a constant).
In other words, PUSHA/POPA already decode internally to a bunch of moves and one ALU op for the stack pointer which can be scheduled OoO. All they needed to do for 64-bit mode was adjust the constant (by multiplying it by two) and emit more uops for the additional registers, but they didn't for some otherwise inexplicable reason. All the machinery to do it was existing.
And then there's the fact that modern compilers are likely already avoiding those instructions because of the timing of saving extra registers that you aren't using in a given function, it probably just doesn't make much sense to keep them anymore.
Compilers won't ever use PUSHA/POPA but lots of other code will --- BIOS, executable packers, OS state-saving code (the perfect example of what these instructions were for?), etc.
See the story of SAHF/LAHF for a similar and even more astoundingly bad decision.
ps: thank you all for the answers
STMFD sp!, {r3-r7,lr}
(I believe the "FD" suffix is to do with the stack growing down - "full descending")Of course this is just implemented with microcode so it's not really any more efficient than a series of PUSHes, except there's a bit less I-cache pressure.
Talking of the ARM reminds me of older heroic feats of assembly-programmed internet software in the Acorn days, such as Ben Dooks' all-asm TCP/IP stack and Jon Ribbens' web browser. Probably not been done that often...
It's actually not vulnerable to a bunch of other attacks as well (e.g. a buffer overflow cannot overwrite the return address on Itanium).
http://landley.net/history/mirror/cpm/z80.html (search for "alternate registers").
> As the two LOADALL instructions were never documented and do not exist on later processors, the opcodes were reused in the AMD64 architecture.[8] The opcode for the 286 LOADALL instruction, 0F05, became the AMD64 instruction SYSCALL; the 386 LOADALL instruction, 0F07, became the SYSRET instruction. These definitions were cemented even on Intel CPUs with the introduction of the Intel 64 implementation of AMD64.[9]
[snip]
> Because LOADALL did not perform any checks on the validity of the data loaded into processor registers, it was possible to load a processor state that could not be normally entered, such as using real mode (PE=0) together with paging (PG=1) on 386-class CPUs.[7]
I noticed that certain options are hardcoded (such as the port number). As someone with very little experience in assembly, how difficult would it be to make this dynamic via environment variables?
Why would anyone want environment variables when you can have config files or explicit options instead?
Consider, for example, wanting different options for less between when it's called by you, or by man, or by git, or by journalctl, or any other utility.
alias man='LESS="X$LESS" man'
alias journalctl='LESS="S$LESS" journalctl'
It's also useful when you want all processes that follow from an invocation to share a certain configuration, and when they follow from another invocation to have another. An example of this is $DISPLAY. I startx on my tty1 and that invokes the Xorg server, exports DISPLAY with the display of the server, typically :0, and executes my window manager. When I press certain shortcut keys, applications will be called, and they, in turn, might call other applications, and they all need to know that when they open up a window it should be on display :0. If I login remotely through ssh, forwarding my X11 service, the applications that I call on the same machine where other processes of the same applications opened their windows on :0, should now open their windows on localhost:10, so I get to see them on the laptop I'm logging in from. If someone else wants to use their account on the same computer, they might switch to tty2 and startx themselves, and all the applications they open should open on their Xorg server process with display :1, not mixing with the other windows I opened locally or on the laptop.This is also why $LANG, $HOME, $USER, and $TERM are useful. They're stuff that should shared by process trees. Their purpose can't be fulfilled by configuration files or command line options.
Pretty much every software supports several config files (system wide, invocation specific) for this reason.
It's not harder to write a line to a file than to set an environment variable.
Environment variables have opaque size limitations. Any non-trivial data is likely to be quoted or encoded in some way, which is going to vary between your applications, and there's going to be quirks in parsing.
It is much easier to programmatically reason about state of standard file formats on disk, instead of resorting to parsing environments of running processes.
Environment variables is inherited to child processes (which is the point of using them). That's going to have security implications for you. Especially if those processes also read their configuration from environment.
And yes, given the choice, I think you will find that most people uses .gitconfig instead of setting their git configuration via environment variables, for these very reasons.
Your post is weird to me, because you seem to have interpreted that I somehow implied that we shouldn't use configuration files and that everything should be done through environment variables. That's not at all the case.
On the other hand, your post seems to imply that environment variables should never be used. Is that so? Because if it is, I'd like to know how you'd think that $DISPLAY or $TERM could be better substituted by another solution. I'm not even going to limit you to config files and command line options. I've already painted a concrete scenario of the use of $DISPLAY in my previous post, can you paint that same scenario with another solution?
Now to paint a scenario of the use of $TERM. $TERM controls how terminal applications communicate with the terminal to use its features like clearing the screen, moving the cursor around, etc. Imagine one server and multiple people connecting to it through ssh in each of their computers. Each person uses a different terminal in their machines. The server terminal applications use $TERM to determine how to use the different features of the terminals that are connecting to it. These are not only the applications that the user launches through the shell, but also the ones that are called indirectly via other programs that the user has invoked. Also keep in mind that the server might simply not recognize the type of terminal you have (it doesn't have the appropriate terminfo file installed), but you might know another terminal that is similar to yours and that the server might know. In that case, you'll want to inform all programs you launch and the ones they launch to treat your terminal as if it were that other terminal (this is when you'd set $TERM in the shell session to something else). Can you rethink this so the terminal applications can use something other than $TERM to determine the type of terminal that's in use and still allow the user to override it?
> Pretty much every software supports several config files (system wide, invocation specific) for this reason.
How would the program know what configuration file to use for invocation specific configuration? If it's not through an environment variable, then I guess you specify it by command line option, and, like I said, that's not going to help when you don't control the options that are passed to the program when it's another program and not yourself that makes the call.
If you have an underlying .so used by an application that hasn't implemented optioning at init of said library it can then let you control things in something that has no runtime configuration source. I've written several such libraries. Also good for controlling things you override with LD_PRELOAD
Global options > user options > env variables > command line parameters
Env variables cascade, which is why they have wider uses than program args. You can organize a system as a group of processes and envs are a good way to share state and abstract out the details like prod vs staging without affecting the program implementation.
This way, my bash environment mirrors the task I am working on
My wife assures me that size doesn't matter, it's actually the way you use your configs to manage your environmental vars.
(sorry, I just couldn't help myself...)
Configuration files don't let us override individual variables easily, in different invocations of the program. We have to generate a custom version of the configuration file for that job and pass that file's name to it.
Environment variables are only visible to children, so they are inherently secure; we don't worry whether permissions are too loose on an environment variable so that another user could see it or modify it. There is no such thing.
In Linux, they're visible in /proc/<pid>/environ, which is permission 400 and owned by the effective user of the process.
> when you can have config files
You still need a way to tell your program where the config file is.
Environment variables are just like a configuration file that's automatically opened by the OS when a process start and automatically parsed as a set of named strings by the libc. In the many cases where one does not need more complex data structures than that, then to introduce a config file is just making things harder.
That being said, I wish envvars had namespaces of some sort.
But if you're already doing the rest of it in asm why would that scare you?
x86-64 is a great architecture for many applications. It's cheap (due to volume), well-understood, has great compatibility, and is pretty much the best-supported architecture when it comes to software.
Edit: that being said, no, this doesn't have any embedded use today. Embedded x86 isn't 386 processors anymore. The typical scenario for an embedded x86-64 is "I need to drive this specific application and I want to leverage the Linux/FreeBSD/Windows environment for programming". That already puts you way ahead of the point where writing assembly code is useful performance-wise, and you probably don't want to put ASM code you downloaded off the Internet in a network-facing position :-).
Everything needs to start somewhere.
(Or rather the weakest, simplest, most dynamic type system you can imagine.)
But many programmer mistakes that a C compiler would at the very least warn about will only be apparent at runtime in assembly: Treating something as the wrong type, forgetting to dereference a pointer (or accidentally dereferencing it), mixing up values of different types, mutating something that's meant to be constant in a context, assigning an integer quantity to a pointer without an explicit cast...
Then on top of that you get a myriad of new potential mistakes that are impossible or at least very unlikely in C: Mixing up registers, accessing the wrong offset in a structure (e.g. easier to mix up reg+12 and reg+14 than foo.enabled and foo.name), messing up the stack layout... And now think about what each of those mean when you're actually changing code: Did you update every register? Did you update every offset?
As for the simpler code being simpler to understand and simpler to keep bug-free, that's kind of a tautology that applies independently of the language. And besides, wouldn't those qualities be even easier to achieve for the same code written in C?
Integer overflows are undefined behaviour in standard C.
That's not true if running complex assembly through an assembler and/or linker. Aside from totally changing the program, there's definitely potential for it to be modified to do things like drop safety checks for optimizations or use instruction patterns that increase covert channels. That's on top of not working correctly at all. I'll add to that point that the Intel documentation is huge with every processor having errata. That's why high-assurance security always considers those things inside the Trusted Computing Base (TCB).
Can you please give some link? It sounds strange, I’ve never read about something like that happening in assembly (“drop safety checks for optimization” — which checks, how can that even happen? Maybe you confused that with the writings about what C compilers do?).
The assembler or linker are programmed to look for and drop checks for stuff like buffer overflow.
A "macro" assembler has some code generation. It might generate the code in an insecure way.
An malicious assembler might swap out instructions that prevent side channels for equivalent instructions that cause them.
There's lots of ways to screw up assembly code in general on correctness side that might at least lead to DOS. The malicious assembler might do any of them.
The linker might do any of that. It could even be a better place for it since few people understand or look at linkers. Although memory is fuzzy right now, you could also look for any security errors that started with the linker. They could be done on purpose.
So, I would expect a formal specification of each component, their correctness/safety/security requirements (policy), the implementation, and evidence the implementation maintains the spec and policy. CompCert and CakeML are example compilers that do that for correctness. I can't recall if someone verified a x86 assembler and linker. Rockwell-Collins is best example of your CPU/microcode being verified in itself and against specs of your application code function by function. I've also seen formal work on linking, mainly with type systems, in Standard ML that a lot of proof systems output. Most just ignore assemblers and linkers for x86 for some reason. Be a great project for some people to jump on! :)
What I understand there verified is however not that the adversary can't add malicious code but that there are formal verification steps that no bugs are introduced as the non-subverted implementation is introduced.
I agree with you that the linkers are a nice place for the malicious code to be hidden. The malicious code can be anywhere in practice... the way I see it, the accidental errors that introduce security issues are much more probable in compilers and the modern linkers (especially as they do more of a job that formerly only compilers did) than in assemblers.
The assemblers are surely the easiest to be verified on producing what is expected of them to produce: a separately developed disassembler is enough.
What they did is a high-level view (formal spec) of exactly what it does, one of what it means to be secure, and proof the design matches that. It's harder to subvert. It can also prove the system can't be attacked if they model those properties. For instance, they do show some protection against leaks from one process to another with the non-interference property of their partitioning mechanism. I just posted a lot more stuff like that in hardware and software here if you're interested:
https://news.ycombinator.com/item?id=17530829
" the way I see it, the accidental errors that introduce security issues are much more probable in compilers and the modern linkers (especially as they do more of a job that formerly only compilers did) than in assemblers."
I agree. I bring it up every time someone erroneously focuses on the Karger/Thompson attack on compilers. The more common error is the compiler messing the code up. An easier subversion would be to put a bug into a compiler that did that which looked like a normal bug. It's why certifying compilers are necessary. Note that TALC is part of a bigger picture at FLINT team where they do certifying compilers for stuff like ML, safe versions of C, intermediate/assembly languages that are typed for extra layer of defense, I/O, interrupts, and many other things. Here's their work:
http://flint.cs.yale.edu/flint/publications/
http://flint.cs.yale.edu/flint/software.html
Yeah, assemblers should be pretty easy if they're just doing straight-forward stuff preserving the data and structure. Those like Hyde's High-Level Assembly (HLA) might be trickier with high-level functionality. However, even those might be divided between a first-pass acting as a compiler for high-level stuff and a simpler assembler.
[1]: https://jamesmunns.com/blog/tinyrocket/ [2]: https://users.rust-lang.org/t/rust-binary-sizes-once-again/1...
Interestingly, the repo notes the binary is only 6kb. I doubt C or C++ will get anywhere close to that.
NOTE: I'm not saying it IS or ISN'T, and although I CAN grok asm, I'm not going to take the time to do it right now...
Just struck me as possible, given the idea of a high-level language is to mainly provide abstraction over implementation details ( point being, you can't really get any MORE implementation detail specific than asm... )
(edit: words...)