Rwasa – A high-performance web server in x86_64 assembly
2ton.com.au
2ton.com.au
I've always thought writing assembly manually was for some very specific edge cases, or to talk to some very specific hardware, but that it was just a waste of time for anything else, especially compared to C (and especially with all the progress in compilers).
Is it the death by a thousand cuts scenario, or are there some big chunks of performances gained thanks to some specific tricks (and if so, could you give some example ?). I'm thinking maybe cryptographic functions ?
The short answer is: there is no single reason, it is the culmination of all of the underlying bits of the library that made it what it is.
I added code profiling to rwasa so that you can run load tests against it and watch call graphs and individual function timings, which makes for interesting inspection of the library itself for specific "tricks" that were employed. (the library's page contains rwasa-specific profiling examples: https://2ton.com.au/HeavyThing/ )
Have you compared it to Intel's optimized implementation of zlib, at https://github.com/jtkukunas/zlib/ ?
If you can improve on that implementation, please consider submitting a patch.
Regarding zlib, is that the fastest implementation that's currently available?
I rememeber stumbling into a guy that claimed his implementation was way faster (2x or more) than the original zlib, and it was a drop-in replacement and scaled fairly well. Unfortunately, I can't find it right now (it should be bookmarked on another computer).
Have you compared my code to [your?] repo yet? I will endeavour to do so, though I am not sure submitting a fasm-based patch to that repo would make sense. Cheers
My strategy, leveraging the Write Great Code book, was to map language constructs onto assembler via macro-assembler or languages such as LISP with good metaprogramming. Then, I hand compiled the code in other's projects with my macro's. So, we took a similar approach there although my goal was to show a correspondence argument between source & asm. Meta tools turned that into full program for assembling and linking.
One wild idea I had for portability was doing optimized routines of a safe HLL in a bytecode like LLVM. That knocks out most of the uncertainty of above layers that limit optimizations the most. Then, the simple optimizations that machines are good at can be performed from there along with generation of assembler. Close to portability of C and efficiency of hand-written assembler with inline available.
For instance, code zlib functions in pretty optimal LLVM and let toolchain do the rest on full-optimization. Think yours will be 25% faster, 10%, similar? And I'm talking what you can code quickly rather than spend 30min-1hr optimizing by hand. Just curious to hear your thoughts as you have way more experience in that stuff.
They're different things. If it's a worthwhile path, then the formal efforts on LLVM IR and verified optimizations could be combined with hand-made, inline LLVM for a safer, portable alternative to different inline assembler for each platform. Complimentary, not contradictory, when in context.
What do you think of that?
I think most beliefs regarding x86_64 asembly is largely "guilt by association" with i386... It's amazing the difference just from making use of the larger register set.
Yeah, but that is a very CPU bound processing pipeline. You would expect that to maximize the impact of any inefficiencies in the compiler.
That you can hand tune for better performance is conceivable. That you can get a 2x win over some pretty tuned code suggests that there is something larger at work than simply tuning lots of little things.
* Allocate registers globally or across large substs. Especially when targeting architectures (like x86_64, but unlike i386) with decent numbers of available registers, this has lots of potential for typical applications that e.g. frequently needs to access common data structures. Compilers for many languages (e.g. C) have a hard time doing this if doing separate compilation per module (since you need to be able to link to code that hasn't been optimized the same way). You need whole-program optimization for this typically, but when programming assembler, it's the natural thing to do if you have enough registers to treat some of them as assigned to specific variables that you know will be frequently accessed.
* Omit stack frames entirely or selectively. Many compilers have options for doing this, but often still ends up pushing/popping more stuff than necessary for things like local variable frames and arguments, where a programmer will often see that a function is not going to use much space and decide to put in extra effort to shuffle things around to keep things in registers only.
* Selectively violate the "normal" calling conventions. E.g. if you have a utility function you often need to use in settings where it's convenient not to clobber certain registers, then you can opt to pass arguments in different registers easily. This again takes whole-program optimization for a compiler to do.
* Avoid saving/restoring certain registers based on the functions you're calling. Again requires whome-program optimization for the compiler to know that the function you're calling won't clobber specific registers.
* Specifically adjust what registers you're using based on what registers may be clobbered by other code you're calling to avoid having to save/restore.
Other things include similarities/patterns in code that are non-obvious in a higher level language because it depends on how the code is translated, which often can allow you to re-write things to eliminate common sub-expressions that are not actually visible/present in the high level code.
It's not that compilers can't do all of these if given sufficient freedom and information, but it often violates expectations of the higher level environment (e.g. separate compilation in C)
E.g. just being able to determine application wide what call sites exists for a given function makes it fairly simple to handle specialized calling conventions, throw away stack frames, avoid saving/loading registers etc. as the biggest barrier against this is that with separate compilation you don't know upfront if a given function will be called from some piece of code expecting standard calling conventions.
You certainly can do really expensive optimizations too that might benefit from saving information, though.
* "unusual" control-flow: return several levels up from a function without requiring extensive elaborate exception-handling mechanisms, coroutines, and techniques similar to continuation-passing-style. After all, functions and procedures are just an artificial construct imposed by HLLs.
* Easily return multiple values from a function by using several registers, even normally inaccessible ones like EFLAGS (very useful for booleans.)
* Generating self-modifying-code, like a simple JIT compiler, is straightforward to do. Works especially well for tight loops that have several variants of their bodies.
Compilers are still bound by conventions and the features the HLL exposes. Asm is only limited by what the CPU can do (and what the programmer can come up with.) I admit that, while I do prefer using something like C for much of the "mundane" code I write, it's quite frustrating in those situations where I can think of a very elegant way to do something that either can't be expressed in C without some extreme compiler-fighting, or is completely impossible because of how it generates code and what the language allows.
(although I don't know anything about a feature comparison)
The -logpath option is fine, but it would be nice if it could create subdirectories too (e.g. $LOGPATH/YYYY/mm/access.log.YYYYmmdd ). Otherwise, over time the log dir is going to get unwieldy.
I'm currently running several sites in alpine+rwasa Docker containers; I'm liking having a set of entirely isolated web servers based on a 10MB container image, each apparently consuming ~6KB RAM while idle.
An example of how people let logrotate do most of the lifting is in the Chef server codebase.
I've always wondered, how many man hours does writing a small web server like that take?
Re: How much time, hard to say from my perspective since a fair amount of rwasa's functionality resides in the library itself. Start to finish for all of the showcase pieces and the library (from 0 lines of code to release) took 13 months of my life :-)
I also want to add I love the documentation.
I'd also like to say, as someone who is quite ignorant of writing x86 assembly, the function hook example is incredibly readable and clear. I'm looking forward to grokking the rest of the code base in attempt to learn more.
Thanks for the hard work!
It would be interesting to see how x86_64 implementation of specific elements compare to equivalent C code, in terms of instruction per clock efficiency, cache miss ratio, etc. (e.g. using the "perf" tool in Linux, or any other tool that use the CPU hardware counters).
1. Is there a config file ? Like "nginx.conf" for example.
2. Is "reverse proxy" possible ? Like nginx reverse proxy.
Some examples would help.
1. There is no configuration file. 2. Yes, you can set an upstream (-backpath).
Good luck with the project!
$ ssh 2ton.com.auJust kidding, nice effort!
EDIT: Here's the header of the file that contains their TLS implementation. I'll let you be the judge: https://gist.github.com/SirCmpwn/ec8aaec128aa3e47ddda
I haven't read much of your TLS implementation, but I'm not a security researcher and I don't think I'm qualified to give an opinion on your particular implementation. However, there are some points to be made here. First of all, almost no one ships their own crypto for a good reason. To trust that something is secure, you need to have lots of people working on it and lots of projects invested in it (something that OpenSSL and such have). "Many eyes makes bugs shallow" is the common phrase, but it holds truer when large companies with highly skilled engineers are putting their secrets and their customer's secrets on the line. Your implementation has the unfortunate problem of being written in assembly. While I don't think there's anything inherently wrong with assembly, many people don't share the same opinion. People will be reluctant to contribute to it (how do people contribute, anyway?). Also, it is easier to make mistakes in assembly, and I would be surprised if there weren't several mistakes in your TLS implementation and in the rest of the server. This doesn't reflect poorly on you as a developer, but is instead a consequence of choosing assembly.
Assembly also inherits a lot of the common security problems C has (like buffer overflows), but makes them harder to identify. I would feel very uncomfortable exposing anything written in assembly to the public net, and doubly so if it used an unproven TLS stack written in the same. Other projects avoid the problem of untested crypto by using tested crypto from an external module like OpenSSL.
Re: how do people contribute, it is on my list of things to do for github's linguist x86_64 support (which is why I didn't put it all on github to begin with).
At the end of the day, trust is a function of time and perceived scrutiny of the stacks at hand. We are getting there slowly but surely :-) Cheers!
Only after massive security vulnerabilties and a lot of media attention more people looked at OpenSSL in detail and eventually decided to do something about it. Which mostly was "let's write a new library or fork it". Thus, LibreSSL, sodium, nacl and such.
What got us in the mess with OpenSSL, was to leave a key component of many software projects to a struggling, small team. It's amazing how much open source relies on a few ancient programs written and maintained by few with little to no financial support (e.g. NTP, GPG).
Right, but those vulnerabilities _were_ found. I worry that they wouldn't be found in this. The only people who'd go looking are the people who see that a specific website is using it and want to exploit it.
Actually, I'd say that it has advantages because the mindset is very different when writing Asm - it naturally forces you to think about things in a low-level and precise fashion, which keeps considerations such as buffer lengths more in the mind than higher-level languages that attempt to abstract it away.
Programming at the instruction level also allows much more fine tuning of instruction ordering and such to resist timing attacks, without any compiler optimisations getting in the way.
This is why people usually prefer to stick to the "regular" libraries even if they're known to be slower.
> BREACH/TIME/etc > > Both the BREACH and TIME attacks rely on measuring the size of compressed response bodies. Since rwasa supports dynamic content compression by default, the HeavyThing library's default setting for webserver_breach_mitigation is enabled and set to 48 bytes. For each rwasa response when TLS and gzip is active, this setting adds an X-NB header that contains a random 0..48 bytes that is hex-encoded to each response header. While this doesn't render response sizing attacks completely useless, it makes a would-be attacker's job much more difficult due to the highly variable response lengths.
It's my understanding that random padding doesn't in fact make the attacker's job "much more" difficult. Only a little more, or not at all?
Could you comment on how integrated the TLS stack is with the webserver? Normally I'd think that using some kind of dedicated SSL terminating proxy, either a new version of HAproxy -- or stunnel/stud or similar -- would make more sense than deploying a new TLS stack that hasn't been through any outside review?
That said, as mentioned by others here - openssl is clearly not a great example of a secure/good TLS implementation. I'm not sure there are any (yet). Hopefully libressl will become one. Personally I'd like to see a minimal library that combined a couple of AES/ECC primitives and implemented TLS 1.2+ only (No SSL), with a sane and clean API on top.
Something along the lines of NaCl but with a goal to support a subset of standard TLS with forward secrecy (and explicitly throw old clients under the bus, Android 2x be dammed).
The BREACH attack verbage at http://breachattack.com spells it out fairly clearly, by adding random bytes to all of the HTTP responses, it makes small compressed HTTP payloads impossible to determine whether guessed bytes were correct or not (well, depending of course on the size variable of the random bytes added).
> Could you comment on how integrated the TLS stack is with the webserver?
The TLS layer is entirely separate from the webserver layer. I built the epoll, TLS, SSH, webserver and client as "IO layers", such that they can be stacked together arbitrarily (imagine epoll/IPv4 listener -> TLS -> SSH -> TLS -> Webserver, perfectly doable, albeit a little nutty).
Hm, ok. At least you didn't "just add some random padding" :-)
Thanks for the comment on structure. Might be nice to try and make ssl/tls terminating proxy as a separate binary I guess.
As for the code, for someone new to fasm it wasn't immediately obvious that to build one had to assemble then link (fasm -m $((bignumber))[1] project.asm project.o && ld -o project project.o # optionally strip project). Might want to but that in a Readme/makefile/build.sh. I found the general recipe in the hello-example - but a short readme in the various project folder and/or top level wouldn't hurt.
[1] ed: from https://2ton.com.au/HeavyThing/#echoserver
fasm -m 262144 echo.asm && ld -o echo echo.o
I think you should just finish the job and implement the entire OS in assembly. ;-)
I'm kidding. To me, assembly programming has always seemed like a true art form. You're forced to think about everything, and if you can successfully fit all the pieces together properly, it's beautiful. Also, not many can hack it through assembly, so there's a huge selection bias too.
It increases the number of requests required, yes, but I don't see why it makes it impossible.
I say "I know" because there may be something out there. I know there exists tools for source code analysis that deal directly with assembler, because I see theses around writing them. I don't know where to get one, though, or how to use it, or how much to trust it. There's a lot more that deal with "C" or "Java" rather than raw assembler, so I can fire several tools, both commercial and open source, at the problem.
And all that said I'm still extremely strongly in the "STOP WRITING C CODE AND PUTTING IT ON THE INTERNET DAMMIT" side, even with that support. Without even that support, I frankly don't care if it's 100 times faster than nginx. nginx is already maxing out my risk tolerance as it is, and I've begun a long, slow program to get it out of my stack too.
I want to emphasize how this is explicitly (but also unapologetically) non-specific, none of this is personal, none of this is directly critical of your code (because if you've gotten the impression I haven't even glanced at it, that is correct), and in particular, please by all means do whatever you like with your spare time. The "problem" here isn't that you have somehow failed to leap my bar, the problem is that the bar is impractically high for code written in raw assembler. I suppose you could provide a math proof but I'd almost argue in that case the server becomes implemented in said proof language rather than assembler anymore.
It's a long slow process partially precisely because I intend to do this carefully, and discarding nginx is not something to be taken lightly, but long term, as I said, I want the C out of my stack.
In the case of an ASM project, I would be very surprised if anyone came along with the appropriate knowledge to ever want to touch the codebase. LibreSSL is currently pulling bits of ASM out of the codebase just to remove that factor.
Like Jerf said, I want to be clear that I'm incredibly impressed you got this project over the line, and I can't make any complaint about how you've done things.