Web server for Linux written in amd64 assembly
github.com
github.com
https://github.com/nemasu/asmttpd/blob/master/http.asm
It dawns on me why we couldn't have shortcuts for several patterns that show up everywhere:
- mov, mov, mov then call/syscall could just be written as call(arg, arg, arg) since it's not that difficult to figure out which argument needs to go to which register if there was a defined order of arguments.
- push push push push <function body> pop pop pop pop <ret> could just be a define really. I realize that there are cases where you wouldn't necessarily do that, but that seems to be the minority. The compiler could just figure out what registers you use in the routine and push / pop those. If you want to keep a register, there could be added syntax for that.
It seems in both of these cases the language optimizes for simplicity and flexibility, ignoring the common case. Neither of these strike me as situations where introducing the abstraction would require a lot of compiler "magic" to guess and optimize. It's almost just string replacement.
Then again, you could just write all this in C and let the compiler figure it out. Interesting stuff!
And there is Randall Hyde's High Level Assembly: http://www.plantation-productions.com/Webster/
There were macro assemblers already available in the 80's!
Born in 1981, still being updated today:
Was this book (the first review complains about the same thing): http://www.amazon.com/gp/product/0130255963/ref=as_li_ss_tl?...
I think the use of m4 needlessly complicated the material.
"Turbo Assembler is a native c64 assembler which was introduced in 1985 by the German company, Omikron."
What?!
Plus it is BSD licensed for those that hate the GPL.
[1] http://git.videolan.org/?p=x264.git;a=blob;f=common/x86/x86i...
xor rax,rax
I used this construct often to zero a register, in the time that memory and CPU cycles were scarce. But nowadays, my time is a more valuable resource, and I tend to write: mov rax, 0
It takes a somewhat longer instruction code, and a few CPU cycles more, but it conveys meaning better.So do webservers written in Python. I don't think that was the point of this project.
Then again almost everything about assembly gets a bit "isn't that what we made other languages for?"
now, you could argue if that were that the case, wth is that person doing there anyway.
also if you compile something like `return 0;` it used to be compiled down to `xor eax,eax; ret;`, but meh barely anyone i know coding these days even knew this to begin with.
For modern out-of-order CPUs - xor reg1,reg2 is problematic because the result depends on the previous contents of the register. Hence it cannot be executed out of order.
However as a special case xor reg1,reg1 (along with sub reg1,reg1) will be detected by the cpu (intel anyway) as a 'zero idiom' instruction and because it has a smaller opcode than mov reg,0 its preferred.
There are also other complex reasons - for detailed info see 3.5.1.8 in the Intel optimization manual.
The initial claim was that mov reg,0 was slower to execute.
You'll get way more bang for your buck by minimizing userspace→kernel round trips and memory copies than by hand-optimizing assembler code. sendfile is one great way to do this.
You'll get even more bang for your buck by eliminating the kernel from the packet processing path by using netmap, PF_RING/DNA, or DPDK, and a user-space TCP/IP stack.
Assembly only really helps alleviate GCC's moronic decisions resulting in excessive stack spills and alignment-unaware loads & stores.
Something about carts and horses.
If you need to hit the metal you can inline assembly in both c and c++.
And why isn't that enough?
I wrote a web server in C. I'm not really a web guy, it was a lot of new stuff to me, I'm a pretty amateurish programmer, and it probably doesn't even deserve to be called a "web server". But it was fun, and really cool to know that something that I wrote (with the help/jump-off point of a tutorial or two) can be used to serve web pages to a client. I can run the program, pop open Firefox, and use the browser to click through a set of test pages as though it was being served by a real web server. That's fucking cool, and that was all the reason that I needed to do it.
Don't Reinvent The Wheel - unless you want to!
I am pretty sure that this was an exercise in practical user space assembly programing rather than an attempt to write yet another http server.
The thing about compilers is that they're leveraging, even if imperfectly, the collective wisdom of their authors and of the companies who actually built the chips and have offered insight, advice, and sometimes even code. It's very probable they know more performance tricks than you do.
One problem is landmines in the ISA, such as instructions that look like they exist to be used, but are really traps implemented in suboptimal microcode for the unwary programmer who didn't look closely at their performance characteristics. Or certain sequences of instructions that might combine to do something ridiculously slow[1].
These landmines vary by microarchitecture. An instruction that's incredibly slow on one line of x86 chips might be a wonder-drug on another. This both increases the probability that your code will hit a landmine on at least some CPUs, and gives you a possible "in": Compilers aren't going to optimize perfectly for every microarchitecture. If you know exactly what you're doing (or spend a hell of a lot of time on trial and error), you might be able to come up with optimal codepaths for specific chips that the compiler didn't.
By and large it's not worth it, though. Hand-tuned assembly still ends up in places, but increasingly rarely, and it's confined to small hot-spots. A particular algorithm or part of an algorithm gets re-implemented in assembly because the compiler just can't get it right.
[1] I could have sworn there was a story about this just recently, but I can't seem to find it. Something like a piece of code running way slower than anyone thought it should, until an AMD engineer piped up and said "Oh yeah, don't do that, it causes a pipeline flush." for reasons that were utterly non-obvious to anyone who didn't know the internals of the chip.
Sometimes, but the biggest case is if you can carefully arrange a tight inner loop, especially one that case make use of SIMD, like some DSP and scientific-computing code. Auto-vectorizers are getting better, but still miss lots of cases, so a skilled asm programmer can beat the compiler. The more "spread out" the performance-critical code is, in general (i.e. performance not dominated by one or two tight loops), the harder it is for hand-coding asm to beat a compiler; humans are not that good at doing whole-program optimization on large codebases. The more cross-platform the code has to be, the worse for the asm programmer as well: beating gcc's code-gen on one architecture is easier than beating it everywhere.
As one example, check out the content-type detection, which is essentially a long chain of repeated strlen + strcmp; assembly language doesn't magically make bad algorithms fast.
Using Document Root: htdocs/ An error has occured, exiting