Lwan: Experimental, scalable, high performance HTTP server
github.com
github.com
And, yes, that's pretty much it, although "rebimboca da parafuseta" is less obscure than "reticulation of the splines" (if you're Brazilian, anyway).
It is used to denote a fictitious part, which name or function are unknown, of a car engine or any other machine. It was coined in the 70s TV show, and later used in some ads aired during the same period; since then, it's an expression used for humorous effect. It's not very common these days but is unlikely you'll meet someone down here that never heard it.
It was coined for a cigar brand TV commercial back in the 90s and also alludes to a fictional part, in this case used by a car repairman to fraud a customer into paying more in said commercial.
* rawmemchr - it might be faster as it doesn't have to decrement the size_t but this only mildly relevant for lwan as many instances of rawmemchr use are simply rawmemchr(ptr, '\0') which is exactly the same as ptr + strlen(ptr) + 1 and even less optimized.
* pthread_tryjoin_np - __linux__ is defined by gcc not glibc, you should check for __GLIBC__ if you want to use glibc specific functions.
* underscore prefixed functions - pedantic I know but it is reserved for the implementation.
Regarding rawmemchr(): both are pretty well optimized. Both are implemented in glibc using the same technique (reading a byte at a time until it is aligned, then moving to multibyte reads). strlen() might be faster, yes, considering that the implementation can hardcode some magic numbers. In other words: some micro benchmarks might help decide here.
Regarding __linux__ vs. __GLIBC__: Lwan works with some alternative libcs (such as uClibc), so relying on __GLIBC__ being defined for things like this doesn't seem like a good idea. In any case, since Lwan isn't portable anyway, one can just assume it is always running on Linux and get rid of these #ifdefs.
I could see a specially written parser outperforming it, just like hand written assembly can still sometimes outperform a compiler.
I agree that it's more likely to have bugs, though.
The HTTP 1.1 RFC requires whitespace be stripped at the end of header values, yet also permits (although deprecates) header folding, giving rise to the following ambiguity (using _'s in place of leading spaces):
Foo: Hello\r\n
______\r\n
________\r\n
________ world!\r\n
Bar: smeg
This requires that the parser buffer all the whitespace between 'Hello' and 'world!' (and the RFC doesn't put a standard limit on header value length) just in case 'world!' never comes and the value of the Foo header has to be stripped back to just "Hello"Here's a related observation from a commit[0] by the Joyent guys, who wrote the streaming parser used by NodeJS: "For http-parser itself to confirm[sic] exactly would involve significant changes in order to synthesize replacement SP octets. Such changes are unlikely to be worth it to support what is an obscure and deprecated feature"
Another example is parsing the Request and Status lines:
GET <uri> HTTP/1.1
Technically <uri> can't contain spaces, but the RFC says you MAY accept them [RFC7230: 3.5. Message Parsing Robustness] ... which then gives rise to the possibility of <uri> containing the literal string " HTTP/1.1", and ultimately opens up bad user agents that send spaces to header injection.Resolving these ambiguities require implementing your own buffering, and dropping down to Ragels 'state charts' feature to avoid your semantic actions being munged... which leaves you to design the top level state machine yourself.
[0] https://github.com/joyent/http-parser/commit/5d9c3821729b194...
Foo: Hello World\r\n Bar: smeg
Foo: Hello_______world\r\n
Here the whitespace has to be preserved, as per the field-content production.So RFC2616 allowed:
Foo: Hello_______\r\n
____world\r\n
to be reduced to: Foo: Hello world\r\n
Incidentally RFC7230 (it's successor) says something subtly different:"A server that receives an obs-fold in a request message that is not within a message/http container MUST either reject the message by sending a 400 (Bad Request), preferably with a representation explaining that obsolete line folding is unacceptable, or replace each received obs-fold with one or more SP octets prior to interpreting the field value or forwarding the message downstream."
So now it's been loosened to one or more SP octets... useful, except there's no mention of TAB octets, which can follow a CRLF as part of an obs-fold... so you still can't just remove the CRLF and preserve all the whitespace... preserving tabs would be illegal. So the new rules don't help streaming either. Joyents parser did this regardless (not sure if it still does).
You'll also notice it says the obs-fold can be replaced with a sequence of SP characters... according to the grammar that's the CRLF and the following whitespace, not any whitespace preceding the CRLF. You'd think that would be helpful because it means any whitespace before the CRLF can always be streamed as part of the field value, right? Except...
Foo: Hello_______\r\n
____\r\n
would still simplify to: Foo: Hello_______[one or more SP octets]\r\n
and then what? Well, presumably you then have to trim off the trailing whitespace to produce a value of just "Hello"... so all the whitespace you buffered before you reached the CRLF (in case you reached another field-vchar like the 'w' in 'world') has to be discarded.However, the easy option is to default to the slow code path on edge cases which considering it's probably a rare enough it's not important to make fast as long it's bug free. IMO, optimizations are always a balancing act between minimizing the computer's effort and minimizing the coders effort while trying to maintain long term readability. But, you can always keep track of more than one buffer so the option is there.
Hm, I actually think the opposite (except for the point about bugs). A hand-crafted parser can easily be faster than a generic parser because:
* You aren't limited to implementing a regular grammar
* You can guarantee no overhead is introduced by the parsing tool
* You have complete freedom to optimize at the lowest levels
Now, I've never used Ragel, so I'm speaking from my experience with other parser generators (and from the perspective of a programming language implementation). I'd be interested to know if Ragel is different with regard to any of these points.
It'll do a better job than you will at reducing DFAs to their optimal representation. Imho, with careful use of its pragmas, it produces pretty much optimal code when using -G2 (goto based code generation). Another really nice feature is being able to dump the DFA to a dot file and render it using Graphviz.
Uh oh.
What license is aimed for it?
We can't use GPL or LGPL at the company, everything is statically linked.
http://lowlatencyweb.wordpress.com/2012/03/26/500000-request...
So can Google's compute engine on a $10 instance:
https://news.ycombinator.com/item?id=6804897
http://googlecloudplatform.blogspot.ca/2013/11/compute-engin...
OTOH, the beefiest machine I have access to test it is a 4 year old laptop, not a 24-core Xeon.
Google's Compute Engine test was using 200 virtual servers, but it does include network connectivity. The response body is a single byte. Their blog entry is a celebration of the performance of their load balancer more than a statement about the performance of each VM.
In March, we were able to exceed 1M requests with network connectivity and without pipelining to a single server [1]. Our project is not testing static web servers, so we don't test with plain nginx; but I expect nginx would also exceed 1M RPS in this hardware environment. This was using a server with 40 HT cores and a single-byte response body.
Similarly, a highly tuned web server such as OP's Lwan should be expected to exceed 1M RPS (network-connected) on a 40 HT core server. 1M RPS with small response payloads is fairly easy on modern hardware.
Incidentally, we see 6M+ RPS with pipelining in our Round 9 plaintext results [2].
[1] http://www.techempower.com/blog/2014/03/04/one-million-http-...
[2] http://www.techempower.com/benchmarks/#section=data-r9&hw=pe...
40 million req/sec with Lua
http://highscalability.com/blog/2014/2/13/snabb-switch-skip-...