Nxweb – Fast and Lightweight Web Server
nxweb.org
nxweb.org
While Nxweb looks very promising, my first question would be 'Why should I use it over eg Nginx?' It would be helpful to have some direct comparison to other servers on the landing page.
EDIT: Ok, there is a link to some odd benchmarks and it includes performance comparisons to Nginx and others which are not understandable (Nginx 141 req/s and Nxweb 200 / 121 req/s while it's not clear when 200 and when 121); moreover they compare it to Mongoose which is an ORM/ODM
However, that doesn't change the why question at all. Except it could be neat to not need the complication of setting up nginx with uwsgi for those who like the built-in Python/WSGI support.
[1] https://groups.google.com/d/msg/openresty-en/aoBL22H8fP4/bJ3...
Technically, there's no good reason why a FastCGI based system would be significantly slower than a custom reimplementation like this.
An HTTP server speaking to a FastCGI application will:
• read the HTTP message
• decode HTTP
• encode FastCGI
• write to application
The FastCGI application will then:
• read the FastCGI message
• decode the FastCGI message
• do application stuff
• encode the FastCGI response
• write to the web server
The webserver then resumes:
• reading the FastCGI response
• decoding the FastCGI response
• writing the HTTP response
Meanwhile, an in-process HTTP server (like nxweb) system will simply:
• read the HTTP message
• decode HTTP
• do application stuff
• write the HTTP response
Less code runs faster; it is obvious to me why this is faster.
In those cases you are correct: parsing and de-parsing is insignificant compared to the amount of energy the computer is using to heat the room.
However in order to do a trillion requests per day you need around 30 machines using a custom web server, or 300 machines using Fastcgi: In this situation the cost is an order of magnitude.
RTB systems have about 30-100msec for the entire transaction (and that includes network to the user), so you need better control of your latency anyway.
Similarly, I've noticed that people tend to get a little silly about web server requests-per-second. It really gets to the point you probably ought to be talking about seconds per request, or perhaps rather, microseconds per request or something.
Because A: as you start talking about these fast servers, you need to contemplate whether your code can run in, say, 2.5 microseconds either; who cares whether your webserver takes 2 or 25 microseconds to handle a minimal request if your minimal response requires 8 milliseconds (i.e. "8000 microseconds")? 8ms would actually be pretty decent performance for a wide variety of non-trivial web requests.
And B: As the webservers get faster and faster, you really need to start wondering what corners they cut to push their reqs/s number up. I can make a blazingly fast webserver that would actually kill nginx's performance stone dead for a "return a constant JSON string response" task... the trick is that I'm not even going to look at the incoming web request, I'm going to just receive a socket, blast out my answer as a constant string buffer without even reading from the socket, and discard the socket. (If you're feeling particularly saucy, hook that up to a user-space TCP stack so you can drop the work of properly setting up and tearing down TCP connections.) There aren't that many real-world tasks for which that is a good solution (though, non-zero!), but it'll look like pure awesomesauce on the benchmark!
Properly handling HTTP is non-trivial problem, and even moreso if it's going to be hooked up to a program rather than a static file system or something similarly easy. I actually start getting nervous about web servers that show excessively high numbers. If your performance is much better than nginx, rather than me cheering for joy, I actually have a lot of questions about how you did that exactly, and what my website's security profile looks like with your way-faster server. I'm not saying these questions are completely unanswerable; perhaps there is a way to safely do a much faster web server. I'm just saying that rather than my default response being celebration and "Oh wowzers cool!", my default reaction is a healthy dollop of skepticism.
You suggest instead that you get the old truck serviced and replace the plugs, distributor and tailpipe. Estimated cost $1000, and should get you from 10MPG to 11MPG. Which is the better deal? Assuming you both drive about 100 miles per week.
Running two web servers (one speaking HTTP and one speaking FastCGI) is necessarily going to be slower than running one web server.
This should be obvious, although it might be "not significantly slower", which is why I provided some real numbers from my experience to show at which point it becomes slower by an order of magnitude.
You might also find that it's easier to debug one webserver than two.
It's very safe as far as I can tell having run it under AFL with no crashes with ASan on as well as having run it in production on the public Internet.
A few of the optimizations I do in "filed" could also be done in nginx, but most would cost too much.
A separate logging thread that is queued to helps a lot and was one of the main reasons for writing "filed" -- my ability to serve files was being slowed by my ability to write logs indicating that I had served something. The downside is that there may be a large queue of unwritten logs in the event of a kernel panic it other unexpected process termination.
Most requests don't even open the file they are serving because "filed" caches open file descriptors -- once the file has been opened it's kept open until cache entry is needed for a newer file.
There are no runtime allocations after startup except for log entries, leading to very consistent performance under loads.
e.g. node.js has no problem getting 20k/sec per core, but a stall at the wrong time kills every pipelined HTTP request that follows (until you tear down the connection and restart it).
Reports of "millions per sec" are usually talking about messages on established channels across all CPUs.
If you're actually aware of a java based web server that can beat even 150k http requests measured by wrk or similar on local host I'd like to see it.
I think it can be micro-optimized beyond that. With predictions to avoid unnecessary syscalls, with syscalls grouped together to make cpu more efficient for the rest of the time it spends in event loop, and if it's possible to modify kernel a bit - with batching syscalls together to make them very cheap.
That is to say, I suspect that if micro-optimisations can double our performance, they will be more complicated than just writing a customised ring0 that implements HTTP directly inside the network driver.
Here is how I'm looking at it:
• 10Gb/sec network port
• 4k max requests and responses
• == 1.3 million HTTP requests per second.
Now the problem is that main memory is not much faster than our fastest network: About 15Gb/sec, so what we're talking about here is code and state staying entirely in L1, and streaming the network buffers across the CPU, and responding in one pass, to get that 1.3 million optimal performance.
My dash server gets ~135k HTTP requests per second on localhost (I should be able to approach 300k/sec over a network if I ever get around to it). That's 22% of our optimal performance, and a lot better than any other HTTP server I'm aware of.
At this speed, one of those micro-optimisations `writev()` is actually slower than `write()` -- likely because the code path is shorter in the simpler codebase -- but it illustrates my concern nicely: That we are close to that break-even point with the optimisations we can make. If we make our server bigger and more complicated, it might not make our programs any faster.
That suggests to me that the solution is actually fewer, simpler syscalls, not more, bigger ones.
All of this is motivation to rewrite everything in C.
I took part in a few projects that replaced high throughput servers handling mobile network traffic from C++ to Java.
Not that they aren't true, rather their validaty depends a lot of programmer skillset and compilers being used.
I work at two ad-tech related companies, one large one medium, both use Java from the beginning. And I know other company using Python/Golang as well.
Didn't know C(not C++) is particular popular until today. Note that ad company are pretty business focused, add or remove features for big clients are pretty common, so development efficiency matters a lot.
> The "stop the world" phase of the collector will almost always be under 10 milliseconds and usually much less.
GC pauses are a killer. They can be worked around, but it takes intensive tuning.
The RTB bidder I wrote many moons ago at a startup was fast as hell, but had problems in the 95th percentile of requests meeting the latency targets, due to GC.
The ad servers at the big ad tech heavy weights are in C++.
Also note I wasn't doing this alone, it was a very big project a mobile operator.
> Limitations: > - only tested on Linux
Unlikely, and given his attitude I'm not going to waste my time trying nxweb.
Yeah, nope. Check out H2O if you haven't already https://h2o.examp1e.net/.
It looks significantly more flexible than anything nginx offers without having to bolt on a server-side language like PHP - unless nginx has something similar in its millions of modules that I'm not aware of.
(I know about and love nginx SSIs, but the templating here looks more flexible than them.)
1) Why you think that java is slower than C++? Server-side JIT compiles much more optimized code as it is really know what and how to optimize.
2) What about security? Almost half of the problems in security in last days came from native code stuff.
While this seems perfectly plausible, would you happen to know some benchmark backing this claim? Thanks.
citation needed
> native code stuff
That is inaccurate. But if you said that it comes from C memory management issues, that sounds more plausible. We could talk about Rust, but I even think that C++ is usually written in a much safer style than C.