How we scaled Nginx
blog.cloudflare.com
blog.cloudflare.com
If you're running Intel DC 3x10 SSDs, check for a firmware update that improves 'maximum latency' in some cases, the update was released some years ago, but people might not have noticed it.
They belong in a desktop/gaming machine not in a production web server.
Here is an old review with enterprise drives and the Samsung 840 pro consumer drive which received excellent reviews at the time it was released, notice how it is the worst performer over time. https://www.storagereview.com/micron_m500dc_enterprise_ssd_r...
The exception is Optane 900p which is blazing fast! A gamechanger!
You can probably figure out the general class of storage you need / can afford, which is interface and bits per cell, but then you probably need to try several well regarded drives (or several dodgy drives, if that's what you can afford) to see how they work for you. Or just go with whatever your hosting provider can supply ;)
Don't get me wrong though, caching/file serving in nginx was always very hacky. And blowing latency by blocking on filesystem and yet relying on vfs cache underneath is just one side effect of that.
“hold memory hostage” really doesn't help make your point. It's inflammatory wording one step above spelling it Micro$oft and it makes it sound like you don't understand the engineering trade-offs which the Microsoft engineers made. It'd be much better if you explained _why_ you disagree with their choice and especially whether there are differences in the type of work which you've used it for which might explain that opinion.
Also in my experience it handles slightly more concurrent connections with a fair bit less CPU usage. There are some weird and not very nice parts + limitations to the API that makes it pretty hard to write cross-platform things with it, granted.
There are ways to do non-blocking disk I/O in *nix (aio/io_submit in linux) but all of which requires you to have an open file descriptor first. Does NT allow you to open a file in an async fashion?
I had not thought about open latency being an issue, that's fascinating. Looking at one of our busy 100G servers w/NVME storage, I see an openat syscall latency of no more than 8ms sampling over a few minute period, with most everything being below 65us. However, the workload is probably different than yours (more longer lived connections, fewer but larger files, fewer opens in general). Eg, we probably don't have nearly the "long tail" issue you do..
That's pretty awesome though, that you have to worry about latency of open().
[1] https://docs.microsoft.com/en-us/windows/desktop/fileio/i-o-...
1. at startup, before threads are spawned, find all static files dirs referenced in the config, and walk them, finding all the files in them, and open handles to all of those files, putting them into a hash-map keyed by path that will then be visible to all spawned threads;
2. in the code for reading a static file, replace the call to open(2) with a look up against the shared file-descriptor from the pool, and then a call to reopen(2) to get a separately seekable userland handle to the same kernel FD (i.e. to make the shared FD into a thread-specific FD, without having to hit the disk or even the VFS logic.)
3. (optionally) add fs-notify logic to discover new files added to the static dirs, and—thread-safely!—open them, and add them to the shared pool.
This assumes there aren't that many static files (say, less than a million.) If there were magnitudes more than that, in-kernel latency of modifying a huge kernel-side FD table might become a problem. At that point, I'd maybe consider simply partitioning the static file set across several Nginx processes on the same machine (similar to partitioned tables living in the same DBMS instance); and then, if even further scaling is needed, distributing those shards on a hash-ring and having a dumb+fast HTTP load-balancer [e.g. HAProxy] hash the requested path and route to those ring-nodes. (But at that point it you're somewhat reinventing what a clustered filesystem like GlusterFS does, so it might make more sense to just make the "TCP load-balancing" part be a regular Nginx LB layer, and then just mount a clustered filesystem to each machine in read-only-indefinite-cache mode. Then you've got a cheap, stateless Nginx layer, and a separate SAN layer for hosting the clustered filesystem, where your SSDs now live.)
No.
As a side note, there was an interesting old hn subthread about async disk i/o of philosophies Windows NT vs Linux:
This is an old new thing post which demonstrates it is easy tot fall off the happy path on nt. Im not sure the windows world is much better off here.
This is a very common way of thinking. But in fact there are only two ways to handle I/O. And no matter what you do, you always end up with one of them:
Path 1, blocking I/O: When you have blocking I/O your process continues to the point where the I/O starts, then sends the corresponding request to the kernel and waits until it gets a respond, potentially forever. This is very low resource usage, but sometimes-hanging-forever is quite a huge price. So usually people put this I/O stuff in a thread/fork and use the parent to have a timeout waiting.
Path 2, non-blocking I/O: In this version when the process hits the I/O it will almost-immediately fail when the desired resource (e.g. file, port, whatever) is not available. So usually you are writing a loop and constantly poll for the resource to become ready. This obviously has a rather high cost on resources, because your code gets more complicated (loops, exceptions, etc) and whatever you are I/Oing to has more activity (e.g. if you poll a webserver you constantly create load on that webserver for each client process). But also an advantage is that you can't hang forever, because you usually break the loop after x seconds or y amounts of retries.
You might feel it sucks (at least I do) but there are not more options. Decide for one version that you can live with more easily, tune the variables you can fiddle with, like timeouts/retries, and then move on to other problems.
https://stackoverflow.com/questions/22780822/linux-kernel-ai...
I've talked to a nginx product manager and he's told me changes that are specific to one customer are unlikely to be accepted.
Also in our implementation we took some shortcuts so it may not be suitable for upstream as is anyway
I think it'd be great if more companies released their own 'opinionated' versions that update with their infrastructure. Like if I wanted to host a OpenStreetMaps tiling server hypothetically using some features cloudflare has in their nginx builds. Makes it easy for white-hats to test, too.
IIRC I've been intersted in HPACK for small http responses where the headers are larger than the body, but if I wanted to use the HPACK patch I have to re-impliment it every time an update comes out that modifies the file.
There are certainly async I/O libraries for Rust (e.g., Tokio), but those are going to be limited by the primitives the OS gives them. (AFAIK, Tokio's core libraries don't directly do async disk I/O; there is tokio-fs, and it does it by shunting the work to a threadpool.) The fact that disk I/O is so uniquely special on Linux effects any language, as it is an aspect of the kernel itself.
I love Rust, and while there are a ton of compelling reasons to use it, I don't think it's fair to say it would have prevented this from happening, in this particular case.
However, development of a new web server in Rust is an interesting project, and could certainly compete with nginx over time if enough effort was put into it.