Thread Pools in Nginx Boost Performance 9x (2015)
nginx.com
nginx.com
"The page cache works pretty well and allows NGINX to demonstrate great performance in almost all common use cases...So if you have a reasonable amount of RAM and your working data set isn’t very big, then NGINX already works in the most optimal way without using thread pools"
So using thread pools in NGINX isn't a general recommendation. The article suggests "a heavily loaded NGINX‑based streaming media server" as the type of situation where it makes sense.
No matter how much you can do with a single thread, in a highly parallel environment (such as a web server), on a cpu with multiple cores (the norm), you can (almost always) do more with multiple threads.
There is no performance benefit to adding more threads, they would increase CPU state/cache thrashing which will slow things down. The main benefit of using a lot of threads is that more programmers know how to design software that way because it was taught for so many years; it isn't good for performance on modern hardware.
Yes, see here for more discussion on this:
https://github.com/ronomon/crypto-async#adjust-threadpool-si...
Source - spent a year performance tuning a matching engine.
So once you are there, the next step is to busy spin at a Thread Per Core (TPC). You have 10 cores, you find 10 (or more realistically 8-9 to leave the OS some spare cores to muck with) threads, and busy spin them at 100% doing work. You never let the linux scheduler touch them.
If this stuff interests you, there are some cool papers by lmax and the disruptor and gil tene that talk about this stuff..
The amount of work you can do with a single 10-20 core CPU is amazing.
If these are "worker threads", new workers are scheduled when all other workers are blocked and work is available. This guarantees that idle CPU will always be scheduled with work. When workers unblock, they continue execution on their work and start picking up new work. The worker scheduler will stop scheduling new workers and the extra workers will go back to the pool of inactive workers. This is the idea behind a worker or thread pool, they get scheduled when its possible to do more work.
A poor implementation of this idea will just hand work off to threads and then expect them to all execute in parallel, which will lead to negative performance impact.
The problem with (some implementations of) async workers is that just scheduling more will result in the latter problem, work split across too many workers. This is especially true for workers implemented as entire processes. The problem this post brings up is that 1 to 1 mapping of workers to cores suffers from idle CPU due to unanticipated blocking. The general solution is to smartly distribute work over a pool of workers, and not blindly schedule a fixed number workers.
So the whole webserver serving clients might occasionally block waiting for disk if you're serving online videos or some other application where your whole website is too big to fit in the OS cache. But if you have multiple threads, then if one of them blocks, it isn't a big deal as the other threads will continue to unqueue requests.
> The asynchronous interface requires the O_DIRECT flag to be set on the file descriptor, which means that any access to the file will bypass the cache in memory and increase load on the hard disks
Worker does use process pools with their own thread pools, but they're still not evented/async.
Apache 2.4 (2012) saw the release of 'event', an async variant of worker, which is like the new nginx model.
"aio_write" -> wasn't the idea of AIO operations to actually be truly async? Maybe it's because it requires XFS. Anyway, I'm confused.
"Although Linux provides a kind of asynchronous interface for reading files, it has a couple of significant drawbacks. One of them is alignment requirements for file access and buffers, but NGINX handles that well. But the second problem is worse. The asynchronous interface requires the O_DIRECT flag to be set on the file descriptor, which means that any access to the file will bypass the cache in memory and increase load on the hard disks. That definitely doesn’t make it optimal for many cases."
"The asynchronous interface requires the O_DIRECT flag to be set on the file descriptor, which means that any access to the file will bypass the cache in memory and increase load on the hard disks"
They even go into how no interface has been surfaced yet that allows determining whether a file is in cache or not. Also how FreeBSD's aio interface doesn't have the same limitations
The headline uses a the word "performance" which is rather ambiguous.
Performance has many aspects: Latency, through-put, resource-usage (doing more with less), etc etc.
A more specific headline wouldn't hurt.
In the storage engine I'm developing/maintaining, I've been using fibers, for both network and disk IO, until people started noticing that, if number of the client grows, and if they ask a lot of data (generating a lot of disk traffic), the performance drops severely. I've moved to the thread pools, dispatching my disk IO operations there (although it's much problematic for the reads, as writes would be performed by kernel anyway, `fsync` and `close` would be operations that you still want to do in the separate thread): https://github.com/sociomantic-tsunami/dlsnode/blob/master/s...
edit: used right link for the freshly opensourced repo.
i guess you are referring to aio(7) ? posix aio on linux is implemented in glibc, and doesn't scale in presence of multiple threads. which is probably what is alluded to...
not quite :) look at aio(7)
There isn't really a "comparison" here. Co-routines/fibers could behave exactly like nginx does for IO or not at all. It all depends on the implementation.
Many HN readers are apprehensive of the use of D because it is garbage collected and conjecture that it is not appropriate for low latency and/or high throughput work loads. Correct me if I am wrong, but that is exactly the kind of load the sociomantic needs to address. I have thoroughly enjoyed the relevant Dconf videos but it would be great to hear some first hand account.
BTW is the move to D2 done ?
Yes, what you describe is exactly the kind of load we're addressing, and we're not getting GC to be in the way, simply by making usage of the reusable memory buffers, which are allocated in the first few requests, and then always reuse/recycled & fetched from the pool, so that we're not giving a chance for GC to kick in (as explained in the great blog post series, in D, GC mark&sweep will kick in only on allocations, completely deterministic: https://dlang.org/blog/category/gc/). Ocean provides some help here: https://github.com/sociomantic-tsunami/ocean/blob/v2.x.x/src... https://github.com/sociomantic-tsunami/ocean/blob/v2.x.x/src..., etc. This makes D completely suitable for these kinds of applications, and from my experience, none of D's features or the misfeatures is making such implementations hard.
D2 move is not yet completely done. We're still writing code that's compatible in both D1 and D2. For example, DHT node project (https://github.com/sociomantic-tsunami/dhtnode) and libraries on which it's depending on such as ocean (https://github.com/sociomantic-tsunami/ocean) and swarm (https://github.com/sociomantic-tsunami/swarm) are able to run in D2 with no performance penalties, but they still don't use D2-only constructs.
What they talk about in the article is that they offload potentially blocking operations to thread pools. That's not about networking, which is already done asynchronously, but about e.g. reading from a file not present in the page cache. Note that detecting that the needed bytes are already in the page cache and thus the read wouldn't block and should not be offloaded is tricky. As far as I can see they haven't done that yet so this new feature (edit: not so new, the post is from 2015) is not without downsides.
Having a always-ready state on regular files is a problem since Linux's non-blocking is going around readiness. On Windows (and I believe on Solaris/FreeBSD), completeness model is in place, so you schedule an operation, kernel does _everything_ and then you're resumed upon completion of IO.
Is it possible for me to use lighttpd as a proxy but at the same time serve static pages for a subset of urls?
You should ask here: https://redmine.lighttpd.net/projects/lighttpd/boards