Show HN: Nginx Sysguard – throttle requests by load, average request time, etc.
github.com
github.com
I'm trying to compare it to alternative solution which prevents the high load in the first place (puts limit on the CPU/mem shares). By limiting the app rather than nginx itself, you can still set the error page for the case where the backend is too busy (handle 502 bad gateway).
The nice part of doing that at system level is that you can prioritise other processes over nginx, but otherwise ensure the app can take 100% CPU if nothing else is interested in using it.
Of note, the module is a port from the nginx tuned by alibaba which is quite a credible source.
I have served over a billion pageviews/month, and cutting off requests at a certain load threshold with a custom error page to users is a good idea. Otherwise the users will just get a huge slowdown or a backlog of requests which will cause a prolonged server recovery.
Also as the author vozlt mentioned, you can use "average of request processing times" and custom http errors. Also, it is much easier to use than the methods you suggest.
It's the "throttle" step here which is important. When overloaded, drop requests; otherwise your server will overload and crash, which can be difficult to recover from.
I don't know that this module is perfect for every situation, but I do know that timeouts don't work great when there's a general overload. Timeouts leave a higher number of concurrent connections there.
I'd prefer a sort of priority queue setup where http clients that already "got in", "stay in", and newer connections are pushed away. Rate limiting per client first might be better as well.
I take your point that this module isn't a panacea. For several things I've seen neither your approach or this approach would work well. An api gateway that throttled per client would make more sense. This sysguard module, though, came from Alibaba...I imagine they are pretty sharp and not prone to making things they don't need.
It also appears to be able to kick off any action, not just 503. So you could have it conditionally enact limit_req on heavyweight requests. That seems to match your desired approach.