35M Hot Dogs: Benchmarking Caddy vs. Nginx
blog.tjll.net
blog.tjll.net
- The sendfile tests at the end actually didn't use sendfile, so expect much greater performance there.
- All the Caddy tests had metrics enabled, which are known[2] to be quite slow currently. Nginx does not emit metrics in its configuration, so in that sense the tests are a bit uneven. From my own tests, when I remove metrics code, Caddy is 10-20% faster. (We're working on addressing that [3].)
- The tests in this article did not tune reverse proxy buffers, which are 4KB by default. I was able to see moderate performance improvements (depending on the size of payload) by reducing the buffer size to 1KB and 2KB.
I want to thank Tyler for his considerate and careful approach, and for all the effort put into this!
[0]: https://caddy.community/t/seeking-performance-suggestions-fo...
[1]: https://twitter.com/mholt6/status/1570442275339239424 (thread)
But disabling metrics is not supported in standard Caddy, you need to remove specirc code and recompile to disable it.
So maybe benchmarking with it isn't fair to Nginx?
We're addressing that quite soon. Unfortunately the original contributor of the feature has been too busy lately to work on it, so we might just have to go the simple route and make it opt-in instead. Expect to see a way to toggle metrics soon!
Update: PR opened: https://github.com/caddyserver/caddy/pull/5042
* How do these things perform by default. This is how they're going to perform for many users, because if it's adequate nobody will tune them, why bother.
* How do these things perform with performance configuration as often recommended online. This is how they'll perform for people who think they need performance but don't tune or don't know how to tune, this might be worse than default but that's actually useful information.
* How do these things perform when their authors get to tune them for our test workload. This is how they'll perform for users who squeeze every drop and can afford to get somebody to do real work to facilitate, possibly even hiring the same authors to do it.
In some cases I would also really want to see:
* How do these things perform with recommended security. A benchmark mode with great scores but lousy security can promote a race to the bottom where everybody ships insecure garbage by default then has a mode which is never measured and has lousy performance yet is mandatory if you don't think Hunter2 is a great password.
Agreed on this one -- today I'm looking at how to disable metrics by default and make them opt-in. At least until the performance regression can be addressed.
Update: PR opened: https://github.com/caddyserver/caddy/pull/5042 - hoping to land that before 2.6.
Phew. Caches are purged, my errors are fixed. I can rest finally. If folks have questions about anything, I'm happy to answer.
[1]: https://blog.tjll.net/reverse-proxy-hot-dog-eating-contest-c...
I can die happy
One of the great mysteries in (my) life is why people think that measuring things is free. It always slows things down a bit and the more precisely you try to measure speed, the slower things go.
I just finished reducing the telemetry overhead for our app by a bit more than half, by cleaning up data handling. Now it's ~5% of response time instead of 10%. I could probably halve that again if I could sort out some stupidity in the configuration logic, but that still leaves around 2-3% for intrinsic complexity instead of accidental.
The HAProxy configuration is just as simple as Caddy for a reverse proxy setup. It's also written in C which is a comparison the author makes between nginx and Caddy. And it seems to be available on every *nix OS.
Does HAProxy have built in support for Let’s Encrypt?
That is one of my favorite features. Caddy just automatically manages the certificates for https.
sub.domain.com {
# transparent proxy + websocket support + letsencrypt TLS
reverse_proxy 127.0.0.1:2345
}
It's a fresh breath of air to have server with sensible defaults after dealing with apache and nginx (haproxy isn't much better in that regard). caddy reverse-proxy --from sub.domain.com --to :2345
Glad you like using Caddy!Where and how parameters are configured is a bit more of a wild card and dependent on the environment you are running in.
See this issue and especially the comment from Lukas Tribus: https://github.com/haproxy/haproxy/issues/1864
Disclosure: Community contributor to HAProxy, I help maintain HAProxy's issue tracker.
> Disks accesses can be unpredictably slow and would block the entire thread which is not something you want when handling hundreds of thousands of requests per second.
This is not something I see mentioned in the issue, but I don't see why disk accesses need to block requests, or why they have to occur in the same thread as requests?
> I wonder if a disk-writing process could be spun out before dropping privileges?
I mean … it sure can and that appears the plan based on the last comment in that issue. However the “no disk access” policy is also useful for security. HAProxy can chroot itself to an empty directory to reduce the blast radius and that is done in the default configuration on at least Debian.
> but I don't see why disk accesses need to block requests
My understanding is that historically Linux disk IO was inherently blocking. A non-blocking interface (io_uring) only became available fairly recently: https://stackoverflow.com/a/57451551/782822. And even then it's a operating system specific interface. For the BSD's you need a different solution.
If your process is blocked for even one millisecond when handling two million of requests per second (https://www.haproxy.com/de/blog/haproxy-forwards-over-2-mill...) then you drop 2k requests or increase latency.
> or why they have to occur in the same thread as requests?
“have” is a strong word, of course nothing “has” to be. One thing to keep in mind is that HAProxy is 20 years old and apart from possibly doing Let's Encrypt there was no real need for it to have disk access. HAProxy is a reverse proxy / load balancer, not a web server.
Inter-thread communication comes with its own set of challenges and building something reliable for a narrow use case is not necessarily worth it, because you likely need to sacrifice something else.
As an example at scale you can't even let your operating system schedule out one of the worker threads to schedule in the “disk writer” thread, because that will effectively result in a reduced processing capacity for some fractions of a second which will result in dropped requests or increased latency. This becomes even worse if the worker holds an important lock.
Go's TLS stack is set to be more efficient and safer in coming versions thanks to continued work by Filippo and team.
In fact I use HAProxy in production pretty regularly because it is solid but its config one of the main reasons I would choose something else.
A basic HAProxy config is fine but it feels like after a little bit each line is just a collection of tokens in a random order that I have to sit and think about to parse.
- Do you want to add a header in this "location" block? Great, you better remember to re-apply all the security headers you've defined at a higher level (server block for instance) because of course adding a new header will reset those.
- Oh, you mixed prefix locations with exact location with regex locations. Great, let's see if you can figure out by which location block will a request end up being processed. The docs "clearly" explain what the priority rules for those are and they're easy to grasp [1].
- I see you used a hostname in a proxy_pass directive (e.g.: http://internal.thing.com). Great, I will resolve it at startup and never check again, because this is the most sensible thing to do of course.
- Oh... now you used a variable (e.g.: http://$internal_host). That fundamentally changes things (how?) so I'll respect the DNS's TTL now. Except you'll have to set up a DNS resolver in my config because I refuse to use the system's normal resolver because reasons.
- Here's an `if` directive for the configuration. It sounds extremely useful, doesn't it? Well.. "if is evil" [2] and you should NOT use it. There be dragons, you've been warned.
I could go on... but I think I've proved my point already. Note that these are not complaints, it's just me pointing out that nginx's configuration has its _very_ significant warts too.
[1] https://nginx.org/en/docs/http/ngx_http_core_module.html#loc...
[2] https://www.nginx.com/resources/wiki/start/topics/depth/ifis...
Unfortunately it's likely I worked in the same building as one of the people responsible for either creating or at least maintaining that mess, but I didn't know at the time that he needed an intervention.
Other than Caddy, Caddy has been great so far but I have only used it for personal projects.
Disclosure: Community contributor to HAProxy, I help maintain HAProxy's issue tracker.
I'm currently on a push to improve our docs, especially for beginners, so feel free to review the changes and leave your feedback: https://github.com/caddyserver/website/pull/263
I last setup a Caddy config maybe 6- 9 months ago, and everything related to client certificates was either scantily documented, wrongly documented, or not documented at all. It might be I got unlucky, as some of the client cert features were fairly new, but it wasn't a great experience.
Still, I much prefer Caddy's config system to Nginx or HAProxy.
Oh, something else I'd really love to see are more fully-featured example configs, as it can be hard to know how to start sometimes.
Generally we encourage examples in our community wiki though: https://caddy.community/c/wiki/13 -- much easier to maintain that way.
The problem with "fully-featured" examples is that people copy and paste instead of learn how the software works. I'd rather our user base be skilled crafting configuration.
Generally we recommend that big examples go into our wiki: https://caddy.community/c/wiki/13
I'm keeping a watching brief on https://github.com/darkweak/souin and its Caddy integration to see if that can step up and replace Varnish for short-lived dynamic caching of web applications. Though I've lost track of its current status.
haproxy started out as a proxy and has gained some web server abilities, but is all about proxying.
haproxy has less surprises as a reverse proxy than nginx does. Some of the defaults for nginx are appropriate for web serving, but not proxying.
I don't think performance is ever going to matter for my use case, but one thing I think is worth highlighting is the quality of the community and maintainership. In a thread I started asking for feedback on my Caddyfile (https://caddy.community/t/suggestions-for-simplifying-my-cad...), mholt determined I'd found a bug and rapidly fixed it. I followed up with a PR (https://github.com/caddyserver/website/pull/264) for the docs to clarify something related to this bug which was reviewed and merged within 30 minutes.
I'm still thinking about that `./`-pattern-matching problem. Will probably have to be addressed after 2.6...
I get wanting to isolate things but this is the problem with micro benchmarks, it doesn't test "real world" usage patterns. Chances are your real production server will be logging to at least syslog so logging performance is worth looking into.
If one of them can write logs with 500 microseconds added to reach request but the other takes 5 milliseconds that could be a huge difference in the end.
If one of them can do 100,000 requests per second but the other can do 80,000 requests per second but you're both capped at 30,000 requests per second because of system level limitations then you could make a strong case that both products perform equally in the end.
This cannot be understated. Caddy is not written in C! And it can even run your NGINX configs. :) https://github.com/caddyserver/nginx-adapter
Given available options, I will take the network software written in a memory safe language every time.
I'm building a service manager à la systemd in Go as a side project, and I really like it - it's not as low level as Rust and has a huge runtime but it is impressively fast.
There is enough juice in compiled managed languages that expose value types and low level features, it is a matter to learn how to use the tools on the toolbox instead of taking always the hammer out.
That's just one example of how OP's "optimized" nginx config is barely even optimized. There are lots of other variables that you can tweak to get even better performance and blow Caddy out the window, but those tweaks are going to depend on the specific workload you expect to handle. There isn't a single, perfectly optimized, set of values that's going to work for everyone.
The beauty of Caddy is that you get most of that performance without having to tweak anything.
I got those results by seriously limiting the junk in http headers. Not with real browsers.
If you have that demand for any commercial service, you have money to distribute your load globally across more than one nginx instance.
I know Nginx doesn't use keepalives to backends by default (and I see it wasn't setup in the optimised Nginx proxy config), but it looks like Caddy does have keepalives enabled by default.
Perhaps that could explain the delta in failure rates, at least for one case?
Keepalives can actually reduce the performance of a server with many concurrent clients (i.e. a benchmark test), and have other weird effects on benchmarks: https://www.nginx.com/blog/http-keepalives-and-web-performan...
A tool that can send a request at a constant rate i.e. wrk2 or Vegeta [2] is a much better fit for this type of a performance test.
1. https://www.scylladb.com/2021/04/22/on-coordinated-omission/
See more discussion here[2].
[1]: https://k6.io/docs/using-k6/scenarios/arrival-rate/
[2]: https://community.k6.io/t/is-k6-safe-from-the-coordinated-om...
Yes it can, via the `constant-arrival-rate` executor[1].
> the overhead of establishing a new TCP connection for a single HTTP request will dominate the benchmark
By default, k6 will reuse TCP connections, and you have to explicitly disable it[2].
I'm not saying that wrk2 or Vegeta wouldn't be a good fit for this test, but k6 is also capable of it, with some minor configuration changes.
[1]: https://k6.io/docs/using-k6/scenarios/executors/constant-arr...
[2]: https://k6.io/docs/using-k6/k6-options/reference#no-connecti...
This is the new gold standard for benchmarks!
OP / Author, stupendously tremendous job. The methodology is defensible and sensible. Thank you for doing this on behalf of the community.
I am also in love with the friendliness and tone of the article. I’m a complete dummy when it comes to stuff like this and still understood most of it. Feynman would be proud.
I only ever use Caddy as a reverse proxy for web apps (think Flask, Ruby on Rails, Phoenix Framework). My projects have never needed high performance, but if my projects ever take off, it's nice to see that Caddy is already competitive with Nginx on resilience, latency, and throughput.
This is a good write up. I was expecting Caddy to trounce Nginx, but that wasn't the case. I'll be back to re-read this with fresh eyes tomorrow.
[1] For the avoidance of doubt, this is not meant as a snarky observation.
But Caddy certainly does in some cases, especially with the upcoming 2.6 release.
I absolutely was, yes. As an observer I see a lot of people saying positive things about Caddy around here, and how it’s superior performance-wise to a variety of ‘classic’ httpd software. Lots of people love Caddy, and they’re quite vocal, so it’s not a stretch to assume there are reasons why they love it. Nginx development has slowed since the events in Ukraine, unsurprisingly, so again it’s not a leap to surmise Caddy is making good things happen in the meantime.
Caddy scales better than NGINX especially with regards to TLS/HTTPS. Our certificate automation code is the best in the industry, and works nicely in clusters to coordinate and share, automatically.
Caddy performs better in terms of security overall. Go has stronger memory safety guarantees than C, so your server is basically impervious to a whole class of vulnerabilities.
And if you consider failure modes, there are pros and cons to each, but it can definitely be argued that Caddy dropping fewer requests than nginx (if any at all!) is "superior performance".
I'm actually quite pleased that Caddy can now, in general, perform competitively with nginx, and hopefully most people can stop worrying about that.
And if you operate at Cloudflare-"nginx-is-now-too-slow-for-us"-scale, let's talk. (I have some ideas.)
I had the case when after push notifications mobile clients wakeup and all of them doing TLS handshake to LoadBalancers (Nginx), hitting cpu limit for minute or so, but otherwise had no problem with 5-15k rps and scaling.
So we find lots of people using Caddy to serve tens to hundreds of thousands of sites with different domain names because Caddy can automate those certificates without falling over. (Huge deployments like this will require a little more config and planning, but nothing a wiki article [0] can't help with. You might also want sufficient hardware to keep a lot of certs in memory, etc.)
Also note that rps is not a useful metric when TLS enters the picture, as it says nothing useful about the actual TLS impact (TLS connection does not necessarily correlate to HTTP request - and there are many modes for TLS connections that vary).
[0]: https://caddy.community/t/serving-tens-of-thousands-of-domai...
Okay, interesting. It seems their operation mode is quite different from what I used for/see around.
I wonder how they do it for active / passive LB setup, internal services (not accessible over internet for http challenge and so on) , probably it's not their case though.
Not saying it's not useful, just so minor part of the other things for my operations burden.
It is, actually!
Caddy automatically coordinates with other instances in its cluster, which means simply sharing the same storage (file system, DB, etc.) -- so it works great behind LB. Caddy's reverse proxy also offers powerful load balancing capabilities similar to and, in some ways, superior to, what you find in HAProxy, nginx, etc. Caddy uses the TLS-ALPN challenge and HTTP challenge by default, automatically fails over to another when one doesn't work, and even learns which one is more successful and prefers that over time.
Caddy can also get certificates for internal use, both from public CAs using the DNS challenge, or from its own self-managed CA which is also totally automated.
It turns out that these abilities save some companies tens of thousands of dollars per year!
Guess I'm waiting for Cloudflare to FLOSS-release their proxy https://news.ycombinator.com/item?id=32864119 :)
To me it seem that Caddy suffer from BufferBloat. Under heavy congestion the goodput (useful throughput) will drop to 0 because client will start timing-out before the server get a chance to respond.
Caddy should use an algorithm similar to : https://github.com/Netflix/concurrency-limits
Basically check what was the best request latency, and decrease concurrency limit until latency stop improving.
I'd probably learn toward CUBIC: https://en.wikipedia.org/wiki/CUBIC_TCP
(I implemented Reno in college, but times have changed)
The nice thing about Caddy's failure mode is that the server won't give up; the server has no control over if or when a client will time out, so it I never felt it made much sense to optimize for that.
This is why in ip switches they drop packet when the outbound queue is full. And letting the caller retry sending the packet.
When there is no congestion it’s ok for the server to not timeout at all and wait for the queue to drain and let the client decide when to close the tcp connection.
but when there is congestion you don’t want the queue inside the server to get too long. Because the server ( including caddy) might be behind a tcp load-balancer so it’s better for the client to queue it’s request inside another Caddy instance that is less busy.
feel free to reach out to me at maxime.caron@gmail.com
I would be happy to try to add it to Caddy if you are interested
> [...]
> Create two EC2 instances - their default size is c5.xlarge
When you're benchmarking, you want a stable platform between runs. Virtual private servers don't offer that, because the host's resources are shared between multiple guests, in unpredictable ways.
But I would certainly agree that, for the utmost accurate results, a bare-metal situation would probably be more accurate than what I have written.
I would say if you are not testing 10kcc you are not pushing the difference between nginx and apache1.3
As soon as you do push 10kcc, kernel tcp buffers and the amount of junk in your browser http headers start to be more import than server perf. Just in the amount of data coming into the nic.
nginx awful use and only make easy accidentally shoot one’s feet’s.