Caddyhttp: Enable HTTP/3 by Default
github.com
github.com
Thanks to Marten Seemann for maintaining the quic-go library we use. (I still haven't heard whether Go will add HTTP/3 to the standard library.)
Caddy 2.6 should be the first stable release of a general-purpose server to support and enable standardized HTTP/3 by default. HTTP versions can be toggled on or off. (Meaning you can serve only HTTP/3 exclusively if you're hard-core.)
PS. Caddy 2.6 will be our biggest release since 2.0. My draft release notes are about 23 KB. We're looking at huge performance improvements and powerful new features like events, virtual file systems, HTTP 103 Early Hints, and a lot of other enhancements I'm excited to show off on behalf of our collaborators!
I have recently gave Caddy a shot, and immediately loved it!.
I am a bit worried about the performance deficit compared to apache or Nginx, but none of my servers really have reached the scale where it actually matters.
Its unfortunate that there is no dynamic brotli support, but oh well, a trade off I'm willing to make.
One of things that aren't clear to me are Caddy plugins. How big of an affect do they have on peformance? Are plugins a core part of Caddy's usage? Or are they just exist to cover some edge cases?
I'm thinking of using the caddy-filter and caddy-crowdsec-bouncer plugins.
Caddy's performance is competitive with nginx, but there are too many dimensions to discuss here. I guarantee that for most people, your web server will not be your bottleneck.
Dynamic brotli is expensive, as brotli compression is inefficient. There are no optimized Go implementations of this yet. But Caddy can serve precompressed brotli files without any separate plugins.
What do you mean by "what affect do [plugins] have on performance"? All features in Caddy are plugins, some just come installed by default: https://caddyserver.com/docs/architecture -- even its HTTP server is a plugin. (All plugins are compiled in natively.)
Oh I didn't know that, thanks for the awesome project!
You have to do your own benchmarking to see what effect they have to performance. Everyone has different needs and configs, so it's not possible to give authoritative benchmarks. Just try it out and see.
Regarding brotli, there is a plugin you can try if you'd like https://github.com/ueffel/caddy-brotli, but brotli's on-the-fly performance is not that good. It works well if you pre-compress you assets though, and you can serve those with the file_server directive's precompressed option.
When I first announced Caddy in 2015 and the server got busy, it started downloading everything as .gz files. (facepalm) (good ol' days)
But that's just a fun story, not actually the reason we don't enable it by default...
The tricky thing about enabling features by default is that turning them off is awkward in config. And while we like to have "magic" in Caddy, we don't like too much magic. Plus, gzip performance in Go is less good (but more memory-safe) than it is in C and assembly. Klauspost's flate implementation is very fast and that's the one we use. But even with a super-fast implementation, gzip requires memory and CPU that busy servers may be in short supply of.
So it's opt-in for now. We can always change it later; going the other way is harder.
(Caddy also supports serving pre-compressed files, if enabled! And starting with Caddy 2.6, those can be teleported at the speed of electricity to HTTP clients using sendfile. HTTPS also sees faster file serving in 2.6 due to optimized copying.)
See also:
https://serverfault.com/questions/296770/why-arent-features-...
But it's obviously a trade-off and there's lots to weigh up...
Not high priority, but if Caddy would offer a report of things that aren't configured but perhaps should be...that could help. Things like Cache-control/Expires/Etag headers, Gzip/Deflate, Caching, and so on.
- An application uses compression
- An attacker is able to supply chosen data to it
- The application compresses the attacker's data and static secret data together
- The attacker is able to monitor the size of the compressed data
- This can be repeated by the attacker a number of times
will be vulnerable to having its secret data stolen by techniques like BREACH. If you want your secret data to stay secret, don't compress it with attacker chosen plaintext where the resulting size could be monitored.
But also, please don't take benchmarks like this as gospel. Run your own benchmarks, on your own hardware, with your own config. Someone else's config may not be ideal for your usecase.
greenpau/caddy-security is fantastic and "just works" for OIDC sso.
mholt, thanks for recently adding the ability to bind to multiple specific IP addresses by default, this help me conserve precious public IPv4 addresses.
*http3/quic gets around this be setting a connection id in the UDP packets so mobile phones jumping towers can still be tracked if they get a new IP.
It's called BGP protocol, as OP mentioned. This is basically how cloudflare works, you have the same IP for dozens of datacenters around the world. If routing changes and suddenly you are routed to different POP then your TCP session is broken but in most cases this is rare. You could "fix" this by routing TCP packets to POP B when detecting on POP A that session origin was POP B but this would incur additional latency for the rest of the session etc. But this has nothing to do with Caddy, you could use any other software same way using anycast IP.
I think I wasn't clear enough. Yes, BGP will route the packet to the nearest POP for that IP if it was announced in different locations, that's how "anycast IP" works, thanks to how BGP works. OP asked:
> or something in front of Caddy
So I've described what is it and how it works in simple terms.
> If routing shifts in the middle of a TCP connection, there's no magic that makes a new machine that didn't SYN+ACK the original connection pick up where the old machine left off.
If you are talking only about BGP then of course you are right but I was not talking about solving this using only BGP. You can tunnel the packet to your other POP which was origin, as I mentioned in my answer.
So if you have a session established after 3 way handshake in POP B
client <-> POP B
and suddenly you are routed via POP A then just tunel all packets for that session:
client <-> POP A <-> POP B
or you could use MPTCP to solve this if supported etc.
Obviously, in reality, people mostly rely on geography to deal with this: you're just unlikely to get randomly rerouted from a Singapore rack to a Sydney rack, and so the issue doesn't come up. But if you care about it --- as the previous commenter does --- you want a real answer.
Yeah, I wasn't clear enough in my first comment, my last comment was more detailed so I guess it should match "a real answer" now.
> Obviously, in reality, people mostly rely on geography to deal with this: you're just unlikely to get randomly rerouted from a Singapore rack to a Sydney rack, and so the issue doesn't come up.
Yeah, most of the time it doesn't come up, anecdotally it's also more specific to some regions, in Europe for example I've seen this to occur more often (between POPs in EU) than in other regions.
1: EdgeRouter4 connected to ISP. This operates as a typical router, nothing special other than some static routes to item 2.
2: Pair of VyOS routers running in VMs. I think of them as top of rack L3 switches in a datacenter. Their primary purpose is for ECMP [1]. They also use VRRP so the can be upgraded without downtime.
3: Three node caddy cluster I mentioned. Running plain old Debian. anycast-healthchecker configures the anycast IP on each VM's loopback interface. Bird advertises the address to the two VyOS routers via BGP.
4: Caddy binds to each address.
All of the hosts in this network are configured with a default route to the VRRP address on the VyOS VM's. That's pretty much it.
I've recently started using Calico in k8s on some separate VM's, and they work the exact same way, advertising Services of type LoadBalancer to the VyOS routers which does ECMP across the nodes. I learned quite a bit configuring anycast-healthchecker with Caddy, then comparing it to how Calico works.
[1]: https://codecave.cc/multipath-routing-in-linux-part-2.html
At first I was scared of how stupid simple it is. It feels like web servers are supposed to have giant config files with a hundred mysterious knobs to twiddle. Now I always default to Caddy, and have yet to find an instance where it didn't fit my needs. Congrats.
Curiously, new versions of Apache include mod_md, which lets you use Let's Encrypt for SSL, without the need for something like certbot: https://httpd.apache.org/docs/2.4/mod/mod_md.html
I actually wrote about it in a blog post of mine, "How and why to use Apache httpd in 2022": https://blog.kronis.dev/tutorials/how-and-why-to-use-apache-...
That said, while the capability is there, things like the DNS-01 validation need additional work and Caddy comes with more sane defaults out of the box and arguably a way easier config format.
It would also be really cool if Nginx came with its own implementation of ACME certificate provisioning out of the box, following the trend of other servers, but certbot is also perfectly passable.
Either way, it seems like most of the "popular" web servers are viable for most use cases nowadays and I can say that Caddy definitely belongs in that list!
For any /foo/ /bar/ /baz/ I want to redirect to /foo, /bar, or /baz. Unless I'm missing something, this needs to be done like so right now:
@trailing_slash path_regexp trailing_slash ^/(.*)/$
redir @trailing_slash /{re.trailing_slash.1} 308It can also do regex replacements on the path portion of the URI. And yes, these are internal rewrites since uri is a directive that wires up the rewrite handler.
We actually have a special section in the docs all about enforcing trailing slashes, including external redirects: https://caddyserver.com/docs/caddyfile/patterns#trailing-sla...
Note that the file_server will automatically enforce canonical URIs, including redirecting to add or remove the trailing slash according to whether the resource is a file or a directory.
To redirect many unspecified paths, your regex is probably the best way to do it for now. Feel free to open an issue to propose an alternative!
I'll follow up with issues tomorrow (I think an issue for updating the common caddyfile patterns bit on the website also makes sense - I suspect a blanket rule for removing trailing slashes is more likely to be a common pattern people seek a solution for than removing a trailing slash on explicitly enumerated paths).
Of course if you're using absolute paths in your HTML then this isn't a problem, but it does become a problem if you want to move to serving this same site from a subpath, because then your absolute paths won't be constrained to this subpath.
I see several posts discussing under which circumstances or use cases one might outperform the other but they never seem to care about having decent metrics.
I might overrate the importance of this, who knows...
The last frontier for these types of comparisons is unusual circumstances like embedded platforms with limited resources.
I'd say that it's possible, just not always easy, due to how many different configurations there can be for any given workload - the same servers might be used to achieve the approximately same configuration in different ways.
However, if you took a common enough use case or workload, such as reverse proxy with SSL and gzip compression for $FOO resource types, serving files for $BAR application and proxying $BAZ API, then it should definitely be possible. You would just need some meaningful real-world test instead of a purely synthetic benchmark, otherwise you'll probably test something slightly different than what your server will be doing in practice.
The other option is just to settle on some bit of common enough functionality and try to increase your sample size as much as possible.
For example, some attempts have been made by OpenBenchmarking and at least give a vague idea of the performance of some servers:
Nginx: https://openbenchmarking.org/test/pts/nginx
Apache:https://openbenchmarking.org/test/pts/apache
I'm mentioning this because while not everybody has the time to reproduce a given setup in any number of technologies, even similar enough tests in a controlled environment can produce meaningful information, at least to let you infer what orders of magnitude you're working with. For something concerning programming languages themselves and their frameworks, one just needs to look at what TechEmpower is trying to do: https://www.techempower.com/benchmarks/#section=data-r21
> Really, you need to do your own benchmarking to determine which solution is best for you. But keep in mind that the web server is rarely the bottleneck, usually your app and database IO are where it takes the most time.
This is well said, though.
In general, I'm inclined to agree: most of the time, the performance of most web servers can be described as "good enough".
The exception to this might be using your application servers (e.g. Tomcat) as a web server and running into situations where serving static assets (that might be baked into your application) would slow down because of API calls being slower and processing them digging into comparatively more conservative HTTP thread limits for the whole thing. Then again, personally I'd argue that you should have one of the popular web servers in front of your applications (and typically serving static assets) to act as an ingress in most cases, but I've seen some interesting things over the years.
That's a perfectly valid use case!
Packaging a front end application (e.g. React/Angular/Vue) inside of a .jar file and making Tomcat or something else serve it, as a part of a larger Java application, though? Perhaps less of an optimal solution than just using Apache/Nginx/Caddy for it, outside of really wanting to be able to deploy everything as a single package.
So, with Caddy or Traefik, a container label can enable HTTP/3 (QUIC (UDP port 1704)) for just that container.
"Labels to Caddyfile conversion" https://github.com/lucaslorentz/caddy-docker-proxy#labels-to...
From https://news.ycombinator.com/item?id=26127879 re: containersec :
> > - [docker-socket-proxy] Creates a HAproxy container that proxies limited access to the [docker] socket
https://github.com/nginx-proxy/nginx-proxy
Note on both projects, if you care about security, you should split the generator/controller container out of the main webserver container, so the docker socket is not used in container that is directly exposed.
So I'm not exactly sure the point you're trying to make. But yes, CDP is an awesome project!
The (unversioned?) docs have: https://caddyserver.com/docs/modules/http#servers/experiment... :
> servers/experimental_http3: Enable experimental HTTP/3 support. Note that HTTP/3 is not a finished standard and has extremely limited client support. This field is not subject to compatibility promises
TIL caddy has Prometheus metrics support (in addition to automatic LetsEncrypt X.509 Cert renewals)
protocols h1 h2
will disable Http/3 but leave 1.1 and 2 on.> Two critical capabilities for security and scalability of web applications and traffic, HTTP3 and QUIC, are coming in the next version we ship
https://www.nginx.com/blog/future-of-nginx-getting-back-to-o...
```
caddy: example.com
caddy.reverse_proxy: {{ upstream 8080 }}
```
Pretty cool stuff.
Disclosure: I'm a community contributor to HAProxy and help maintain the issue tracker.
Changing certificates (https://docs.haproxy.org/2.6/management.html#9.3-add%20ssl%2...) and adding servers (https://docs.haproxy.org/2.6/management.html#9.3-add%20serve...) using the stats socket is both officially documented.
If you find the documentation insufficient or unclear, then this can be considered a documentation bug that I recommend filing here: https://github.com/haproxy/haproxy/issues
Unfortunately there's no best answer for this: Using a single socket will allow for connection migration, but it will end up being a bottleneck in terms of scalability since it will serialize access on a lot of kernel and driver datastructures (just a single transmit/receive queue). Connected sockets avoid that, but don't allow for address migration. And doing external load balancing gets far more complex than just starting a binary - even the most simple solution requires running XDP code.
I think I'm stumbling over terminology here, but UDP is connectionless, right? I think the QUIC protocol re-implements connections on top of it, but I've never heard of "connected UDP".
A socket pool does seem to make sense from a performance standpoint!
Go implementation of HTTP/2 already took a /5 hit over http/1.1 (Go http/2 implementation is 5x slower than Go http/1.1)
With HTTP/3 our early benchs indicate /2 ot /3 from HTTP/2 (so /10 from http/1.1)
Most users do not operate at nearly the volume required to feel the impact, including enterprises.
Would be interested in repeating your experiments and knowing your real world use case and seeing how similar they are.
Yeah I agree this is a edge case, 99,99% of Caddy users don't push so much data ;)
But Caddy can't be used as a big files download server for instance, unless sticking with HTTP/1.1 (that and sendfile + kTLS not been supported last time I checked?) (at least with a single http instance).
Another issue for HTTP/3 is https://github.com/lucas-clemente/quic-go/wiki/UDP-Receive-B... Current Linux in WSL for instance has this issue. This will probably resolves over time as more and more Linux distros will take HTTP/3 requirements into account.
Btw: A good recommendation for anyone doing benchmarking is to both measure throughput for a single connection, but also throughput for the whole server when using lots of connections (e.g. 100 connections per core, and each connection doing multiple concurrent requests). The latter will point out further scaling issues.
In terms of full system efficiency HTTP/2 vs HTTP/1 (with TLS) is actually not that bad for good implementations - might have 10-50% more cost but in the end the performance critical parts (doing syscalls for TCP sockets) are the same for both. QUIC can be much worse. Well optimized setups (which e.g. use optimized UDP APIs with GSO or even XDP) somewhere between 50-100% more expensive than TCP+TLS setups. And simpler implementations can be 5x as expensive, and might not even scale beyond a single CPU core.
- aioquic supports HTTP/3 only now https://github.com/aiortc/aioquic
- httpx is mostly requests-compatible, supports client-side caching, and HTTP/1.1 & HTTP/2, and here's the issue for HTTP/3 support: https://github.com/encode/httpx/issues/275
Or perhaps simply use the number of half-open/embryonic connections as the metric.
Ref: https://github.com/caddyserver/caddy/blob/50748e19c34fc90882...
This attack is the QUIC equivalent of a SYN flood, which results in half-open connections, because the attacker is unable to complete the connection by responding to message the server sends. RequireAddressValidation corresponds enabling syn cookies.
Unfortunately I don't use caddy myself, and don't have time to get involved deeper in improving this.
Does a QUIC connection work without a CA (corporate) root cert being referenced?
QUIC/HTTP3 use the same TLS stack as H1/H2, same things are trusted. So there's no difference there.