HTTP/2 zero-day vulnerability results in record-breaking DDoS attacks
blog.cloudflare.com
blog.cloudflare.com
The novel HTTP/2 'Rapid Reset' DDoS attack - https://news.ycombinator.com/item?id=37830987
The largest DDoS attack to date, peaking above 398M rps - https://news.ycombinator.com/item?id=37831062
(Caddy just uses Go's HTTP/2 implementation.)
Why do I say this? Because it breaks nearly every optimization that's been made to serve content efficiently over the last 25 years (sendfile, TSO, kTLS, etc), and requires that the server's CPU touch every byte of data multiple times (rather than never, for http/1). Its basically the "what if I do everything wrong" case in my talk here: https://people.freebsd.org/~gallatin/talks/euro2022.pdf
Given enough time, it may yet get close to HTTP/1. But its still early days.
HTTP/2 and HTTP/3 both have a limit on the number of simultaneous streams (requests) the sender may create. In HTTP/2, the sender may create a new stream immediately after sending a reset for an existing one. In HTTP/3, the receiver is responsible for extending the stream limit after a stream closes, so there is backpressure limiting how quickly the sender may create streams.
¿Por qué no los dos?
:)
If a lot of people use the thing, it must provide some value to them.
No, it isn't. This whole article seems more like a marketing sales pitch than a disclosure.
- HTTP/2 as implemented by browsers requires HTTPS, and some people don't like HTTPS.
- HTTP/2 was "designed by a committee" and has: a lot of features and complexity; most of those features were never implemented by most of the servers/clients; most of those advanced features that were implemented were very naive "checkbox implementations" and/or buggy [0]; some were implemented and then turned out to be more harmful than useful, and got dropped (HTTP/2 push in browsers [1]) etc.
[0] https://github.com/andydavies/http2-prioritization-issues
https://en.wikipedia.org/wiki/HTTP_pipelining#Implementation...
But establishing a connection is extremely expensive compared to sending data on an already established channel. With this method they need to open far fewer connections for the same qps.
There's no need to confuse the issue by trying to diagram multiple connections at the same time.
> When Cloudflare's reverse proxies process incoming HTTP/2 client traffic, they copy the data from the connection’s socket into a buffer and process that buffered data in order. As each request is read (HEADERS and DATA frames) it is dispatched to an upstream service. When RST_STREAM frames are read, the local state for the request is torn down and the upstream is notified that the request has been canceled. Rinse and repeat until the entire buffer is consumed. However this logic can be abused: when a malicious client started sending an enormous chain of requests and resets at the start of a connection, our servers would eagerly read them all and create stress on the upstream servers to the point of being unable to process any new incoming request.
Sticking with HTTP/2, or going with grpc/similar is also possible. It depends on which corner of the Internet you inhabit. (Cloudflare isn't the whole Internet, yet)
Actually that's not true, it was already suggested here as a way to circumvent the max_concurrent_streams setting an it seemed particularly obvious: https://lists.w3.org/Archives/Public/ietf-http-wg/2019JanMar...
As soon as you start to implement a proxy that supports H2 on both sides, that's something you immediately spot, because setting too low timeouts on your first stage easily fills the second stage so you have to cover that case.
I think that the reality is in fact that some big corp had several outages due to these attacks and it makes them look better to their customers to say "it's not our fault we had to fight zero-days" than "your service was running on half-baked stacks", so let's just go make a lot of noise about it to announce yet-another-end-of-the-net.
It never occured to me that it could be used nefariously!
The exploit has more to do with their implementation than the protocol.
Is it? I imagine that implementations can do things like make creating/dropping a stream faster but how would an implementation flat out mitigate this?
It's called programming under soft real-time constraints.
In the real world, pipes cannot put arbitrary pressure, so your constraint is more bounded than this. So if you receive 2000 psi but your pipes can only handle 1000, you just need a small component that can handle the 2000 to split the pressure in two, and you can handle it all without releasing any.
The same applies to digital logic; it's always possible to build something such that you can guarantee processing within a bounded amount of time by optimizing and sizing the resources correctly.
As the word "digital logic" suggests, these sorts of guarantees are more often applied when designing hardware than software, but they can apply to either.
This is pretty much impossible unless you make the client do a proof-of-work so they can't send requests very quickly. Okay, you could use a slow connection so that requests can't arrive very quickly, but then the DoS is upstream.
The concept of a TCP session is purely virtual, it's just a 16-bit integer in the header grouping the packets together.
https://pentestmag.com/good-bad-and-the-ugly-of-http-2/
I genuinely don't know if this is real a zero day, or if it's a known protocol vulnerability that nobody was mitigating.