SYN cookies ate my dog – breaking TCP on Linux (2018)
kognitio.com
kognitio.com
Couple of caveats:
- you can jam syn cookies enabled with tcp_syncookies=2 sysctl
- syn cookies are generally bad because they prevent to negotiate window scaling. Window scaling is important unless you are doing low bandwidth like telnet :)
- you can somewhat negotiate window scaling when tcp timestamps are enabled. But enabling tcp timestamps in general case brings little benefit and wastes 12 bytes of each packet for basically no gain.
- for a bonus point, consider what happens when both syn cookies and TCP_DEFER_ACCEPT are enabled.
More about syn packet handling in linux https://blog.cloudflare.com/syn-packet-handling-in-the-wild/
As suggested by majke, CloudFlare's blog is another great source of material regarding protocol wrangling at scale.
If anyone else has similar blogs, bookmarks, etc please do share :)
I think timestamps are used as a possible source to compute one-way delay in latency based congestion controllers.
Reading the RFC, couldn't the requested window scale be jammed into the sequence number? Section 2.3 limits the window scale to <= 14 bits. This article suggests that the MSS category fits into 2 bits, so we take up 6 of the 32 sequence number bits for mandatory additional data.
The syn cookie would then be (MSScat | WSCALE | HASH(key | saddr | daddr | sport | dport | sseq | MSScat | WSCALE)). The hash could then be 26 bits, leaving a very small probability of forgery of 2^-26.
To re-address the problem of replay attacks, the server could rotate keys rather than include a timestamp in the data component. The cookie would need to verify with one of the (small) K most recent keys. If K were 2, then any ACKs younger than COOKIE_AGE would verify, ACKs up to twice COOKIE_AGE might verify, and older ACKs would be rejected. The probability of forgery would marginally increase to 2^-25.
But, I'm no TCP expert. I'm sure there's something misguided or wrong in the above?
MSS and Windows Scale are each turned into 3-bit table entries, SACK gets one bit, and one bit indicates which MAC secret was used; secrets are rotated every 15 seconds, two are kept, so client ACK needs to come in 15-30 seconds depending on where in the rotation cycle you are; 24-bits are used for the MAC. 15 second round trip covers the vast majority of internet connections.
Details nicely commented here: https://github.com/freebsd/freebsd/blob/master/sys/netinet/t...
In this context, it's easy to get a valid hash --- when the system is in syncookie mode, send a SYN from an address where you have visibility of the SYN+ACK responses.
Then, you could use that cookie (sequence number) to spoof ACK packets with other sources, and they've estimated the number of packets you need to spoof before you'll have probably generated a connection. That number of packets is significantly fewer than when the syncache has not overflowed recently, and you need to have sent a SYN, and have an exact match of the sequence number.
I disagree; TCP timestamps are awesome. Linux enables these by defaults.
Quick search gives me some measurements from 2012 [1] that indicate that TCP timestamps are enabled on 83% of the top 100k web hosts.
You can afford to waste 12 bytes; the bottleneck isn't these 12 bytes but how well you get congestion control to work. And congestion control relies on getting an accurate estimate of the round-trip time
[1] https://link.springer.com/chapter/10.1007/978-3-642-36516-4_... (paywall)
Edit: typo
Edit: Also, just because 83% of web hosts having it enabled does not imply that it is a good idea to do so in general. They could just all be running the linux defaults and these could be just wrong
It's hard to overstate how expensive TCP timestamps are. The thing is that they bloat every single packet including control packets. 2% of the world's bandwidth is being wasted on this.
The only reason for anyone to implement TCP timestamps today is that iOS clients have horrible receive window scaling if timestamps are disabled. (Well, that was the only reason a few years ago when I was still in the game of keeping up with the quirks of different TCP stacks.)
I wrote more on the subject at the time: https://www.snellman.net/blog/archive/2017-07-20-s3-mystery/
Kudos on having the integrity to point this out even when it superficially weakens your argument.
Also, a more relevant percentage would be: If you have 64KB packets (which you want, to get the best throughput) 12 bytes is a 0.018% overhead, less than one five-thousandth of your packet.
Edit: and apparently it's 0.8% with the default MSS of 1460B. Ugh.
initial
the word 'initial [sequence number]' is a term of art, not an adjective that can be substituted. very commonly called ISN. odd that the article never used the acronym.
> syn cookies are generally bad because they prevent to negotiate window scaling.
this is wrong thinking. they degrade performance under "attack", yes. the alternative is instant death. "attacks" need not be from the big bad internet either. my experience is that in the last 10 years most synflood activations are internal buggy or poorly throttled clients.
> enabling tcp timestamps in general case brings little benefit
disagree, but sibling comment addresses this.
From running a large messaging service I agree; most of the perceived attacks (synflood or otherwise) were actually coming from our own clients. But there was a periodic stream of strictly abusive traffic coming from who knows where; mostly UDP reflection, but a SYN flood every once in a while.
I guess you would have a similar lack of dogs, but also, if a connection was opened and closed quickly, a re-transmitted packet from the client would satisfy the SYN cookie calculation, and the server would re-open the connection, but at it's original sequence number.
The details are a bit hazy, but the client would get an ACK with SEQ behind where it had ACKed, and would send an ACK probe with it's latest values. The server would see an ACK ahead of where it had sent, and send an ACK probe. If the hosts had low enough round trip times, the number of ACK probes sent could be tremendous. For those unfamiliar with FreeBSD, the localhost interface on FreeBSD runs full TCP, and under high load can drop packets and retransmit. We ran into this on localhost first, but then later across the internet with external clients.
[1] https://github.com/freebsd/freebsd/commit/56ba0a68edde7b3832...
[2] https://github.com/freebsd/freebsd/commit/fc2be30b217171175b...
Then every packet other than the first data packet will be discarded as invalid, and eventually the client-side retransmits will take care that everything works properly.
The problem that is impossible to solve, however, is the lost third ACK (acknowledging SYN-ACK) from the client , if the client doesn’t send any data to server upon the connect. It’s sufficiently rare in today’s protocols, though.
Another problem that the above approach will create afresh is that it assumes that the retransmitted client SYNs will have the same ISN, which isn’t the same in practice with e.g. some load balancers (who also try hard not to keep the state). And that behavior is kinda a slightly gray zone in the TCP spec, IIRC...
Edit: (I wonder if the last paragraph above is the real reason or I missed something else)
Edit2: oh, thanks to Majromax’s mention of the DJB’s write-up, the above has the problem of not complying with “sequence numbers increasing slowly”, and indeed brings up a real-world scenario where that approach was an issue - using rcp/rlogin protocols, which reused a very narrow range of source ports, so the 5-tuple reuse was common.
I don't think this is a big problem with SYN cookies. If you get a SYN with initial sequence X, you send an appropriate SYN+ACK, and if you get a retransmitted SYN (because the other end didn't get your SYN+ACK), you send a new SYN+ACK appropriate for that one. If you then get an ACK for either, you would form a full connection; which should work fine.
I would have to review the RFCS, they might say that if you had room in your syncache to hold the data, you should send a RST to the second SYN or the first SYN, because the states are conflicting; but since you don't have the information you don't have the information.
Anyway, unless the client end is really messed up, it shouldn't send both the ACK on the first SYN, because it received your SYN+ACK and a new SYN, because it didn't receive your SYN+ACK. I acknowledge that there are plenty of really messed up TCP stacks on the internet though :)
This is a race condition; hypothetic sequence of events:
send SYN-0, wait for reply or timeout
timer interrupt fires
timeout to resend SYN(-1) is ready, start running that
packet interrupt fires (interrupts resend)
got SYN+ACK-0, construct and send ACK-0
iret
finish constructing SYN-1 and send it
iret again
This is clearly a bug, but it could easily work >99.99% of the time (especially if the timeout is high enough that normal RTTs never hit it, which is probably how the person setting the timeout would try to set it).Overall, I'd rank that like a 3 out of 10 on the scale of tcp bugs in the wild.
--- a/fs/dlm/lowcomms.c
+++ b/fs/dlm/lowcomms.c
@@ -1209,7 +1209,7 @@ static struct socket *tcp_create_listen_sock(struct connection *con,
log_print("Set keepalive failed: %d", result);
}
- result = sock->ops->listen(sock, 5);
+ result = sock->ops->listen(sock, 128);
This is mostly a theoretical issue, because upstream RHEL used to support only up to 16 hosts in a cluster.Network is fascinating. I initially introduced myself to the subject by reading Michal Zalewski's Silence on the wire [1]. I really recommend it to anyone who wants to do the same.
From that security-based perspective, this bug seems to belong in a common category of data escaping the hash -- here where the (sequence number + MSS category) sum has hash collisions.
Sure, this time it's a software design bug in the endpoints, but next time it might be a cosmic ray, or an evil middleman, or a buggy proxy. If data isn't encrypted and authenticated, then you shouldn't care what form it arrives in.
You're referring to things that happen in a different OSI layer
Yes, it needs a failure detection mechanism, because next time it might be a cosmic ray. But encryption and authentication alone doesn't help necessarily.
To lump this into an existing category of bugs, syn cookies are a kind of HMAC, only the implementation is custom and nonstandard. It isn't a surprise that a bespoke HMAC leaks, but to the credit of kernel developers syn cookies the initial 1996 specification pre-dates the common understanding. (But to its demerit, it looks like the DJB spec (http://cr.yp.to/syncookies.html) would have not had this issue, since the MSS was encoded in the top bits of the cookie and not the bottom bits.)