Why TCP needs 3 handshakes
pixelstech.net
pixelstech.net
In total, there are four exchanges (two questions and two answers). However, if you look closely, the second person's reply of 'Yes' already confirms that they can both hear and speak. Therefore, the second 'Can you hear me?' is unnecessary. With just three exchanges (one question and two answers), both people know that they can send and receive messages.
A: Can you hear me?
B: Yes
A: What time is it?
B: 5 o'clock
A: Thank you, goodbye!
B: Goobye!
Nothing is lost compared to:
A: Can you hear me?
B: Yes
A: Yes
A: What time is it?
[...]
Both sides need to know to "commit" to the connection.
That conversation can also go like;
A: Can you hear me?
B: Yes
A: What is the full unabridged story of War and Peace?
B: <...>
A: I am going to a take a nap for now, goodbye!
B: ... Goodbye!
A: Can you hear me? B: Yes B: What time is it? A: ...
At the point that B has replied Yes, B knows that it can hear A and that it can send to A but it doesn't know that A can hear B. As long as A makes the first move in the rest of the conversation that's fine - the next message from A confirms that B's "Yes" was received, but if A has nothing to say then B has to send it's next query and hope that A received the Yes successfully. If it didn't then B thinks the connection is established but it actually hasn't been.
A: If you can hear me, what time is it?
B: Yeah I can hear you; it’s 5.
This sounds like a variation of the Two Generals' Problem: https://en.wikipedia.org/wiki/Two_Generals%27_Problem
In the two generals problem, the channel can fail at any time (which is what happens in real life), so no amount of handshakes can assure you. Because of the above, I don't agree with their conclusion that more handshakes is better. Either you assume immutable channels, so you only need three, or mutable channels that can fail any time, so you need infinite.
This is true if you only care about having 100% confidence. Sending more handshakes allows you to do statistical analysis to give you more confidence that the connection is reliable.
And that sounds bogus to me. Connection initiation isn't about testing if packets can reach or not. It's about acknowledgement, building a two peer consensus about how to handle upcoming packets from a peer, not reliability checks. In that sense, I don't even think the third step is necessary, but apparently, it's needed to handle the case of both endpoints going into a timeout loop, this article explains it perfectly: https://www.baeldung.com/cs/handshakes
So reaching consensus and testing for reachability in both directions are the same.
In a client/server scenario, the client knows the connection is good when it receives the SYN+ACK, but the server doesn't know until it receives the resulting ACK. So the third packet is necessary to communicate consensus to the server; it doesn't need to be a pure ACK though, it can have data, if the client's stack makes it possible to queue outgoing data before the SYN+ACK is received.
No, that's not the intent of a handshake. Assume a hypothetical Internet where every node has guaranteed connectivity to every other node that never fails. Do we suddenly lose the need to do a 3-way handshake? No. It's not about testing connectivity, that's semantically wrong. And what's the meaning of testing connectivity for a connection of an arbitrary length and with a quality of unknown degree throughout? It doesn't make any sense.
There is no "knowing the connection is good", there is a process of building up consensus. The peers are only interested in an answer to this question: "Does the other party assume a connection?".
If we knew the answer to that question beforehand, we wouldn't need a handshake at all. Reachability, transmissibility are all irrelevant. And, UDP actually works like that. Both endpoints assume willingness to connect. That's why you don't need a handshake with UDP.
TCP-FO worked liked that too, removing the need for a TCP handshake completely, because it could persist the consensus information.
Lack of consensus/lack of handshake on UDP is why protocols with large responses become tools for DDoS reflection, and TCP protocols with large responses don't. TCP protocols with large responses can still be used for DDoS, but only when the target takes an active part --- either filling its outbound bandwidth to serve the large response to distributed requestors or being induced into making a request with a large response that fills its inbound bandwidth.
TCP Fast Open is interesting, but didn't seem to go anywhere. Speculatively including one MSS worth of data in the handshake process seems like it could be useful in some circumstances, but the server side implications of processing a request without an expectation that the client will receive the response are hard. (Of course, TCP can't and doesn't guarantee the client will receive the response, but when you have recent consensus on communication, you can expect it will). The crypto context shows evidence of past communication, but does not show evidence of current communication.
The reason this isn't done in reality is mainly that using a server's resources to send a lot of data blindly, in response to a very small request, is wasteful (if the reverse path is not accessible, the server is using all those resources retrying for nothing). It is also dangerous, if the source IP of the SYN is forged, then the server ends up sending lots of data to a victim's IP. This often happens with DNS, for example.
I think now there are ways to do the 3-way handshake in hardware at hardware speeds, and only involve software if the connection has been vetted. This can protect against Denial-Of-Service attacks.
Do we need hardware for this? Syncookies have provided a software method to handle large volumes of inbound syn without memory restrictions since the late 90s, and it's been in all major platforms except Mac since the late 2000s; Apple forked FreeBSD's tcp months before FreeBSD added syncookies, and last I checked, Apple never pulled them in. My testing is a bit old, but I had more trouble generating line rate syns at 2x10g than handling them several years ago. Are syncookies in software enough at 100g? I'm not sure, but I'd assume so. There's plenty of things to hardware accelerate on a NIC, but syn handling doesn't seem worthwhile IMHO.
The problem appears if the server would like to be the first to send data on a new connection. If the server included its data with the SYN-ACK, that would work perfectly well with benign clients, but it would be a vector for DoS attacks. An attacker could send a small packet with a forged source IP, and cause the server to send a large response to the victim's IP. So, the server can't safely send data until it receives an ACK with its secret SEQ number from the client.
The baeldung article is just wrong
> The client, in turn, receives the delayed ACK and assumes that it refers to the last sent SYN message
ACKs have acknowledgement numbers, so this sort of confusion can not happen.
Particularly, the two-way handshake presents potential problems when the ACK message from the server delays too much. Thus, if a connection timeout occurs, the client sends another SYN message with a new sequence number (Z, for example) to the server. However, if the server previously sent an ACK (which is delayed), it’ll discard this new SYN message. The client, in turn, receives the delayed ACK and assumes that it refers to the last sent SYN message. Here’s where the error happens: the client will send messages with the sequence number Z, while the server expects messages following the sequence number X.
In TCP every segment has a SEQ number and an ACK number, regardless of it being a handshake segment or a data segment. This completely negates the described problem: the server's first response to a SYN includes "SEQ=Y, ACK=X". If the client which just sent "SYN seq=Z" receives the ACK for SEQ=X, it will drop this ACK and wait for a new one. Also, a server which receives a new SYN has no reason to drop it, it will send a new ACK instead.
If it is the client, then a two-way handshake is enough: client sends a SYN, server sends a SYN+ACK, then the clients data which ACKs the server's SYN. These days that is the most common model. HTTP and TLS work like that.
However, what if the server sends first? For example the banner of SMTP. In that case the client sends a SYN, the server sends a SYN+ACK, the client sends an ACK and only then the server start sending data.
In general, an operating system doesn't know what is going to happen. So the client's kernel will just send the ACK immediately even if it will be followed by a data segment shortly after.
We take for granted the inter-connectivity of most modern equipment, but to this day companies still try to create synthetic technology monopolies to cash-in. i.e. to sustain a tenuous service commodity out of something that has essentially been free since the mid 1990s.
Philosophically it doesn't matter TCP is imperfect, but rather that the inter-connectivity is compatible with the inertia of the installed infrastructure.
One can indeed optimistically ignore the TCP connection drop and syn part of standards to tunnel/reverse-proxy though certain censorship firewalls... but it still does not make it safe for the people that live under such regimes.
Does this make it more or less clear? =3
This is why as a network engineer I always advocate for open standards everywhere I can to avoid vendor lock-in. The classic one is using OSPF instead of EIGRP on Cisco routers (or their other proprietary protocols). Nowadays this is much tricker with cloud computing and black box stuff like SDN/SDWAN.
The drama that goes on inside Cisco could fill a soap opera season. =3
The delta-t protocol is also used in RINA, which was invented by John Day. It is also used in Ouroboros (https://arxiv.org/pdf/2001.09707), and I can confirm it works. ;)
Also related and of interest in this connexion: CurveCP and its handshake, https://curvecp.org/packets.html
> John Day has been involved in research and development of computer networks since 1970, when his group at the University of Illinois was the 12th node on ARPANet (precursor to the Internet) and has developed and designed protocols for everything from the data link layer to the application layer. Also making fundamental contributions to research on distributed databases. He managed the development of the OSI reference model, naming and addressing, and a major contributor to the upper-layer architecture. He was a major contributor to the development of network management architecture, working in the area since 1984 and building and deploying LAN products and a network management system, a decade ahead of comparable systems. Mr. Day has published Patterns in Network Architecture: A Return to Fundamentals (Prentice Hall, 2008), which has been characterized (embarrassingly) as “the most important book on network protocols in general and the Internet in particular ever written.” The book analyzes the fundamental flaws in the Internet and proposes what appears to be the only path forward. Today Mr. Day splits his time between making this new path a reality and teaching at Boston University. Mr. Day is also a recognized scholar in the history of cartography focusing on 17thC China, and is past President of the Boston Map Society.
Pilot: "You have the plane." Co-Pilot: "I have the plane." Pilot: "You have the plane."
The third statement is the pilot acknowledging that the transfer of control has completed. Without it, the co-pilot doesn't positively know that the pilot knows that control has been handed over.
As an aside, Air France 447 was a crash where the co-pilot was pulling back on the stick without the captain knowing. The captain couldn't understand why his controls weren't having the intended effect. Both pilots were making inputs, unaware of each other.
> Theoretically, even more than three handshakes would not guarantee a "completely reliable" TCP connection. However, through three handshakes, it can at least be confirmed that the connection is "basically usable." Increasing the number of handshakes would merely increase the confidence level in the "connection availability."
3 is an arbitrary number, but one that works well in practice.
Can Can
Transmit | Receive
Client SYN-ACK SYN—ACK
Server ACK SYNThe author made a mistake in the title, but the content is mostly correct (fails at "First Handshake", "Second Handshake" and "Third Handshake").
He should call them stages, phases or something like that.
TFA appears to use both terms sort of interchangeably, which is uncool when describing a formalized protocol standard!
Segments are encapsulated in IP datagrams, or packets, which are, in turn, encapsulated in other PDU types, such as Ethernet frames.
IP fragmentation is typically avoided, and in IPv6 it is only allowed at the originator of a packet - if a router is trying to route an IPv6 packet larger than the MTU of the outgoing link's MTU, it must drop the packet and send a signal about the MTU back. An IPv4 router may fragment the packet, but this rarely works well in practice.
You're referring to "segmentation" which is also defined in that RFC. https://www.rfc-editor.org/rfc/rfc9293.html#name-segmentatio...
Its not latency related. Its because there 3 packets involved in completing the handshake.
> If the delay sending a packet between the client and server is k, there are 3 hops involved in establishing a connection so the time is 3k.
Also I think calling it "hops" is not correct. A hop occurs when a packet passes through a router.
> Confusingly, this has nothing to do with how many devices are involved.
In a TCP handshake there are only two devices involved (disregarding any middleboxes that meddle with the connection).
what it really amounts to is a three phase transaction.