What developers should know about TCP
robertovitillo.com
robertovitillo.com
That means whatever you pass to a send() call is not necessarily the same amount of data the receiver will observe in a single read() call. You might get more or less bytes, since the transport layer is free to buffer and to fragment data.
I have seen the assumption of TCP having packet boundaries on application level being made too often - typically in stackoverflow questions like: „I don’t receive all data. Is my OS/library broken?“
The idea being that if your software is actually written to relevant standards, and actually handles things properly outside the golden path, then it should still work fine. If, however, you accidentally did something implementation-defined, or that only worked by coincidence, this system will break it.
That will most likely not help newcomers which directly write their code agains the OS socket. But once you get a better understanding of the topic and start adding tests to your codebase it's rather easy to add.
The other linking/OS problems can probably be automated with some simple integration tests and a bunch of different docker containers to compile the code in. Should be possible to squeeze it into a CI/CD flow somewhere with some clever tricks.
In between the emails I googled a bit and found the changelogs for the RTOS they were using. Turned out that it was a bug in the upstream HTTP server. This also meant that the platform they were using had all the security holes from those five-plus years. The bug was later silently fixed when they acquired a newer release from upstream.
Currently I'm having a similar issue with the very same vendor. This time they don't understand why client-side authentication means no authentication at all and why passwords must not be stored in plain text in the database that can be remotely backed up from the device.
I would assume, however, that there is no law forcing minimal security so you can class A them, can you?
Because nobody gives a shit about quality unless it hits their paycheck.
"Onnnngggg. They pay me hourly. Onnnngggg."
They cut lots of checks.
The longer it takes for this problem to become public, won't the more harm be caused when it does become public?
I've seen this... with an intern! I can't imagine dealing with a whole team like that.
The network layer on which the transport layer rides is packet switched. The TCP uses segments with each segment having its own header and sequence numbers. Streams are just a series of segments populating across a single established handshake without a prior defined termination segment.
That is still a bit imprecise. Userland applications won't directly see TCP as they are just looking at an application protocol. Typically it's the OS that packages and unpacks the application protocol data into a TCP segment, so of course the userland application won't see it since its not managing that part of the communication.
https://en.wikipedia.org/wiki/Transmission_Control_Protocol#...
There are some exceptions where some application platforms allow developers to write custom TCP protocols, such as Node.js, but these exceptions generally apply to network services and don't commonly apply to the end user application experiance.
https://nodejs.org/dist/latest-v14.x/docs/api/net.html#net_n...
The thing to turn off is delayed ACKs. See "TCP_QUICKACK". Delayed ACKs were a feature which is only useful for things like Telnet, where the payload in each packet is one character when the user is typing. The fixed timer for delayed ACKs is for keyboard typing speeds, and for networks so slow that human typing could congest them. There's a reasonably good explanation here.[1]
As others said above, TCP is not a message protocol. It's a stream protocol. If you're sending messages over a stream, you need something that's reading data from the stream, and when it has a full message, it send that off to be processed. There is no set of TCP options which will reliably cause one write at the sending end to result in one read at the receiving end. If there were, it would be inefficient for small messages and would fail for large ones.
[1] https://www.extrahop.com/company/blog/2016/tcp-nodelay-nagle...
Not sure it says anything you haven't, but a StackOverflow answer on fragmentation (framed by asker as Go not behaving like C) is one of the more-read ones I've written: https://stackoverflow.com/questions/26999615/go-tcp-read-is-...
> That means whatever you pass to a send() call is not necessarily the same amount of data the receiver will observe in a single read() call.
Yes, this. For god's sake, listen to them.
I had to fight a coworker on this. I had quickly created some client code just to validate that the server was working. Due to some quirk, all the messages were arriving in full in every read call. He told me to ship it.
I said no! "I need to check if there's more data and if so add a loop to read again" "But it is working, release it". That went on for a while, to no avail. Wouldn't look at documentation either.
Eventually he head to leave for the day, and I took the time to implement it correctly.
I started including basic TCP questions on interviews. Not many people even get past the TCP handshake (if they even know about that).
Famous last words :-)
If you want to use TCP to implement some (higher-level) message-based protocol, you need to parse those out for yourself.
Edit: On second thought, I guess OP meant that all of the results were coming back "complete" which doesn't obviate the issue of needing to do a check that handles the "not done" case.
You would be wrong.
That naturally leads to the question of why it's not the default, which can be answered by understanding the history of TCP and computer networking; and more interestingly, how might things have been different if MSG_WAITALL was the default from the beginning.
>I had to fight a coworker on this. I had quickly created some client code just to validate that the server was working. Due to some quirk, all the messages were arriving in full in every read call. He told me to ship it.
>I said no! "I need to check if there's more data and if so add a loop to read again" "But it is working, release it". That went on for a while, to no avail. Wouldn't look at documentation either.
I was just going along with that.
I didn't check up on TCP further to verify whether this was actually true; if you are saying that's not a well-defined concept, you might want to reply to that comment to say so.
(Edit: also, it would help if you said what the correct interpretation would be, since it has the phrase "messages were arriving in full in every read call".)
Not at the TCP level they don't.
They just give you some bytes. It's up to you to decide whether you have the full message or not and if you want to try to read more. The TCP read functions just give you the data they have. There is no concept of one write at the sender end translating to a complete message at the other end. It's just a stream of bytes.
The caller absolutely must concern itself with this crucial “detail”. If you do otherwise, your code is broken, full stop.
You can implement a higher-level protocol on top which handles this kind of thing internally and presents a higher-level interface (e.g. not passing any partial data along to its caller until a full message has been received), but if you are just working with TCP directly, what you get is just a stream of bytes. The guarantee you get is that the bytes will be in order and without any gaps.
If you e.g. send UTF-8 encoded text, you must be prepared on the read side to have your stream of bytes cut off arbitrarily in the middle of a character.
Let function B call TCP read() and never return anything until it's received all the bytes of the message.
Both of those seem (IMHO) like functions you could have in a TCP library. Neither seems (IMHO) like a higher level protocol.
So the job of separating the stream into messages is left to the application layer. (unlike for example in UDP, but then you have to worry about dropped messages)
If your point is just that TCP doesn't have a concept of a "message" (a bytestream with a clear beginning and ending), then that's fair, but, as I said elsewhere [1] the original comment took for granted that TCP does have a well defined notion of "you've reached the end of the message", or at least, "there is no further data to receive". No one seemed to have a problem with that there, and I was just working off that assumption.
As before, I haven't checked whether this is true (I can't quickly verify from descriptions of TCP).
And, interestingly enough, there's this comment [2], which says that what I described does exist, but isn't the default. So ... I'm at least not getting a consistent answer to my question, and people who think they know what they're talking about are inconsistent with each other.
No, the original comment was complaining about a coworker who didn’t understand (and refused to listen when told otherwise) that there is no such notion in TCP. It was a response to another comment complaining about people on the internet (e.g. Stack Overflow) too often making the same mistake.
You’re more or less playing the part of that coworker here. It’s unclear why.
It's responding to the last part. It is a higher level of protocol. In the traditional TCP/IP model, it's in the application level. There are many libraries with an API like you asked, they are just in a higher level.
(and TCP_WAITALL is a partial solution, applicable only if you know in advance the exact size of the message you are about to receive)
The OS will buffer bytes received from TCP packets for you until you read from the socket again to drain the buffer. Your application needs to determine how to semantically chop those bytes up into the protocol it's expecting (e.g. http request).
My low-level networking chops are a little rusty so please correct my understanding if I'm off-base somewhere.
I don't see anything in the comment you linked that implies that.
Sorry you got confused, but there's no need for anyone to go reply to that comment to say so.
I've always wondered: What's the best/defacto way to delimit this back into packets at the application level on the receiving end?
I would think the obvious approach would be to insert some magic word into the stream so that you can re-sync.
Or is this not an issue since you know that once you're connected, you'll never drop a single byte, therefore, the only way to get out of sync would be a program error?
If you need some packet-oriented messaging, you could use something like http://jsonlines.org/ (i.e. JSON messages separated by newline characters), or https://github.com/protocolbuffers/protobuf if it's more performance-critical.
I like zeromq to get to a packet based system.
For example if the message is x bytes long then you first send 'x' then you send the x bytes of the message.
Or your messages have a defined header that contains the length of the message payload.
The best approach is typically put a length in front of every message. The good things about that approach are:
1. The receiver can allocate buffer that is exactly the size it needs to fit the message. 2. The receiver can check whether the message is too long before seeing the entire message.
The only disadvantage is that you have to know the length of all messages in advance.
Or see if there's an existing protocol you can abuse for what you want. If it's transactional, you get a pretty big ecosystem of battle-tested clients/servers/proxies/etc if you use HTTP.
or just use http.
The first quick fix was to unconditionally realign the fifo contents after every write (the fifo had a realign method), but that ran into a computational complexity problem when you had lots of small lines (e.g. the application caller dumped a huge message into the buffer and then flushed it out in one go) and a high-latency connection that resulted in many short writes; you were constantly memmove'ing the megabytes of remaining contents in the buffer for every tiny write you did. So then I ended up having to add a new interface to the fifo that returned a slice up to a limit but always ending with a specified delimiter (e.g. "\n") if the delimiter was within the maximum chunk size.
Of course, none of these fixes would have completely remedied the issue as lower layers (the TLS stack, the kernel TCP stack) could have still potentially split logical lines, and I'm sure did on occasion. But it at least seemed to put us on equal footing with everybody else in terms of how often it happened, which is really the best anybody could have done. Complaints did die down.
If you're going to read it, though, find a used copy of the original Stevens' first edition, not that terrible desecrated second edition.
This is one of the best written textbooks, if not the best, I have ever read.
* An Engineering Approach to Computer Networking: ATM Networks, the Internet, and the Telephone Network by S.Keshav
* TCP/IP Illustrated Vol -I by Richard Stevens (any edition will do).
* Effective TCP/IP programming by Jon Snader.
While so, so, so much of this is rarely the network, knowing how to look under the covers and see what's actually hitting the wire (versus what the API call asked for) leads to far, far faster resolution of problems.
It's frustrating to me that so many people see this as a mystery of "knowing networking" when it's really just basic protocol analysis.
That's fine. But every developer should have a basic understanding of networking. But that can also be dangerous.
I still have people in the company who swear you can't have more than 65k incoming connections to a machine, because "that's how many ports there are". Don't get me started on all the misconceptions on TCP_TW_REUSE AND TCP_TW_RECYCLE. Lengthy discussions because apparently "TIME_WAIT is bad and uses up ports! "(see also, 65k). For context, these are servers, with multiple clients, from different source IPs.
I wonder, why would somebody use 204 bytes -> 1632 bits, why not less (why not more for e.g. jumbo frames). Is there some data sheet / source that you would recommend?
Turned out our cloud provider's networking gear had a bug that disabled ECC and there was a bit flip happening. Convincing the provider's support that we had found faulty hardware in their datacenter was an interesting journey.
write(2) syscall returned without a error means that data has been placed in OS kernel buffer. OS kernel then will try to send it to a remote host. If couple packet will be lost it's not a problem - kernel will retry a few times. But if power will be lost shortly after a write, data may never hit the wire. Then there is possibility that network link will be broken for a long time. OS will retry, but for a limited time and then will give up. Also remote host can crash at any time before remote application actually will read the data.
So if you need reliable delivery you need acknowledgement on application protocol level despite the fact that TCP already have acknowledgements.
I swear a high percentage of their calls were questions about how the load balancer wasn't working and sending all the traffic to one server and then after some investigation we discover all traffic is in fact directed to that lone server... because the client code has the IP of that server hard coded. A tedious discussion would then ensue about how that is not how to do it.
The next week? Same angry call...
Partly that is what inspired my decision to change careers. "Man if these developers can't figure out basic networking, maybe I could be a developer...?"
there is work that shows that higher RTT connection do statistically suffer a smaller fair share, but that's a subtler if related issue. actually, I really wish the author would have shown the sawtooth.
Being closer means faster initial 'slow start', but also faster 'slow start' on congestion, which is why you get a bigger share.
So the articles (unstated) conclusion seems to be that, as long as there isn't network congestion, it is smooth sailing after that.
But that congestion reduces bandwidth. But of course, that applies just as much to a national backbone as to last-mile.
So I'm curious: where does most packet loss occur? Is it last-mile, at your ISP, or along major backbones? Because that has major implications as to whether caching video content closer to users actually results in higher-quality video (e.g. supporting 1080p instead of 720p) or not.
Here's an interesting paper from SIGCOMM (it won best paper at the conference in 2018, FWIW) that attempts to figure out what links are congested without direct access to ISP networks: https://www.caida.org/publications/papers/2018/inferring_per...
So yes it would really help to have more decentralisation. Like putting the content closer to the user.
I recently started doing off-site backups, which requires my entire internet uplink to be used for uploading said backups for about a week at a time. The internet basically becomes unusable because all the packets end up in a buffer on the router and latency spikes to 5000ms.
[1] https://www.bufferbloat.net/projects/bloat/wiki/What_can_I_d...
In any case, the solution to bufferbloat is queue discipline, not congestion control.
What do you mean by queue discipline?
What about the queue discipline?
Most of the congestion control algorithms use packet loss as the only indicator of congestion. In a network with oversized buffers, congestion will result in delay and not packet loss. If the delay gets large enough, recieve and congestion windows will restrict the effective bandwidth, but the latency at that point is terrible.
There are some alternate congestion control algorithms which do use latency as a signal, but they aren't universally available, and may not be a good fit for all flows.
For your backup use case, probably the simplest thing is to reduce your sendbuffers for the backup sender process. Although allowing packets to drop instead of queue at your router/modem would really be best, often that's difficult to acheive.
Note also, Apple is using MP-TCP and ECN in iOS, and the world didn't stop. It might not work everywhere, and I don't praise Apple lightly, but there's a pretty clear path to using things like this. Send a syn with it enabled, wait a bit, and send one with it disabled. Keep track of networks where it doesn't work and stop trying it there. If you have leverage, yell at people to not do dumb things, otherwise, let them figure out why expensive things work better on their competetors' networks. You can't rely on being able to use these things, but you can use them for progressive enhancement.
It can. Enable BBR + fq/fq_codel on the box in question and CAKE on your router.
There are algorithms that try to use increased delay as a signal that the link is full. This approach has multiple problems, one of which is that delay can be really noisy on wireless networks; another is that if you have a loss-based and a delay-based connection sharing the same link, the delay-based one will get much less than a fair share of its bandwidth. People have been trying to make an algorithm that both coexists with Reno/CUBIC and does not induce bufferbloat for the last 25 years or so, and there's been some progress, but none of it has reached the point where it could be used as a default congestion control for all operating systems.
The problem of "I have files to transfer in background, but I want my connection to yield to more important traffic" can actually solved using a special congestion control algorithm called LEDBAT [1]; it's used by Apple for things like software updates, and BitTorrent uses it too. Unfortunately, I think only Apple implements it in its TCP stack, so anyone who wants to do that would have to roll their own thing using UDP.
This gives the sending TCP algorithm the wrong impression. It's waiting to hear about a dropped packet to indicate that there's congestion. When your router holds on to those packets (instead of dropping them), the TCP algorithm doesn't get any feedback, so it keeps shoveling data into the connection.
This leads to the bad state you're seeing. And that's where the advice on "What can I do about Bufferbloat?" comes in.
There's no benefit to having more than one packet buffered by the router. (Hanging on to more than one packet per connection only causes the latency/lag you're seeing.)
There are routers that actually check the time the router has held packets. If packets have been queued for "too long", the router discards them immediately, giving the vital feedback to the sending TCP. Those routers use the technique known as SQM (Smart Queue Management) and the fq_codel, cake, PIE algorithms to keep the queues within the router short - typically less than 5 msec.
To solve your problem, investigate getting a router that implements one of those SQM algorithms. They're listed on the "What can I do..." page. I am a fan of OpenWrt (use it at home), but have installed a bunch of IQrouters and Ubuquiti devices for friends.
https://tools.ietf.org/html/rfc793
TCP state machine diagrams can be useful too.
So, I hope people learn to check their http client/server implementations to have proper connection handling. Client should have a thoughtfully sized bounded connection pool with reasonably large idle timeout. It shouldn't close the connection after every application request (say, http request). There shouldn't be sockets in TIME_WAIT state accumulating at the client end.
Server should accept thoughtfully limited number of connections per client. Server should never close the connection except when it is shutting down.
There should be tcp keepalive messages to keep the connection alive with intermediate hop stateful firewalls (connection tracking table entries in firewalls expire when the connection is idle for too long) and to detect stale connections and re-establish them.
All of these things can be verified by analyzing at a packet capture. You can get a manageable sized pcap file by filtering on client/server ip/port-range pairs for at least 330 seconds.
Knowing tools to understand/debug tcp issues is an essential skill. sock stat command - ss, wireshark/tshark with Lua scripting is super useful. Knowing higher level application protocols like TLS and http is essential too.
I am super puzzled why something like websockets not solving this problem, simple heartbeat could solve the problem, but no one implements it.
SO_KEEPALIVE is available on all relevant OS.
I don't see how this breaks the whole Internet. Yes the server might not behave, but that's not a TCP problem.
The libraries I use tend not to enable it by default, but they are generally implemented.
Even MDN is not mentioning it: https://developer.mozilla.org/en-US/docs/Web/API/WebSocket
Pings are sent from the server to the browser. Browsers are supposed to automatically send Pongs back to the server. So a client implementation has no reason to know about Pings or Pongs—it's handled by the server or the browser.
Networking issues are always on customer side, you have to detect it from the app otherwise it will literally freeze.
It was really fun expecially because it allows you to understand better all networking layers.
I did some tests about network topology to minimize lost tcp packs as possible, given different network traffics
Turning on TCP_NODELAY was a quick-n-dirty fix, but the real fix was to rewrite the handshake to be more compatible with the inner workings of TCP.
Wait, this isn't TCP, this is protocol level above TCP, right? TCP doesn't shape traffic by itself through rate limiting and congestion analysis, does it? I thought the layer above it used TCP to send/receive the buffer size, and that has nothing to do with TCP.
Am I wrong?
Does http3 fix this?