TCP Puzzlers
joyent.com
joyent.com
Fun anecdote, at Blekko we had people who tried to scrape the search engine by fetching all 300 pages of results. They would do that with some script or code and it would be clear they weren't human because they would ask for each page right after the other. We sent them to a process that Greg wrote on a machine that did most of the TCP handshake and then went away. As a result the scrapers script would hang forever. We saw output that suggested some of these things sat their for months waiting for results that would never come.
Not necessarily. If you have a good RPC layer, it will abstract away most of these details. For example, gRPC automatically does periodic health checking so you never see a connection hang forever. (But you do see some CPU and network cost associated with open channels!) gRPC (or at least gRPC's predecessor, Stubby) has various other quirks which are more critical to understand if you build your system on top of it. I expect the same is true of other RPC layers. Some examples below:
* a boolean controlling "initial reachability" of a channel before a TCP connection is established (when load-balancing between several servers, you generally want it to be false so your RPC goes to a known-working server)
* situations in which existing channels are working but establishing new ones does not
* the notion of "fail-fast" when a channel is unreachable
* lame-ducking when a server is planning to shut down
* deadlines being automatically adjusted for expected round-trip transit time
* load-balanced channels: various algorithms, subsetting
* quirky but extremely useful built-in debugging webpages
It's valuable to understand all the layers underneath when abstractions leak, but the abstractions work most of the time. I'd say it's more essentially to understand the abstraction you're using than ones two or three layers down. People build (or at least contribute to) wildly successful distributed systems without a complete understanding of TCP or IP or Ethernet every day.
In my experience an abstraction is only as strong as an organization's ability to hold its invariant assumptions, well invariant. And what I took away from that experience was that knowing how an abstraction was implemented allowed me to see those invariance violations way before my peers were starting to ask, "well maybe this library isn't working like I expect it to."
Huh. I don't see bug reports like that. "This RPC with no deadline hangs forever" definitely but I wouldn't call that a Stubby problem. I'd call it a buggy server (one that leaks its reply callback, has a deadlock in some thread pool it's using for request handling, etc.) and a client that isn't properly using deadline propagation.
So, a simple setsockopt(fd, SOL_SOCKET, SO_KEEPALIVE) call?
Sometimes the heartbeats are sent by the server, sometimes by the client, it depends. But they always end up in the application layer protocol.
Not really. Writing reliable network programs always had to take this into account. Take the scenario (not mentioned in TFA) where a firewall somewhere in the network path suddenly starts blocking traffic in one direction. Or handle a patch cable being pulled from a switch (or from your computer which has a different result). These scenarios always were real and resulted in comparable connection error states.
Do agree it's a good post though, explains it rather nicely.
Today most servers have shorter timeouts, mostly as a defense against denial of service attacks. But it's often the HTTP server, not the TCP level, that times out first.
[1] https://blogs.technet.microsoft.com/nettracer/2010/06/03/thi...
This is not true, in general. I think you're describing connections which have TCP KeepAlive enabled on them.
e.g. my home router has a connection timeout of 24 hours.
[1] https://tools.ietf.org/html/rfc0793 [2] https://tools.ietf.org/html/rfc5482
> The TCP user timeout controls how long transmitted data may remain > unacknowledged before a connection is forcefully closed.
As I understand it, this only applies if there is data outstanding. In the puzzler, there was no data outstanding. You're right that if there had been, the side with data outstanding would eventually notice the problem and terminate the connection. The default timeout on most systems I've seen is 5-8 minutes.
By contrast, the previous article you linked was about KeepAlive, which will always eventually detect this condition, but by default usually not for at least two hours.
All TCP keep alive does is send a packet every so often, which is actually something which is rarely actually set.
The socket should be marked ready for reading, but when you try to read you'll get zero bytes back: Something in your framework may not realize that -- truss/strace the process and I'd guess you'll see a 0 byte read followed by not closing it; alternatively you may not be polling the socket for read availability?
Some things would change if you intended for the socket to be half closed, but I don't think you do?
If you need to know sooner that your data isn't going to be sent, it's pretty trivial to set up a short timeout that overrides the system defaults.
[1] https://www.freebsd.org/cgi/man.cgi?query=setsockopt&sektion...
Setting those options isn't much different from setting a timeout in your poll() call.
- If you have a port-forward to a machine that is switched off then you can get ICMP network unreachable or ICMP host unreachable as the response to the a SYN in the initial handshake.
This can also happen at any point in the connection. Other ICMP messages can also occur like this (eg. admin prohibited).
It's always worth remembering that the TCP connection is sitting on an underlying network stack that can also signal errors outside of the TCP protocol itself.
Do you know at the socket API level how the ICMP unreachagle is manifested. Looked at connect() error and saw ENETUNREACH -- guessing that's the one?
sudo dtruss -d -t bind,listen,accept,poll,read,write nc -l -p 8080
dtrace: failed to execute nc: dtrace cannot control executables signed with restricted entitlements
If you copy nc over to /tmp/ and run dtruss again, it will work. Hope that helps others. Fun article, OP!
$ strace -tebind nc -l -p 8080
14:27:44 bind(3, {sa_family=AF_INET, sin_port=htons(8080), sin_addr=inet_addr("0.0.0.0")}, 16) = 0
$ strace -teconnect nc 10.88.88.140 8080
14:28:29 connect(3, {sa_family=AF_INET, sin_port=htons(8080), sin_addr=inet_addr("10.88.88.140")}, 16) = -1 EINPROGRESS (Operation now in progress)
Does anyone know if the various dtrace-based dtruss implementations can be made to do this?Are there any tools for 'and suddenly, the network disappears/goes to 99% packetloss/10s latency' etc that you can use to test existing software? I'm imagining some sort of iptables/tc wrappers that can take defined scenarios (or just chaos-monkey it) and apply them to traffic, possibly allowing assertions as to how the application should behave for various situations.
[1]: https://github.com/aphyr/jepsen
[2]: https://news.ycombinator.com/item?id=9417773
[3]: https://news.ycombinator.com/item?id=6451885
[0] https://www.snellman.net/blog/archive/2015-10-01-flow-disrup...
For some reason I remember that when a process crashes the OS's TCP stack sends an RST to the opposing side, rather than a FIN, preventing puzzler #3.
Puzzler #2 is an unclean shutdown with a client-side TCP timeout since the server never responds.
However since the minimum allowed heartbeat is (usually) two hours host reboot/shutdown is usually detected via a heartbeat at the application layer instead, which sends some no-op traffic every n seconds or minutes.
"It's the responsibility of the application to handle the cases where these differ (often using a keep-alive mechanism)."
Removing that using chrome's inspector made this text legible again.