Why TCP over TCP is a bad idea (2001)
sites.inka.de
sites.inka.de
Edit: Obviously, it still reacts doubly as bad to the base connection dropping packets. I recall testing once how a TCP connection behaves when you subject it to a random % of packet loss (regardless of bandwidth it tries to use), and it resulted in completely stalled connections at a surprisingly low drop rates.
Edit2: In case anyone wants to try themselves, here's how: https://www.pico.net/kb/how-can-i-simulate-delayed-and-dropp...
That was with BitTorrent uploads of Linux ISOs to Taiwan. (Why do they download so many copies of Ubuntu 14.04 LTS?)
But since I didn't do controlled tests of multiple congestion types I could just be seeing things.
What exactly do you mean with "completely stalled connections"?
Do you mean, that the sending side is queueing up to-be-send messages and can't clear the queue because it is working on correcting packet losses all the time, so the queue will just grow and never shrink?
Do you recall at which % of packet loss this behaviour started?
I can't quote a number for the drop percentage, honestly I've forgot. Discovering the limit was a side effect, we were simply looking to test how a piece of software would behave over a bad connection. I just remember being surprised that everything just stopped when I put in packet loss that wasn't anywhere near 100%.
My since-then-adjusted expectations would put any double digit percentage (yes, starting from 10%) of random packet loss as unusable conditions for TCP.
However, firewalls are usually much more permissive to TCP than to UDP. I wonder if there is any project that encapsulates UDP-semantic datagrams into TCP-looking segments?
My devices tend to try to connect back via
* UDP port 443, sometimes works
* an sstp vpn
* SSH to tcp/$highnumber, sometimes they blck/MITM port 80, 443, but leave the standard
* DNS
I can't think of a time that one of them didn't through.
But all implementations I know use a much shorter timeout/keepalive period for UDP than they use for TCP because of firewalls/NATs. (I think the RFCs even recommend something like 300 seconds for TCP, but only 30 for UDP as a default?)
This has pretty significant implications on power consumption for mobile devices.
A better option would be for the tunnel to generate fake ACK to avoid retransmission happening?
Imagine what happens when, in this situation, the base connection starts losing packets. The lower layer TCP queues up a retransmission and increases its timeouts. Since the connection is blocked for this amount of time, the upper layer (i.e. payload) TCP won't get a timely ACK, and will also queue a retransmission. Because the timeout is still less than the lower layer timeout, the upper layer will queue up more retransmissions faster than the lower layer can process them. This makes the upper layer connection stall very quickly and every retransmission just adds to the problem - an internal meltdown effect.”
[1]: https://yggdrasil-network.github.io/2018/07/13/about-mtu.htm...
* Unless it's OpenVPN, which can run SSL over UDP.
The key thing to remember is while TCP over TCP is a bad idea, bad connectivity to your workplace network is not worse than no connectivity at all, which is why SSL VPN solutions exist.
Sadly OpenVPN over UDP does not work as good - DNS queries are failing and so on. Wireguard falls even more behind.
TCP can use this to detect the difference between delay, and loss, and treat them differently.
Unfortunately, not all transports are in-order, and there is no way to know from the endpoints, hence the large number of TCP retransmit algorithms trying to find some optimal heuristic...
...or just use UDP?
If you absolutely must do TCP over TCP the only sensible thing to do is have the outer tunnel terminate and buffer the stream internally. This is fairly memory intensive so routers won't do it, but applications can get away with it.
So then how to VPN's work performantly?
I'd always assumed my VPN implemented TCP and UDP over a TCP connection. Do they not? Is the VPN connection actually just UDP?
See IPSec (most common VPN implementation): https://www.cloudflare.com/learning/network-layer/what-is-ip...
See Wireguard (increasing popularity): https://www.wireguard.com/
TCP is protocol 6, UDP protocol 17; while IPsec uses protocol 50 for ESP (encrypted) and 51 for AH (authenticated).
For pragmatic reasons (firewalls, NAT, that sort of thing) it is nowadays mostly tunneled over UDP.
Modern VPNs will usually attempt UDP first, falling back to TCP only if that does not work; unfortunately, OpenVPN doesn't seem to have that option and requires manual configuration in my experience. This means that many sysadmins configure it to use TCP for the higher success rate in most environments.
IPsec is UDP only, as is Wireguard.
(ignoring IPSec NAT-T)
Also, UDP and ESP can almost be used interchangeably for this discussion (flow control, congestion control, segmentation etc. or a lack thereof).
The average wired connection has 0% drop rate, losing less than a packet in a million. Try to use TCP on a shitty connection and it's a different matter.
Could this be explained by TCP over TCP?
Still, TCP over SSH works. It's not perfect, but the problems are almost not noticeable. It used to work by 2001 too. There is something that makes all of that to not apply most of the time. I believe that unused capacity makes the problem possible to recover from.