tcpcp - passing TCP connections between hosts (2005)
tcpcp.sourceforge.net
tcpcp.sourceforge.net
I keep hoping one day I'll be able to do the checkpoint operation in one chained io_uring op series :-)
If the purpose of using tcpcp is for 'load balancing' or w/e, then why does it seem to need the connections behind a single machine for routing? I assume this is what redirects inbound traffic from the peer to the new machine. But if it still requires a single machine for routing then it will be just as congested and not load-balanced.
What am I missing here?
[1] https://en.wikipedia.org/wiki/Common_Address_Redundancy_Prot...
The servers accept every 2nd connection + any established+related connection packets. All other packets are dropped.
went something like this:
- a routing node (a kernel extension) would receive packets and forward them (verbatim, same src & dest IP) to a backend worker
- backend worker was configured for same IP, which meant: a) arp had to be disabled to avoid conflict, or b) the loopback interface had to be aliased to same IP (using loopback avoids arp conflict and worker still happy to accept the packet and respond)
- response then goes back directly from worker to original source, which was a good approach for scaling web traffic back then, small http requests going through routing node, large responses going back direct from worker to source.
more details:
It might even be possible to run a service fully on a moderate fleet of ephemeral instances and save over half of the EC2 costs.
b) sure ok, if it's an http load balancer, you can finish up the existing requests and any new requests will go to new servers; but if it's a tcp load balancer, those don't generally let you swap the server in the middle
b) In general your right, but this is already a fairly niche use case. But I do wonder if reprogramming the conntrack entry in linux would work, or for lvs which already has a userspace daemon for state synchronization, how it would behave if reprogramming an existing state. It's at least not implausible to do that rewrite somewhere real time... if you control the edge system.
Also, again in a you control the network scenario, OpenFlow switches might allow you to reprogram the state tables while they're live.
The Linux kernel stack is crazy.
I'd actually be really interested in a write-up on such ideas. SSH, TLS or wireguard (voluntary) 'takeover' or Checkpoint&Restore. I don't remember whether QUIC had multihoming (might help with checkpoint restore) but since most APIs are in userland and it's udp it might be faaar easier.
There isn't anything out of the box that I could find, but there was some discussion/prototyping around adding an API for exporting all the necessary key material and metadata to the mbedtls API. With that it would have been "relatively" "easy" to do the TLS bits :)
See https://github.com/Mbed-TLS/mbedtls/issues/3141 and linked ML posts.