TCP Connection Repair
lwn.net
lwn.net
The technique sounds messy but actually it involves not much work when the target is Linux, a single half-open SYN is good for 63 seconds with the default sysctl settings, which seem to be used with almost every Internet service you might want to reach (including e.g. Google)
I was playing with this during an interview last year and intended to write it up, but never got around to it. The technique seems to work as intended, I made a little prototype reverse proxy for it in Python using a temporary listening socket with a drop-all SO_ATTACH_FILTER to allocate a port number and prevent Linux on the initiator side from responding with a RST to ACKs for a half-open connection it knows nothing about
Edit: oh here's a Meta talk on their QUIC CDN doing DSR[1].
The original "live migration of virtual machines"[2] paper blew me away & reset my expectations for computing & the connectivity, way back in 2005. They live migrated a Quake 3 server. :)
[1] https://engineering.fb.com/2022/07/06/networking-traffic/wat...
[2] https://lass.cs.umass.edu/~shenoy/courses/spring15/readings/...
Normal applications would typically just make sure the client can reconnect as fast as possible (and QUIC can do it with 0 to 1RTT), and then have suitable application level semantics that limit any availability issues when a reconnect happens (e.g. for large downloads you can restart using a ranged request. For persistent connection the server can tell the client with a GOAWAY that it might shut down and the client can reconnect early to avoid running into the availability issue).
My understanding is that for [1], their frontend proxy still is a single QUIC peer which contains all state of the actual connection - otherwise they also couldn't do connection level flow control and overall connection congestion control. That layer now just instructs another server about packet layouts to send, but doesn't make the other layer handle QUIC transmission completely on its own.
but given the place where we ended up, maybe host addresses make more sense than interface addresses (ignoring the effect that would have on routing table aggregation)
CRIU is a project to save an application or containers complete running state to a file, and then restore it elsewhere.
It comes with lots of caveats.
The actual details are far more funny and interesting (we could talk about checkpoint not being an atomic operation for the kernel, how you need to do some magic with "plug" qdiscs and qdiscs being applicable on egress only you'll look into IFBs and I love Linux it is so versatile and full of amazing little features). Don't forget to hot-update conntrack too...
And since libsoccr is GPL you might need to do this yourself, and you'll want to do it anyway, because it's interesting and you'll learn so many things.
My only gripe is the checkpoint still being a bit slow and maybe if I keep annoying Jens Axboe on twitter maybe soon it'll be a io_uring chain <3.
Apparently there are some workloads in finance that use similar hardware.
Not surprising, I've heard from someone who worked at a financial institution doing high-speed trading that basically every ms counts for them.
- you have a human in the loop to take a split second decision, every millisecond counts
- you have a very short time to perform 'looped' operations - where the result of one measurement must be taken into account to effect the next measurement (adaptive optics, some radar systems, some mechanical control loops) and you can't wait.
You'd think 'oh but you got more than 1 ms for that' Well not always since one must take into account the time to detect the failure, and the time to switch other parts of the system (which sometimes must be done in sequence with the connection takeover).
I'd say 'forget tcp' there but we don't always get to decide the comm layer...
If the window of unavailability was instead 1ms, there would be dramatically less noise, potentially none.
(But in practice, we mostly rely on the fact that if you are in Google Cloud and specify a minimum CPU version, Google masks CPUID in that VM to exactly match that CPU, even if the underlying hardware is newer. They do this so live-migration of the whole VM works, but it also helps for process migration. I'm not sure where this is specifically documented but see e.g. https://stackoverflow.com/a/44507857 .)
When you are elbow-deep into the state of the interface, those other issues should be pretty trivial.
SR-IOV, which you want for lower latency network connections can't do it cleanly. vDPA will hopefully save us from all that insanity.
Migrating VMs works pretty decently actually, biggest latency is from telling the network equipment where the traffic should now go.
The attached hardware is the issue, as you said. If all you talk is paravirtualized stuff it generally works, just need to make sure target CPU supports required stuff (we had to downgrade exposed CPU temporarily to migrate between intel and AMD hypervisors), but good luck with having GPU attached...
Moving to "migrate the process" (which is essentially what is required to migrate containers) adds a whole level of mess to deal with.
The whole use case of "I have a container that CANNOT EVER GO DOWN even when I migrate hardware" seems to be vanishingly small one
I don't have much context here, what's the use case where one would migrate a running container from one physical host to another?
In practice of course that’s easier said than done (eg what if there’s a proxy in between that had some data buffered in application buffers), but it’s a neat idea I want to revisit some day. And it’s totally possible that using something like QUIC is a better model than relying on migrating TCP kernel stack.
> It is natural to want those connections to follow the container to its new host, preferably without the remote end even noticing that something has changed, but the Linux networking stack was not written with this kind of move in mind.
Yes, it was written in the C language.
So, my point is the ability to take running state, serialize it, and reinstate it elsewhere is only impressive to those who have misused computers for so long that they don't understand this was something basic in 1970 at the latest.
To play some silly semantics games, this isn't so much about _serialising_ a connection as it is about _deserialising_ the connection and having it work afterwards. That act has literally nothing to do with programming language.
It is because UNIX is written in the C language that there are even multiple flat address spaces instead of segments or a single address space systemwide. The fact that the kernel exists at all is also due to this. It has everything to do with the implementation language.
> Languages "such as Lisp" will have the exact same problem, for the same reason.
Under UNIX, yes.
> Collecting all of the "components" of the connection and sending them to a different host won't make the other host start sending packets to the new recipient, or replay the in-flight packets (which is state on intermediate routers, different computers than the connected ones entirely), or fix the ARP tables on the neighbouring hosts. None of that is available, and certainly isn't writeable, to the host doing the serialising.
It may very well require some specialized machinery, but not nearly so much as one may think to be necessary.
> To play some silly semantics games, this isn't so much about _serialising_ a connection as it is about _deserialising_ the connection and having it work afterwards.
That's implicit. I needn't write of deserializing when writing of serializing, as one is worthless without the other, at least in most cases.
> That act has literally nothing to do with programming language.
Look at what Lisp and Smalltalk systems could do before UNIX existed and tell me that again.
That is flat out wrong. C supports multi-programming in a system that has one address space (that includes the kernel too). Programs just have to be compiled relocatable.
You know, like what happens with shared libraries: which are written in C, and get loaded at different addresses in the same space, yet access their own functions and variables just fine.
> Programs just have to be compiled relocatable.
Yes, and with unrestricted memory access, one program can crash the entire system.
> You know, like what happens with shared libraries: which are written in C, and get loaded at different addresses in the same space, yet access their own functions and variables just fine.
That is except when one piece manipulates global state in a way with which another piece can't cope, and at best the whole thing crashes. Dynamic linking in UNIX is so bad some believe it can't work, and instead use static linking exclusively.
So do MS-DOS, Mac OS < 9, and others: any non-MMU OS.
> Yes, and with unrestricted memory access, one program can crash the entire system.
That's true in any system with no MMU that runs machine-language native executables written in assembly language or using unsafe compiled languages.
Historically, there existed partition-based memory management whereby even in a single physical address space, programs are isolated from stomping over each other.
https://en.wikipedia.org/wiki/Memory_management_(operating_s...
This problem is the same with both static and dynamic linking.
And lisp too!
> UNIX breaks down quickly without multiple fake single address spaces for each program.
Citation needed. I don't think my programs very commonly try to go completely outside their address space. The closest thing I see is null pointer crashes, which are still not very common, and those would work the same way in a shared address space.
Edit: Yes, fork doesn't work the same. That's a very narrow use case on the vast majority of machines.