Using the FreeBSD Rack TCP Stack
klarasystems.com
klarasystems.com
I know of one other operating system which has a somewhat similar feature, but not quite the same. z/OS supports running multiple TCP/IP stacks concurrently on the same OS instance [0]
Whereas this is multiple TCP stacks, but still only one IP stack (or maybe one for v4 and one for v6)
[0] https://www.ibm.com/docs/en/zos/2.2.0?topic=overview-conside...
Similarly, in z/OS, a single process can have sockets belonging to multiple TCP/IP stacks. There is a system call (setibmopt) which can be used to choose a default stack, and thereafter all sockets created by that process will be bound to that stack only (the "stack affinity" is inherited over fork/exec; can also be set with _BPXK_SETIBMOPT_TRANSPORT environment variable). Alternatively, you can call ioctl(SIOCSETRTTD) on a socket to pick which TCP/IP stack to use for that particular socket. There is also a feature, CINET, where the OS chooses which TCP/IP stack to use for each socket automatically, based on the address the process binds it to. CINET asks each TCP/IP stack to provide a copy of its routing tables, and then uses those routing tables to "preroute" sockets to the appropriate stack.
But I get the impression VNET doesn't allow a single process to use multiple IP stack instances simultaneously? If VNET is bound to jails, a single process can belong to only one jail.
One reason why z/OS has this multiple TCP/IP stack support, is historically the TCP/IP stack has been a third party product, not a core part of the OS. So instead of IBM's stack, but some people used third party ones instead, such as CA TCPaccess (at one point resold by Cisco as IOS for S/390). One can even use both products on the same OS instance, primarily to help with piecemeal migrations from one to the other. Other operating systems with a history of supporting TCP/IP stacks from multiple vendors include OpenVMS and older versions of Windows (especially 3.x)
>"However, when the loss is at the end of a transmission, near the end of the connection or after a chunk of video has been sent, then the receiver won’t receive more segments that would generate ACKs. When this sort of Tail loss occurs, a lengthy retransmission time out (RTO) must fire before the final segments of data can be sent."
I believe this whole passage is just describing TCP fast retransmit vs a retransmit timeout expiring. However if the final TCP segment from the sender is lost wouldn't the receiver also start sending duplicate ACKs as well? This sentence seems to indicate duplicate ACKs would not be sent if the last segment was the TCP segment that was lost. In other words a duplicate ACK from the receiver is lost and so the RTO expires.
The many video-streaming type workloads, the connection will go idle, for seconds or even minutes at a time. If the loss is at the tail end of some activity, before a period of idle, the recovery takes a lot longer than it would if there are further activity on the connection.
Anyone know if there's one that can deal with severe buffer boat? I have a connection where I control both ends and I have seen ping times exceed 20s under load. Throughput is highly variable so I can't just throttle the connection.
That being said it sounds like this connection is a good candidate for using BBR anyways so I'd give that a shot and see if anything changes.
https://blog.apnic.net/2020/01/10/when-to-use-and-not-use-bb...
Either way, the packet is buffered in the sense that it is stored in a buffer on the switch; otherwise the packets would be dropped, not eventually make it through.