2. TCP doesn't have encryption either. They should probably show DTLS working though.
4. You can't just say QUIC is faster without testing it.
2. TCP doesn't have encryption either. They should probably show DTLS working though.
4. You can't just say QUIC is faster without testing it.
That is a sender-driven resend and as such relies on the sender identifying the “send lost” condition. However, the entire protocol is designed around not doing that and thus you can only feasibly rely on “should have received a reply by now” as your timeout.
But Homa is intended to be a RPC protocol. So you send to server, wait for server to process command, then wait for server to send reply. Your timeout depends on the variable and heterogeneous command processing time.
Even if you were able to give a separate correctly tuned timeout for every possible RPC that is still awful. Any RPC with long processing time should not trigger the timeout until the expected reply time, but if that is far larger than the RTT then you are waiting a tremendous amount of time.
For instance, a select query that is only a few bytes long (and thus fits in one packet) could take seconds on a large database even though the database server is physically nearby and only microseconds away. In that case you would have to wait seconds before timing out instead of just microseconds like a ack-based design could achieve. The worst-case transport latency becomes application-specific instead of related to the application-independent transport parameters like RTT. That is troubling.
2. TCP can be forgiven for that given it predates asymmetric cryptography and even DES. Not considering it relevant in a new protocol this side of the millennium is much less reasonable.
4. MsQuic at 7.5 Gbit/s: https://microsoft.github.io/msquic/
I don't see anything particularly Homa-specific in these issues. A protocol like TCP that doesn't implement RPCs doesn't have to face these issues; it punts them up to the application, so the application has to deal with them. Homa implements all of this in the transport, so applications don't have to worry about it.
In a lossless RPC the client should expect the reply in 1002 ms. When should the client resend the request if it got lost?
Given that the client would not expect to see the first byte of the reply until 1002 ms and receives no other feedback until then, they should not issue a resend of their request until 1002 ms.
That means if the network lost their first request, they will not send a resend until 1002 ms and then get the reply to their resend at 2004 ms.
In contrast, in a ack based approach, the client would expect a server ack in ~2 ms as the server gives feedback on receipt rather than on processing completion/reply start.
That means if the network lost their first request they will not send a resend until 2 ms and then get the reply to their resend at 1004 ms, nearly half the latency.
In ack based approaches, packet loss requires one additional RTT latency to recover per repeated packet loss. But Homa incurs one additional RTT + processing latency to recover per repeated packet loss.
As a secondary downside, this requires the client to hold Tx resources for RTT + processing time since the client does not know when the server has received all of the data. In ack-based approaches, you learn the server has received all of the request when it acks all of the data in the request which is on the order of the RTT and thus can release Tx resources at that time. This is similar to the reason Homa has ACK/NEED_ACK packets except applied to client side resources instead of just server side resources.
The case you are describing is more like what if the server explodes, but you still need the results/modifications related to your RPC to occur. In that case, whether the Tx got through or not is irrelevant because your RPC still did not complete and you need it to complete, so there is no advantage to the ack-based approach. But by that logic, what if it explodes before persisting changes? Or what if all your servers explode? At some point we should call that out of scope for a transport protocol.
I prefer setting that line at “bytes got from one end to the other”. If you want higher guarantees then you can layer that. As you point out, it is easy to just layer a “server dead” timeout ping which basically solves the problem in the same way Homa needs to.
I agree that a transport cannot hide all possible failures and that at some point things have to get kicked back to the application. But if the only guarantee provided by the transport is byte delivery in one direction, that makes things significantly harder for apps: every app now has to implement its own timers and pings to detect server failures. This basically duplicates timers and retries already present in the kernel. I think Homa's approach makes life easier for apps. There's a clear and simple division of responsibility. Homa detects all failures, including both lost packets and server failures. It recovers from packet loss but not server failures. Server failures are reported to the app, and the app decides how to handle them.
And, Homa's approach cuts the number of packets in half in the common case of short RPCs with fast service times. With acks it takes 4 packets for each RPC: one for the request packet, one for its ACK, one for the response packet, and one for its ACK. With Homa there are only 2 packets in the common case: one for the request and one for the response. The response implicitly acks the request, and the next request acks the previous response.
Hmmm, I see now that QUIC can delay ACKs and piggyback them on subsequent data packets, so it looks like it can get similar efficiencies to Homa.
I agree that Homa holds Tx resources for longer than would be needed in an ack-based scheme (I don't see that as particularly problematic).
And then in basically every other respect it is worse including compatibility.