1. Our throughput benchmark is designed to measure data transfer bandwidth, and the comparison point is RDMA writes. In the benchmark, eRPC at the server internally re-assembles request UDP frames into a buffer that is handed to the server in the request handler. The request handler does not re-touch this buffer, similarly to RDMA writes.
We haven't compared against RPC libraries that use fast userspace TCP. Userspace TCP is known to be a fair bit slower than RDMA, whereas eRPC aims for performance RDMA-like performance.
2. eRPC uses congestion control protocols (Timely or DCQCN) that have been deployed at large scale. The assumption is that other applications are also using some congestion control to keep switch queueing low, but we haven't tested co-existence with TCP yet.
3. 75 Gbps is achieved with one core, so there's no need to distribute load. We could insert this data into an in-memory key-value store, or persist it to NVM, and still get several tens of Gbps. The performance depends on computation-communication ratio, and we have tons of communication-intensive application
4. Packet loss in real datacenters is rare, and we can make it rarer with BDP flow control. Congestion control kicks in during packet loss, so we don't flood the network. Our packet loss experiment uses one connection, which is the worst case. An eRPC endpoint likely participates in many connections, most of which are uncongested.