Multipathing, delay based congestion control, predictable performance under load (even with multi-tenancy!), builtin encryption, loss recovery while maintaining in order behavior even when doing IB Verbs, modular pluggable upper layers to allow apps to use NVMe and RDMA directly, support for smooth operations even with >100k connections, separation of course based management and data flow for rapid iteration / programmable congestion control, ordered and unordered processing, connection type awareness to allow further optimization / taking advantage of unordered wins.
The results sections speaks clearly to their wins. A tiny drop rate or pre-order rate causes enormous goodput losses for RoCE. Falcon hardly notices. P99 latencies go to 7x slowdown over ideal at a 500 connections for RoCE versus 1.5x slowdown for Falcon. Falcon recovers from disruption way faster, whereas RoCE struggles to find goodput. Multipath shows major wins for effective load.
If you don't care about p99, and you don't have many connections, yeah, what we have today is fine. But this is divinely awesome stability when doing absurdly brutal things to the connection. And it all works in hostile/contested multi-tenancy environments. And it exposes friendly fast APIs like NVMe or RDMA directly (covering much of the space that AWS Nitro does).
It is also wildly amazing how they implemented this. Not as its own NIC but by surrounding a Intel E2100 200gbps NIC with their own ASIC.
The related works section makes comparisons versus other best of breed and emerging systems. Worth reading that to see more of what wins were had here.
The NIC's job these days is to keep very expensive very hot accelerators screaming along at their jobs. Wasting with congestion and retransmits and growing p99 latencies costs incredible wastes of capacity, time, and power. Falcon radically improves the ability of traffic to make successful transit. That keeps those many kilowatts of accelerators on track with the data and results they need to keep crunching.