Hardware Acceleration of Key-Value Stores [pdf]
zhehaomao.com
zhehaomao.com
I think this paper vastly underestimates memory constraints of higher performance systems.
From a quick skim, the Xilinx work synthesises the entire network stack (including TCP) in hardware, unlike the above study, which only supports UDP in the HW traffic manager.
The numbers in the Xilinx paper are more attractive than what this study found and the paper also includes power measurements (since Joule/request is one metric on which FPGAs and hardware do quite well compared to software).
That said, much of the latency gain likely comes from bypassing the OS kernel and its generalised network stack. There is plenty of existing work (unfortunately also not referenced in this paper) that does this and which achieves very low latency, albeit -- to be fair -- on x86 hardware. (Examples: Arrakis [1], IX [2] and MICA [3].)
[1] -- https://www.usenix.org/conference/osdi14/technical-sessions/...
[2] -- https://www.usenix.org/conference/osdi14/technical-sessions/...
[3] -- https://www.usenix.org/conference/nsdi14/technical-sessions/...
Our initial evaluation with a realistic workload shows
a 10x improvement in latency for 40% of requests without
adding significant overhead to the remaining requests.
And then I re-read this claim, did a little math, and realized that they only reduced it by 36%."10x for 40% of requests" is a skeezy way of saying 36%.
I don't intend to criticize this paper in particular, but, generally, I don't see small performance improvements in such software to be very useful for society. Academia just becomes a research arm of corporations that might even be a net negative for society: eroding privacy rights (Facebook et al) or introducing volatility into stock markets (HFT could use this paper's insight just as fruitfully.)
Less than 400ns per switch hop (http://www.arista.com/en/products/7150-series) excluding congestion.