In ~2010, I was benchmarking Solarflare (Xilinx/AMD now) cards and their OpenOnload kernel-bypass network stack. The results showed that two well-tuned systems could communicate faster (lower latency) than two CPU sockets within the same server that had to wait for the kernel to get involved (standard network stack). It was really illuminating and I started re-architecting based on that result.
Backing out some of that history... in ~2008, we started using FPGAs to handle specific network loads (US equity market data). It was exotic and a lot of work, but it significantly benefited that use case, both because of DMA to user-land and its filtering capabilities.
At that time our network was all 1 Gigabit. Soon thereafter, exchanges started offering 10G handoffs, so we upgraded our entire infrastructure to 10G cut-through switches (Arista) and 10G NICs (Myricom). This performed much better than the 1G FPGA and dramatically improved our entire infrastructure.
We then ported our market data feed handler's to Myricom's user-space network stack, because loads were continually increasing and the trading world was continually more competitive... and again we had a narrow solution (this time in software) to a challenging problem.
Then about a year later, Solarflare and it's kernel-compatible OpenOnload arrived and we could then apply the power of kernel bypass to our entire infrastructure.
After that, the industry returned to FPGAs again with 10G PHY and tons of space to put whole strategies... although I was never involved with that next generation of trading tech.
I personally stayed with OpenOnload for all sorts of workloads, growing to use it with containerization and web stacks (Redis, Nginx). Nowadays you can use OpenOnload with XDP; again a narrow technology grows to fit broad applicability.