PacketShader: GPU Accelerated Software Router
shader.kaist.edu
shader.kaist.edu
On the other hand, this research is limited by PCI-E bandwidth (quote "performance of our system is limited by the dual-IOH") so CPU/GPU does not seem to be bottleneck in most of the tests.
One thing that puzzles me about this design, they use it to boost the memory bandwidth in NUMA systems but isn't that more usually done using switches or special cards with crossbars or point-to-point links rather than routers?