Every person that works on kernel knows the overhead that comes with abstractions. Everybody that has worked with DPDK knows that you can get 20Mpps+ on a single core.
All they have done is to "frame" the usage and the term RPC differently ... i.e., it's all story telling and no real meat :)
Look at the previous set of publications by the same author: e.g., Achieve a Billion Requests Per Second Throughput on a Single Key-Value Store Server Platform, etc. they are all based on a single assumption that if you don't implement X and Y in your stack you can get better performance. Of course, you can. If you use your calculator to only compute 2+2 you might as well hardcode 4 as the output of your calculator.
They are not setting a lower bound. They are hardcoding and bypassing the parts they don't find useful.