According to some benchmarks I've seen (sorry no link), a system call in the middle of a memory heavy inner loop can impact performance for 100 us before performance is back to steady state. This is 100x longer than the system call alone. Of course there is work being done, but at a reduced throughput. This is in line with my practical experience from working with performance sensitive code.
These things are difficult to benchmark reliably and draw actionable conclusions from the results.
This is not a criticism to the author of this article, the article clearly describes the methodology of the benchmarks and does not suggest any wrong conclusions from the data.