For many of the things we work on, this is not just slow, it is atrociously slow. For analysis of real-time systems I want to put many data tap points throughout the processing chain. If I add 50 (not uncommon), at 5ns a pop most systems (again, IME) are not going to care about 0.25 microsecond extra latency. At 600ns I add 30 microseconds to their processing chain and will likely have to spend some of my time explaining to the owners why my tap points will not affect the system I am helping fix.
It's all relative. It may sound fast but on the other hand it's less than two million times a second. There are many cases where spending an extra 600ns in a tight inner loop would absolutely obliterate a program's performance, whereas an implementation taking just 3ns would not.
Compared to 4 ns on Windows, many calls to the method can create significantly different performance profiles across the platforms.
if you do e.g. HFT, additional 600ns might kill you