Except they are and your claims are trivial to disprove: simply run the benchmarks under perf. You'll find that most of the time is spent on the rwsem which is described here onwards: https://www.kernel.org/doc/html/latest/filesystems/path-look...
the fs cache does most/all of the optimizations you're doing manually
bypassing the fs cache is highly atypical for user-space code