Also no need to theorize: run the benchmark I linked for yourself. It clearly shows a massive advantage to having each thread work with its own directory.
Also no need to theorize: run the benchmark I linked for yourself. It clearly shows a massive advantage to having each thread work with its own directory.
this synchronization is handled for you by the fs, specifically the fs cache
inode alignment and errors are managed by this intermediating layer
your benchmarks are not demonstrating what you think they are demonstrating
the fs cache does most/all of the optimizations you're doing manually
bypassing the fs cache is highly atypical for user-space code
2. the overhead of modifying the dirent is statistically zero compared to the costs related to manipulating the files on disk
$ hyperfine --warmup 3 -N "./test /dev/shm 8 zip" "./test /dev/shm 8 chain" Benchmark 1: ./test /dev/shm 8 zip Time (mean ± σ): 118.5 ms ± 11.6 ms [User: 92.9 ms, System: 726.6 ms] Range (min … max): 103.6 ms … 143.4 ms 23 runs
Benchmark 2: ./test /dev/shm 8 chain Time (mean ± σ): 235.7 ms ± 11.0 ms [User: 116.4 ms, System: 1537.7 ms] Range (min … max): 220.1 ms … 258.3 ms 13 runs
Summary './test /dev/shm 8 zip' ran 1.99 ± 0.22 times faster than './test /dev/shm 8 chain'
i mean ignore me if you want, no skin off my back
but you're not benchmarking what you think you're benchmarking