Using https://github.com/c-blake/batch/blob/master/examples/closes... from the examples/ directory of https://github.com/c-blake/batch, I get:
L2:bat/examples$ BATCH_EMUL=1 ./closes|awk '{print $3,$2,$1*$2}'|fitl -s0 -c,=,n,b -b100 1 2 3
$1= 3.815 *$2 + 695.6826 *$3
bootstrap-stderr-corr matrix
0.8872 -0.7856
0.04174
L2:bat/examples$ BATCH_EMUL=1 ./closes|awk '{print $3,$2,$1*$2}'|fitl -s0 -c,=,n,b -b100 1 2 3
$1= 4.528 *$2 + 695.3265 *$3
bootstrap-stderr-corr matrix
1.174 -0.7640
0.05302
L2:bat/examples$ a (695.3265 +- 0.05302)-(695.6826 +- 0.04174)
-0.356 +- 0.067 # i.e. run-to-run consistent
L2:bat/examples$ ./closes|awk '{print $3,$2,$1*$2}'|fitl -s0 -c,=,n,b -b100 1 2 3
$1= 779.67 *$2 + 48.861 *$3
bootstrap-stderr-corr matrix
1.550 -0.7495
0.2023
L2:bat/examples$ ./closes|awk '{print $3,$2,$1*$2}'|fitl -s0 -c,=,n,b -b100 1 2 3
$1= 787.77 *$2 + 48.745 *$3
bootstrap-stderr-corr matrix
1.350 -0.6207
0.1635
L2:bat/examples$ a (48.861 +- 0.2023)-(48.745 +- 0.1635)
0.12 +- 0.26 # i.e. run-to-run consistent
L2:bat/examples$ a (695.3265 +- 0.05302)/(48.745 +- 0.1635)
14.265 +- 0.048 # i.e. batching makes it 14X faster!
One can go further to get a mean +- std.err of ~695.52 +- ~0.04 nanosec hot-cache time per syscall overhead or even do point-weighting there, but it's probably more important to check fit residuals for serial autocorrelation (very do-able with https://github.com/c-blake/fitl, EDIT: oh, and if you care `a` is basically shown in https://github.com/SciNim/Measuremancer/pull/12 which uses standard error propagation).The bigger point would not be ever more careful measurement of the cost if you do not, but rather the ease of "batchifying" if you do of any highly regular call interface similar to syscalls with high latency and also the "follow on" API design, granularity-wise. The inner code of a mini-assembly language letting you implement, say, open/fstat/mmap/close all in one syscall crossing is only about 30 lines of C. { Yes, yes..It could use multi-CPU-arch and syscall auditing integration.. and sure, ebpf & io_uring alter the Linux landscape these days }. Also, @exDM69 is very correct that pure hot cache-hot loop numbers are only one part of the cost story (https://news.ycombinator.com/item?id=39188551).