You can try to re-vectorize the code for larger vector size.
51 karma · joined April 26, 2015
230 M/mm2 translates to 33nm "half-pitch".
Of course, transistors aren't square and aren't so densely packed, but these numbers are more real IMO.
Very odd - notice how the latency is high when the read IOPS are low. When the read IOPS climb, 95th percentile latency drops.
Looks like there is a constant rate of high latency requests, and when the read IOPS climb, that constant rate moves to a higher quantile. I'd inspect the raw results but they're quite big: https://github.com/scylladb/diskplorer/blob/master/latency-m...
Background: https://github.com/scylladb/diskplorer/
This may work for sequential reads, but not for random reads.
Note you'll need gcc 4.9 or later.