Benchmarking High Performance I/O with SSD for Cassandra on AWS
techblog.netflix.com
techblog.netflix.com
It's not a good idea to scale disk writes using threads when you can use aio. You've introduced unneeded scheduling and locking penalty from threads. Additionally, such threads will be accessing and modifying kernel structures frequently and that's going to cause additional locking inside the kernel. They'll all be modifying same resources and structures while writing.
Also, sixty threads on such a powerful machine is not really a big deal. Even though 0.06ms service time seems small, it is large enough such that you won't have more than a handful of threads competing for CPU resources (they'd waiting for some I/O completion)
Last I checked (at least a couple of years ago admittedly) glibcs POSIX aio_* implementation was implemented using threads despite kernel support for it.
TL;DR
What follows is a more detailed explanation of the benchmark configuration and results. TL;DR is short for "too long; don't read". If you get all the way to the end and understand it, you get a prize...
I believe it's "Too Long; Didn't Read"