Visualizing AWS Storage with Real-Time Latency Spectrograms
sysdigcloud.com
sysdigcloud.com
I'm surprised that there's no asynchronous way that the FS cache will flush itself i.e. when it reaches 50% capacity, and rate-limit incoming requests if it's too full. The idea that an FS cache is so dumb that it can't do anything while it's flushing its entire self is a bit scary - I'd expect that circular buffers and granular locking mechanisms could be used to great effect here. Is this kernel code? Userspace code? Is there research into this? Fundamental tradeoffs that I'm missing?
Red implies problems, green implies "normality", but here this association is misplaced. Perhaps a typical "fire" palette would be better - from dark brown to red to orange to yellow and, ultimately, to white for the extremes.
In the meantime, it's very easy to tune the colors your own: just modify this line https://github.com/draios/sysdig/blob/master/userspace/sysdi... in your local version of the script, using this as a reference http://misc.flogisoft.com/_media/bash/colors_format/256_colo....
I believe the issue raised isnt the palette range itself, but rather that it is the reverse of what it is typically expected. The current red area "should" be green indicating there are many calls in the fast region while the current trailing green blocks "should" be red indicating problem issues
This color of green=good and red=bad I believe stems from Triage tags: http://en.wikipedia.org/wiki/Triage_tag
Sometimes white is used below green as 'dismiss/not an issue'
[1] http://dtrace.org/blogs/bmc/2013/11/10/agghist-aggzoom-and-a...
As for SSD vs Magnetic EBS, I can't say that I'm surprised. I'd assume that EBS implements some sort of cache in between you and your actual disk on the other side of the network so that the writes can return even faster. Try doing this again with reads and I'd bet you'd get some interesting results.
Edit: Also, did you pre-warm your EBS volumes? http://docs.aws.amazon.com/AWSEC2/latest/UserGuide/ebs-prewa...
And yes, there are several interesting workloads that I didn't test, including read only and read+write. It's potential material for another blog post.