A Microscope on Microservices
techblog.netflix.com
techblog.netflix.com
http://en.wikipedia.org/wiki/Little%27s_law
You have a request rate in requests per chronometer-second, and a response time in stopwatch-seconds per request, and you multiply them to get a "demand" figure in stopwatch-seconds per chronometer-second. It's sort of dimensionless, and sort of not, because the seconds on either side of the division operator are sort of orthogonal (very vaguely like how joules per newton-metre is not dimensionless).
How do i use a number like this? Does it make sense to compare the numbers from two different instances of the same system? From instances of two different systems? Should i worry if it goes up? If it goes down? What can i do about it, either way? Is it meaningful to calculate it for component parts of my system, and is there a way to critically relate the values in parts to the whole? Is there a way to relate it to other quantities in my system?
The flamegraph code is on github (https://github.com/brendangregg/FlameGraph). There's other implementations too (see http://www.brendangregg.com/flamegraphs.html#Updates).
We're using them primarily to analyze CPU usage of the Linux and FreeBSD kernels, Java, and Node.js. We had an earlier post about the Node.js ones: http://techblog.netflix.com/2014/11/nodejs-in-flames.html
Do we have any idea how massive Netflix's scale is, in terms of end-user requests per second, or some other metric?
And, probably more relevantly for me, how big a scale can one get to while using most monitoring and analysis tools?
I used carbon-relay in one job, and if you're using that, i'd guess you have 1000 machines serving users, and 30000 collecting metrics!