Brubeck – A statsd-compatible metrics aggregator
githubengineering.com
githubengineering.com
There are some details of how Etsy uses statsd that are not well-communicated. Etsy samples metrics aggressively to limit the amount of total traffic. And they monitor the packet error rates on the statsd boxes like hawks to keep the loss rate in check. Back when I was working there, if you added a high-volume counter without sampling it, the alarms would sound and you'd have an ops person tapping your shoulder pretty quickly. If you use statsd and skip either of these steps, the 40% loss that github experienced is what you get.
AIUI Etsy's moved to a consistent hashing scheme that's at least vaguely similar to this.
Node was not then, nor is now Etsy's area of expertise. We were going through an adolescent "let's just use every language" phase when we built statsd. I think the problems outlined here are solid supporting evidence that you should use a smallish set of tools and master them (a point of view which is very on-brand for Etsy engineering as it exists today).
Having whatever library is sending data to statsd, instead keep counters/gauges in memory and then expose that on a regular basis would greatly reduce the data volumes involved as it's O(timeseries*frequency) rather than O(events).
This is the approach we take with Prometheus, and based on the statsd setups of some people who've come talking to us there's scope for a reduction in network load of at least an order of magnitude without having to do any downsampling.
The statsd design choices here are mostly explained by the fact that Etsy uses it to collect from PHP. PHP doesn't afford a great way to aggregate in the client. (These are design choices that serve PHP well systemically, although it's limiting here.)
You're getting into IPC then, which is a fun topic alright e.g. https://github.com/prometheus/client_ruby/issues/9 and https://github.com/prometheus/client_python/issues/30
http://stackoverflow.com/questions/12871642/scaling-statsd-w...
(the first answer)
I am really happy to see another project, one that will probably reach more people than my web framework, honor the same wonderful musician I was trying to honor.
Proud to say that it all happened in RC's inaugural batch too. :)
"After three years running in our production servers (with very good results both in performance and reliability), today we're finally releasing Brubeck as open-source."
* Brubeck only runs on Linux. It won't even build on Mac OS X.
…
Brubeck has the following dependencies:
* A Turing-complete computing device running a modern version of the Linux kernel
…
The are several ways to interact with a running Brubeck daemon.
Though I guess that jazz guy is more famous.Any plans for pluggable backends in addition to BRUBECK_BACKEND_CARBON? Pls consider making it agnostic, i.e. using a wire protocol.
Am I think about this the right way? Would channels have helped handle some of the load effectively?
Any gophers care to comment? :D
"I was young(er) and stupid(er) then.."
> We like C. We're going to need a few more years of
> research, real production experience, and
> language/library maturity before betting critical
> infrastructure on something else by default.
That's a really bizarre opinion to hold in the year 2015. Don't get me wrong: I love C, too. But certainly not for home-grown infrastructure at a web company — the incentives just don't align. And hiding behind the "we just don't know" bugbear doesn't parse, either. Go, for example, powers gargantuan-scale infrastructure at Google, and has for half a decade. And there's whole fleets of organizations as big or bigger than GitHub that report the same experience.Note that he is not saying they're not open to alternatives, nor that they're not experimenting with alternatives, but that they need more experience first before "betting critical infrastructure on something else by default".
> Google also has a gargantuan-scale dev team that
> includes the people behind Go. It's ridiculous to
> compare.
It's not ridiculous. There are probably reasonable arguments against using Go for things like this, but "insufficient developer capacity" isn't one of them. Becoming a Go expert is a task measured in weeks.One of the first things people abandon when scaling golang servers is the channel abstraction and I expect the same would be true in this case.