Show HN: Pup, real-time app metrics with Statsd
datadoghq.com
datadoghq.com
I solved the same problem in another way with:
* collectd
* collectd to graphite plugin (for server-level metrics, from all your servers)
* graphite
* statsd (for custom app metrics over UDP)
* graphene - my own UI toolkit for graphite, also based on D3.js (http://github.com/jondot/graphene)
All in all, I got a better, more robust and detailed solution using these components. Using Chef, the time to set it up is very minimal.
Pup has a narrower scope than Graphite (which we also use and love), but optimizes for shortest time to app metrics viz.
(Oh and Datadog itself is all free up to 5 hosts - but remains completely optional)
I'm not sure I understand what "shortest time to app metrics viz" means. Are you referring to the time Pup takes to create charts? Or the lowest time frequency (minutes, seconds etc.) it can handle?
Graphite defaults to reporting at 1-minute intervals, and can report at a finer granularity (1 sec. intervals) if you tweak a few settings and set up a cluster. I'm just wondering how Pup differs in this respect.
By default, pup aggregates statsd data in 10s intervals, and the UI refreshes it continuously.
Still, it's pretty good that there is an alternative frontend, if nothing else but to expand the graphite ecosystem.
[1] https://github.com/obfuscurity/descartes
[2] https://github.com/paperlesspost/graphitihttp://graphite.readthedocs.org/en/0.9.10/carbon-daemons.htm...
If you're watching a number go up, use a statsd.Counter. Want to track request arrival rates? Use statsd.Gauges. Want to record times for request fulfillment? Use statsd.Timers.
Statsd isn't necessarily better or worse than carbon-cache via UDP, but it provides a handy solution for the above use cases.
To build custom UIs they wrote javascript libraries Crossfilter (http://square.github.com/crossfilter/) & Cubism (http://square.github.com/cubism/)
If that doesn't cut if for you, we're happy to help - let us know what you're trying to do!
Pup's feature-set is limited today, though, compared to graphite. You may also check-out the full http://datadoghq.com service for more. It's SaaS, and free up to 5 servers.
Does this always need to run on the same server as the Application? Or is it possible to run it on a different server and push the data to it?
disclaimer: i have no experience with statsd/graphite etc.
- Our goal with pup was to make it super-easy to see app metrics. It is therefore much narrower in scope than services like New Relic or projects like Graphite. It's open-source, though, and you can take it anywhere you want.
- We ourselves operate a service that can consume data from Pup and other sources, and provides metrics + events aggregation and correlation, fancy graphing, alerting, etc.. You can check-it out at http://datadoghq.com
- DogStatsD is running on my Server and is collecting data. - I can use "dogstatsd-ruby" to collect data from my (i.e.) Rails App. - DogStatsD reports Stats to your Service Datadog HQ (optional) - Pup is a small version of Datadog HQ that is running locally on my Server and connects directly to DogStatsD. - If i feel Pup doen't satisfy my needs, i can switch from Pup to Datadog HQ, get more functions and you my money? :)
right?
We're on a mission to provide monitoring that doesn't suck, and we believe that making it easy and rewarding to instrument your app is an important step on the way.
Once you get addicted to metrics and want more aggregation / graphing / alerting / analysis capabilities, there's a number of open-source components you can pipe your statsd data into. Or you can use our own http://datadoghq.com service for that.
To my knowledge nobody has really solved this problem yet in a satisfactory way, outside some closed-source solutions at big internet companies.
Our service isn't free beyond 5 hosts, but is quite a bit cheaper than rolling and running your own or flying blind and facing the consquences. It's a hard problem, and we're on a mission to solve it for companies that don't start with goo* or end in *ook.
Without some understanding of the dimensions of the data, it is very difficult to compose dashboards or aggregation rules that have arbitrary filters or can update automatically when new components are added.
I really ought to elaborate in a blog post someday :)
You can attach any arbitrary set of tags to metrics or events - on a per-datapoint basis, and slice / dice / alert based on those tags. Datadog will automatically tag your points by chef role or AWS availability-zone, for example, and you can add any other tag you want. Tags also don't have to be tied to a host and can also relate to a specific volume, mysql index, etc...
(Note to HN, this is a feature of datadoghq - pup will gladly collect and filter on tags, but won't aggregate them, yet)
Suppose I apply the tags "ord", "foo.example.com" and "bar.example.org" to a metric. How does the dashboard builder or the aggregator know what they are? They're values without the fields.
And you can then ask for "my.metric" summed over "availability-zone:us-east-1a" grouped by "role".
Here's a screenshot to show what it looks like https://skitch.com/oliveur/emywc/sobotka-datadog
You obviously have a lot of experience with monitoring tools - I'd love to connect. I'm @oliveur on the twitters.
Don't be shy if you have questions when running pup.
So, no logs to parse to get any app metric, everything already graphed for you in real-time, quasi-nil setup.
We chose statsd because it means that you don't have to change your code to get it to production. You can then rely on pup for display (but no historical data), or on your internal tool chain (e.g. graphite et al.) or give our service (http://www.datadoghq.com) a try if you don't want to set your own stuff up.