Grafana 4.0 with alerting is released
grafana.org
grafana.org
Oh, and if your in New York tomorrow, signup for GrafanaCon: http://grafanacon.org
We struggled a bit with whether or not it really "belonged" in Grafana, but we believe in alerting while "in the flow".
It makes a lot of sense (from an experience standpoint) to 'manage' alerting while you're 'managing' your dashboards, visualizations, and queries; you already have a sense of the data _right there_.
Collectd, telegraf. etc, can be configured to send the same metrics to your favorite TSDB and Alerting system (like riemann) in parallel.
This means that the lowest value for the last 5min of the serie have to be above 80% before the alert triggers.
With alerting in place now, I'm even happier than ever. A huge thank you to the Grafana team for solving a huge pain point!
(Haven't tried telegraf yet, setuping a prometheus at the moment)
So, I meant encryption in transport, authentication, etc. as many solutions work well if you're monitoring "in the clear" from the backend, but not so much over the internet.
I love everything about using Influx but it would die and never restart and every time it would some crash on semacquire. I'll have to try it again since I need to check out this Grafana update anyway.
The two game changers are using the UDP line protocol instead of HTTP, and making sure you are batch-processing inputs. Fixing these settings is the difference between an instances that crashes all the time, and a purring one.
Shameless plug - I recently published a log router in Golang. It sends data to influx too ! (github.com/agnivade/funnel)
I checked my graphite setup once. We had 27% of metrics lost over UDP. That was bad.
pro-tip: "netstat -anus" and look at the error counters.
- statsd
- collectd
- graphite
- whisper
- carbon
- prometheus
- grafana
- seyren
- riemann
- nagios
- icinga
- zabbix
There are multiple modern SaaS software that will do all of that in a single tool with better integrations, more polish, less work and no maintenance.
1) See https://www.datadoghq.com and last news https://techcrunch.com/2016/01/12/investors-feed-datadog-a-h...
2) https://signalfx.com/ and last news https://techcrunch.com/2015/03/12/signalfx-emerges-from-stea...
3) http://www.bmcsoftware.uk/it-solutions/truesight.html if you're not anti entreprisey (that was the "Boundary" startup, bought by BMC a few years ago and integrated in their offerings).
And don't think that they are "new" fancy tools. They've been around for many years.
The OSS tools costs a fortune in human to maintain them, and another fortune in hardware to run it.
(I picked 600 because that was the approximate number of machines we had at my last job, where we used Graphite maintained by one guy, part time).
You included a LOT of redundancy in your OSS list. Multiple timeseries databases. Multiple collection daemons. Multiple dashboards. Multiple alerting systems (Who in their right mind would use Nagios AND Icinga?). You're effectively arguing about maintaining multiple monitoring stacks, some of which are quited aged.
Let's say statsd + collectd (metrics collection) + graphite (aggregation) + carbon/whisper (graphite storage) + icinga (alerting) + grafana (graphing). That doesn't exactly come easy.
No offense but a single graphite is not a monitoring solution. It's just the tip of the iceberg. Monitoring does take a lot of engineering work and a lot of maintenance. You won't get away operating 600 hosts on the cheap, just think about how much are the hosts themselves.
Note: I am one of the maintainers of Diamond, a metrics collection tool written in python. https://github.com/python-diamond/Diamond
A typical system doesn't use all of the tools above. You use what fits you and many of the tools play pretty well together. I've had luck with Icinga2 and Grafana lately, for example, which integrated quite smoothly out of the box.
What you call a 'clusterfuck' is really a wider ecosystem. It would be pretty crazy for a single organization to use all or even most of the tools that you list.
Right now, people accept high degrees of cost (especially for at scale users) and lock-in, in exchange for the convenience of SaaS. Or, they go open source (which to your point, certainly is an investment in time)
Watch out for what team Grafana will be doing in 2017. Our plan is to provide a fully turnkey, hosted offering based around Grafana (and a handful of other open source tools). OpenSaaS.
We hope that for many users, this can be a third choice, and in some ways the best of both worlds.
Having nice graphs is nice... until they fall apart because the source is unavailable.
And that doesn't help with alerts either. (I tested the alerts in the v4 beta, it's just not comparable to the better alerting tools out there).
The alerting in v4.0 is just the beginning. Torkel and the team have tried to optimize for the “relatively simple" 80% of alert use cases.
We are fans of other, more sophisticated open source alerting tools like Bosun, and you can be sure that we'll be both improving our alerting capabilities in 4.x
- customization will be very expensive, if not impossible
- you must have people for the procurement process (x10 more costly if you are in a gov agency),
- weird failures due to not finding the license,
- your cheap personal that install software won't be able to do it,
- you'll have problems creating testing environments because you don't have licenses
- you won't be able to do some things immediately because there aren't enough licenses.
And these are just the problems that came to my mind right now. All of them are real problems that I'd found in commercial software.
You've got a full API and integrations with a hundred different tools and services out of the box.
Seriously, my coworker was skeptic at first too (so was I). Then we configured the full integrations with AWS/the-agent/statsd/postgre/mysql/cassandra/elasticsearch/riak/nginx/haproxy/redis/memcache/pagerduty/slack and some more.
My co-worker concluded in front of my CEO, "it was 2 orders of magnitude faster [than anything else we've ever tried for monitoring]". And that's not even talking about the additional features and customization we couldn't even dream of.
> - you must have people for the procurement process (x10 more costly if you are in a gov agency),
True. That's the only major problem I can see: People who can't buy the software they need. That's a social problem, not a software problem.
> - weird failures due to not finding the license
It's only one API key to put in the agent config file.
> your cheap personal that install software won't be able to do it
I don't know who you're talking about. Monitoring has our best people working on it. At other places I've seen, it's done by devops consultants raking up £600 a day.
There is no cheap personal involved. (Maybe you're thinking about of cheap interns who add alerts? that's an anti pattern).
- you'll have problems creating testing environments because you don't have licenses
Same license. Put a tag environment=<environment> in the config and done, all metrics all servers and all alerts will be tagged.
- you won't be able to do some things immediately because there aren't enough licenses.
Not applicable. It's not a limited license by seats.
You pay the bill at the end of the month depending on the number of hosts in your package. There is a hourly price for ephemeral hosts and overrun.
I think having one pane of glass to do all passive monitoring tasks is an incredible step forward.
I am yet to see if the Active monitoring of Grafana is any good, but it does look very promising
You get all of this for free with Crate.io without giving up the flexibility of a general purpose SQL database...
Wanna start storing log data in crate as well? No problem! Just design your table schema, and API ingest layer (My favorite is NodeJS) but you can use any language you like.
Or if security (facing the public) isn't an issue (if you're on a subnet safe from the public internet) then you can certainly just use the built-in REST API which crate exposes.
With Crate, I've been able to store hundreds of GB of systems log data without worrying about silly things like table-bloat (the autosharding of partitioned tables handles the spectre of bloated table shards for me for free).
Thanks to the amazing developers over at Crate.io for taking the best of Elasticsearch and making it sane, fast, and chock-ful of SQL goodness!
Also a big thank you to the Grafana team for recognizing the potential synergies that Crate.io & Grafana could catalyse for unifying time-series & log data streams.
Is it correct that Grafana works best with Graphite? At least that seems to be my impression, and it is a bit sad, since I think Graphite is cool, but it really has a lot of moving parts.
They all worked great. I can't think of any reason to use Graphite.
We also use it against mysql. Annoyingly there is no driver for that so we had to build a nasty layer than translates from influx to mysql.
Very glad I saw this announcement on HN.
What have other people had success with?
I think Grafana will fill the basic GUI alerting needs, though. When you need more than a simple flat treshold you usually want to get out of the GUI and ask the ops team for help anyway.
And using a single paid tool that does the job better AND doesn't kill me in maintenance work.
See https://www.datadoghq.com/ as leader or https://signalfx.com/ as the second comer, or http://www.bmcsoftware.uk/it-solutions/truesight.html if you're enterprisey.
$15/month/host gets expensive fast. Datadog doesn't start providing discounts till you are at 1000+ hosts.
$15 * 500 hosts = $7500 per month.
If you think it's expensive, I can only advise you to check how much the hardware will costs on EC2 to run the free tools, plus how much work it will take to get the 8 different and independent OSS tools to work not only alone but integrate together, plus how much additional work and maintenance to keep it working without hiccups (war story: there is nothing worse than a monitoring tool that is less reliable than the thing it monitors).
Oh that's ridiculous pricing. Any team running a server installation at any sort of scale would scoff at the pricing, and that's even before you take into account the implications of sending so much of your metrics data to a third party to be held hostage there if you decide to leave the service.
You can use Splunk if you have money. That's the de facto standard. Beware that it's one of the most expensive software license on the planet :D
Here's a short howto + video: https://sematext.com/blog/2015/12/14/using-grafana-with-elas...
(Writing those "ALERT ..." requires a steep learning curve.)
> I repeat: Your alerts and dashboards belong into your SCM, not a random SQL database!
(And I 100% agree, particularly for alerts)