How does that work?
How does that work?
If you emit metrics as necessary to a time series database, then you should be able to build alerting based on the time series metrics. Your monitoring systems should be good at building alerts based on a stream of metrics and visualizing the time series data.
Sometimes you might have to look at the visualizations to find something, but ideally you then set up an alert on the thing you looked at so you have the alert for the next time it happens. A great monitoring system lets you turn graphs into alerts right in the interface, so if you're looking at a useful graph you can make an alert out of it.
Sometimes logs can be useful, but only after your monitoring system has told you which system is not behaving, and then you can turn on logs for that system until you've solved the problem, but you shouldn't need access to old logs, because if the problem was only in the past, then it's not really a problem anymore, right? If you have an ongoing problem, then maybe have the logs on for that service while you're investing that problem, but then turn them off again.
But having a ton of logs always generating and being stored tend to be fairly useless in practice with a good time series database at hand.
Also the cost of the infrastructure to search the logs and view the logs.
> Sometimes logs can be useful, but only after your monitoring system has told you which system is not behaving, and then you can turn on logs for that system until you've solved the problem, but you shouldn't need access to old logs, because if the problem was only in the past, then it's not really a problem anymore, right?
Some things happen rarely, but can still have large impact. E.g. Imagine a once a day job of moving files which fails twice a month, rendering those files inaccessible.
Likewise, turning logs on only after you've seen a problem means you miss out on troubleshooting the root cause of it - if there was a spike of badness this morning but you don't have logs for it, you're missing out on diagnostic information that may have protected you from repeats of that spike in future.
I've also had business guys want to analyse things like access logs in ways that they didn't know previously. Logs provide a datastore of historical activity, which in smaller shops is a cheap data lake.
Perhaps the 'no logs' thing works for your setup, but I think it's bad general advice. And your position is not that logs are useless ("turn on logs for that system until you've solved the problem"), but that retaining logs are useless - quite a significant difference between the two.
That's an important distinction, one that I agree with, and I should make clearer.
Logs do have a purpose, but I'm not sure that retaining them does.
Sure, for a very small shop, throw them on a disk, use awk, sed, grep, and perl to look through them, and call it a day. But once you get to the point of "spinning up a cluster of log servers" or something like it, I'd say you're probably better off investing in monitoring instead.
More than once I've run the entire corpus of requests to a system, ever, through a dummy rebuild as a pretty great integration test. It's a powerful SRE tool. Spelunking through all historical data is just icing on that cake, honestly. As the author says, SRE is basically just an information factory; I'd be haaaaaard pressed to agree with you on throwing away a lot of information -- you don't know what you don't know until you want to know it -- and betting all-in on monitoring. Retaining logs is not the hardest problem SRE deals with, either, but SREs turn around and force unrealistic latency requirements on the query side (I see a lot of ELK deploys running into this).
You have to look at it as an Oracle. Oh, great Oracle of a pile of meaningless logs, cook off this map/reduce and tell me an interesting number that I can put in a Keynote for executives. Definitely not dashboarding from logs data. That's an impedance mismatch that Google gets away with because of the nature of their logging.
Nonetheless, splunk is the most expensive software license on the planet. More than Oracle, yes.
Interestingly, we never used a cluster of log servers. I was always skeptical of their utility. It was grep, plus some hand-rolled utility scripts to interrogate. One was a thing of beauty I spent years on, which saved us a ton of time.
A monitoring system has a lower barrier to entry.
http://datadoghq.com/ => will do ALL of that and much more. You can deploy it in a few hours to thousands of hosts, no problem.
Direct competitor: http://signalfx.com/
Have no money to pay for high quality tool? graphite + statsd will do the trick for basic infrastructure. However it's single host, doesn't scale and only basic ugly graphs are supported.
Similarly, basic infrastructure health is not giving you the same sort of information (what the software is actually doing) that logging does. In order to do time-series monitoring of your software rather than your system, you need to spend time thinking about what metrics you need to track and how you're going to obtain them.
I run both an ELK stack and a Prometheus stack, and I find they're good for different things.
Sumologic, papertrail, logentries (and many more) for cloud logs. Graylog or ELK or Splunk for self hosted logs.
However, logs should never be send to the cloud, they are too sensitive information to outsource. Server metrics + stats are more reasonable.
Agree, stats and logs cover different things. Need both.
Saying that logs should never be sent to the cloud is overly absolute. Some logs should indeed stay behind firewall, but lots of organizations have logs that can be shipped out to services whose features derive all kinds of interesting insights from logs.
Since you mention Papertrail specifically in the context of costs - Papertrail is actually a bit pricey relative to the competition. For example, compare https://papertrailapp.com/plans to https://sematext.com/logsene/#plans-and-pricing . I think Sematext Logsene is 2-3 times cheaper than Papertrail.
Lastly, I was at a Cloud Native Conference in Berlin last week. A lot of people have the same setup as you - ELK for logs + Prometheus for metrics. We're running Sematext Cloud where we ship both our metrics and our logs, so we can easily switch between metrics and logs much more easily, correlate, and troubleshoot faster. Seems a bit simpler than ELK+Prometheus...
That's what Grafana [0][1] is for -- i.e. creating nicer displays for Graphite.
However it's single host, doesn't scale
It may take some effort, but it can be done, and much of the heavy-lifting seems to have been done and been made available as open-source.
Here's a blog post from Jan. 2017 [2] from a gambling site about scaling Graphite.
And here's a talk [3] from Vladimir Smirnov at Booking.com from Feb. 2017 about scaling Graphite -- their solution is open-source (links in the talk and slides available at the link):
This is our story of the challenges we’ve faced at Booking.com and how we made our Graphite system handle millions of metrics per second.
(And this [4] is an older, but more comprehensive, look at various approaches to scaling Graphite from the Wikimedia people with the pros and cons listed).
[1] https://github.com/grafana/grafana
[2] http://engineering.skybettingandgaming.com/2017/01/13/graphi...
[3] https://fosdem.org/2017/schedule/event/graphite_at_scale/
It can be done but at what costs? Better get a tool that gets the job done out of the box and does it well.
For starters, if you're operating in the cloud, you cannot get servers with FusionIO drives and top notch SSD. That limits your ability to scale vertically.
Discalimer: I work at Sumo Logic and enjoy it.