Guide to Logging
loggly.com
loggly.com
But don't aggregate every code exception and warning. If they aren't affecting your customer experience, who cares? And if they are affecting it, then it should be expressed as an application metric in your monitoring system.
Edit: To make my point a little clearer, I believe that central logging and aggregating should be available, just not turned on all the time.
The only way to find many issues is by looking at logs of unexpected (and especially unhandled) exceptions.
If you have so many of those that they're impossible to dig through, FIX YOUR CODE.
Keep in mind I consider the existence of exceptions to be an application metric that should be logged, so if there is a security issue causing exceptions that should show up in the monitoring, and then you can look at the exceptions that happen going forward.
If the box is compromised, your local IDS should catch that (assuming you are allowing writes on the box at all)
I will concede if you have a perfectly configured all-knowing security oracle in front of your application then you don't need proper logging. :)
immutable infrastructure for the fronted so an attacker can't do much damage if they do get in.
I've seen attackers get access then just dump the site databases or copy all the site's code sitting on the compromised servers. Immutable infrastructure doesn't protect against reading things over the network.
The class of security problems is larger and less clearly defined than what can be pre-programmed into static monitoring or analysis tools. It's always good to store logs for after-the-fact forensics if something does go wrong. How often do we see attacks, but then the attacked company doesn't even know what data was compromised or exfiltrated due to lack of logging? It happens daily.
If you do happen to have a magic all-knowing security oracle, you should probably productize it and make a trillion dollars.
Heh, that's not quite was I saying. :) I was just saying that any analysis you'll do on application logs you can also do with an application firewall.
> How often do we see attacks, but then the attacked company doesn't even know what data was compromised or exfiltrated due to lack of logging? It happens daily.
Again, I would contend that you won't find this information in application logs anyway. You'd find it in your IDS logs that are monitoring outbound network traffic.
Anecdote: I've previously caught an in-progress exploit by seeing mysql errors logged from the application because the exploits were doing dumb things like SELECT * against a table with tens or hundreds of millions of rows. Sometimes the little things let you know.
IDS is nice in theory, but it's not really de rigueur for non-specialized platforms these days. It's about as practical as requesting companies keep complete netflow logs: great in theory, but almost nobody does it.
If these security approaches were deployable in a complete drop-in fashion, we could push for industry wide adoption, but right now everything is custom tailored to individual architectures and environments. ain't nobody got time for that when there's hustlin' to be done and worse is better and perfect is the enemy of good and fail fast and be lean and flaunt your tail feathers towards all the VCs.
Wouldn't that require you to actively watch the logs going by? In a sufficiently large system, you can't really watch the logs scroll by and gain anything from it. In which case you would need some real time filtering, which basically means hand coding an IDS. :)
lol, in that case, yes. we had 500 servers reporting centralized syslog and many people would just tail -f the central log to make sure not too much crazy shit was happening.
(in a perfect software system (spherical cow) obviously all errors would be tagged and categorized with proper monitoring and alerting thresholds. but, in the real world you have a 800,000 lines of php spread across 1,000 files written by 300 mostly junior people (who only stay for 6 to 14 months at a time) over the past 10 years. you make the best of what you've got.)
Accept the fact that remote logging is necessary (and cheap) for both security and stability reasons.
This was a nice ad hominem attack but I'll respond anyway. I actually have multiple certifications in computer forensics, and have done forensics and incident response for eBay, PayPal, reddit and Netflix.
> Accept the fact that remote logging is necessary (and cheap) for both security and stability reasons.
I'm an open minded person and I'm willing to change my opinion in the face of new facts, but you haven't actually presented any new facts. Do you have any use cases that support your statement?
I have a few facts that counter them. Central logging is definitely not cheap. It costs a lot of money to store those logs at rest, and more money to store them in a way that is searchable, as those data structures expand pretty quickly. It also isn't necessary to stability, given that we made stability go up after we ditched central logging at Netflix (I will be the first to admin this is correlative and not causative, but still, it isn't necessary for stability).
I'm genuinely curious how the security team goes procedures of analysis there.
Hint: google his username.
Here's an example of what you typically want to do: "Give me a list of all customers whose contact information was viewed by customer representative X and who's had a Paypal withdrawal made since that event".
How are you going to accomplish that without logs in sufficient detail?
You might call your audit trail something else if you wish, but you are required to do it (immutable logging inaccessible by applications) if you have a nontrivial application with any sort of compliance requirements.
That's not the kind of data you would get with stuff being spit out to syslog.
You can totally use syslog as a transport protocol for these if you like. You just want to ensure that you're using the reliable form of the protocol (i.e. TCP, preferably with sender-side disk buffering to guard against unavailability of the collector).
When set up correctly, the practical differences between using, say, syslog-ng PE and some other message bus (e.g. Kafka) for recording events becomes relatively small.
We need to take care not to conflate metrics, audit logs, transport mechanisms, encodings, indices, and storage formats. They're all very different pieces of the complete picture and deserve separate scrutiny.
That doesn't work very well if you're using stuff like AWS Lambda functions, short-lived VMs, Docker containers that come and go regularly, etc.
Plus, off-server logs are invaluable if the server is compromised.
Actually that's the use case where it works best. If you have constantly changing infrastructure, how useful are those logs anyway? Unless its something that is happening across the fleet, in which case it should get picked up by your application metrics. Also this is the perfect case for stream processing of the logs where you look for things in real time and then throw away the actual logs (no need to keep them around after you've processed them).
> Plus, off-server logs are invaluable if the server is compromised.
Only if your systems are mutable. If they are immutable then you're much better off looking at the incoming data through an application firewall or looking for unexpected data changes in the data store. The application server should just be a conduit for the user to interact with the data store.
Sometimes I care about the bulk behaviour of my application. I hit 2757 RPS and suddenly find a catastrophic gap in the performance landscape of my application -- that's useful. I don't care much about the particulars of the requests. What matters is the statistical behaviour.
But sometimes the application crashes when a particular input is given. These are outliers, they don't show up clearly in statistical behaviour. In that case I care about the details.
Metrics are warnings that smokers die younger, on average, than non-smokers.
Logs are a biopsy showing which kind of lung cancer you have.
No sensible production system or PaaS treats them as identical, even if you can derive metrics from logs.
But my question is, why do you care about the outliers if the customer experience is not harmed by them? And if it is harmed, then it should should up in your 99th percentile metrics, even as an outlier, at which point you can turn on aggregated logging to look for that one thing. And if that one thing doesn't happen again, then did it really matter?
Parent makes a good point imo. Systems are now sufficiently complex that logging everything is a useless indicator. And if you're only using a subset of your logging anyway (by generating metrics), why collect anything more before you know there's a problem.
Tl;dr: avoid premature log verbosity
The effect of errors is not necessarily linear to their frequency.
Distributed systems of any consequence are not meaningfully modelled with smooth, continuous functions suitable for highschool calculus. There is a great deal of chaos and catastrophe.
Metrics -- including customer visible ones like "zero HTTP responses per second" -- raise the question. Logs provide the physical clues.
Exactly my point. First your monitoring tells you that there is zero HTTP response, then you turn on central logging until you find the problem.
If the problem goes away before you turn the logging on, then it wasn't really a problem, right? And if it happens again, then you can turn on central logging until you find it.
I guess my point is that if a rare bug is rare enough, then it doesn't really matter, especially if the effect isn't catastrophic.
> And since you need the added resources that logging requires to be available at all times anyway, you might as well just leave logging on. What's your objection to having recent logs?
There are a lot more resources required to log everything all the time then to log a subset of things some of the time. Even having recent logs of everything is a huge overhead compared to more detailed logs of just some things.
For example, a lot of people just assume that a bank can't be eventually consistent, but if you make them stop and think about it, they realize that they can and are. The ATM isn't always connected to the bank's servers. That's why your card has a withdrawl limit. That's how much they are willing to risk for eventual consistency.
I feel like the same is true for logging. That other checks can be put in place to mitigate risk to an acceptable level.
Say that you've built your system like any good highly-available system so that you rigorously check consistency of the program invariants, log if there are any anomalies, and then either try to recover or back out with a user error if something goes wrong. You've configured your monitoring to alert if any of the consistency checks fail. Now you get an alert that you've errored out and served a 500 on one in every million requests.
For Reddit, Facebook, Google, or other free services, it's no big deal: one in a million means that you're serving a couple hundred, maybe a few thousand 500s per day, max, and then the user just refreshes and gets over it.
But in a financial transaction app, this is a big deal. The basic assumption you have to make is that if you have a consistency problem that your monitoring caught, you probably have consistency problems that you didn't catch. And the right thing to do is investigate until you understand what is going on and what the impact is. So you pull the logs of all surrounding requests, you pull the logs of any requests that were in flight at the same time, you pull the logs for the user that initiated the request and see exactly what they were doing. And then you try to reproduce the problem, test for it, and fix it.
The difference is entirely in the tolerances built into the system, and how that dictates that you respond to errors. In some domains, a one-in-a-million error is a "well, we'll catch it next time" event. In others, a one-in-a-million error is a "drop everything you're doing, fix it, and ensure that no similar errors exist" event.
BTW, the strictest tolerances usually are not in B2C companies, they are in suppliers of large B2C companies (or the government). The bank is completely willing to risk a few hundred dollars per customer on an ATM card, but if the ATM vendor loses that money without it being part of the spec they gave the bank, they've lost the contract. Most of these enterprise B2B contracts are based on trust, and if the customer observes you losing money accidentally, they'll wonder what else you're doing accidentally.
That's only true if you run a trivial application. If you are a bank for example, it matters very much why the outcome of a transaction came out wrong, even if the effect wasn't catastrophic. Every circumstance and every bit of identity to the transaction is important. "It's rare" or "I don't believe in logging" doesn't cut it.
> There are a lot more resources required to log everything all the time then to log a subset of things some of the time.
Resources are a matter of budgetary concerns, mostly. What does a bug in production cost you? Are you liable? How fast? What's your SLA? Those things matter.
True, but most people unfortunately don't have unlimited budgets to work with.
Also, there is more than a monetary cost. Every system in your architecture incurs a non-zero cost to maintenance, reliability, cognitive load, and management. For example, if you are in AWS, every unused box adds time to the API call for listing the instances. Every extra box that is there for "just in case" does too. It's important to balance these things.
For certain classes of software, zero bugs is not an unreasonable or unrealistic goal.
After making my own logging system, which was expensive in dev time + maintenance, I switched our company over to papertrail ( https://papertrailapp.com ) they handle quite large volumes of logs for less than $100 a month.
They provide an interface to search the logs with low latency.
I'm not sure what type of organisation can't afford it.... (cost scales with usage)
I've given a more detailed example of this in example 2 here: http://henrikwarne.com/2013/05/05/great-programmers-write-de...
I'd expect in a well instrumented system that a single request could easily affect hundreds of metrics. Trying to get that level of information from logging alone isn't practical due to the volume of data it'd require.
If you want to track overall behaviour of hundreds of codepoints then metrics are good, if you want to track behaviour for all of your events at a handful of codepoints then logs are good.
If I were to play the devil's advocate, the real needs for raw log data in a centralized location is for folks outside of Ops: data analysts and data scientists.
What you don't seem to realize is that the cost of centralized logging is not always worth it. Machines are ephemeral and so are many application related problems. It's one thing to counter the OP saying that centralized logging has merits (and I believe the OP agrees with that statement) and another to say centralized logging is always a must.
Half the point of logs is that they're the examinable blast-wave of something that might no longer exist.
If it happens repeatedly, then you add central logging until you find the problem.
The severity of a bug has nothing to do with its commonality. Sometimes there's something that happens once, and bankrupts your business. Security vulnerabilities are such.
However, I'm more confused by the statement "add central logging"—how are you doing this, and how much time does it cost you? If you mean enabling logging code you already wrote using ops-time configuration (in effect, increasing the log-level) then I can see your point. If you mean adding logging code, then you're making your ops work block on developers.
Either way, what is the cost you're imagining to central logging, that you would consider adding it in specific cases where it's "worth it", but not otherwise? It's just a bunch of streams going somewhere, getting collated, getting rotated, and then getting discarded. The problem itself is about as well-known as load-balancing; it's infrastructure you just set up once and don't have to think about. It doesn't even have scaling problems!
Logging is not the only solution, but neither is metrics / monitoring only.
One problem is that we run across multiple services, on multiple machines, so we would have to have some way of passing a customer ID into our AZ level logstash forwarders, and find a way of collating ERROR level logs, and sending those transactions up.
We also have the problem of customers asking for retrospective issue debug, so having all the logs is useful.
Am I the only one that stash GBs of low-level log messages in order to diagnose bugs that occurred in the past?
For logging for devops, I 100% agree with you. Looking at application metrics rather than raw logs is far more productive, and the raw logs should only be consulted after you have triaged the situation based on the metrics monitored.
However, there is another kind of logging, and that's for data science and analytics. Here, it's hugely helpful to have centralized logging. Hell, it is a must. The last thing you want is to have data scientists with a shaky Linux knowledge to ssh into your prod machines. At the same time, logs are the best source of customer behavior data to inform product insights, etc. By centralizing these logs and making them available on S3 or HDFS or something, you can point them there and have everyone win.
Among Fluentd users, we definitely see both camps. As a matter of fact, one of the reasons that I think people like Fluentd is that because it enables both monitoring and log aggregation within it.
Stuff going to syslog isn't generally going to be used for data science.
And if you're not logging it, once you do figure out that 50% of your customers are leaving because your payment endpoint throws an exception on Amex cards, how do you know what's causing that? Now it's just a silly detective game for something you could have been just logging in the first place.
And if you're only able to find out via your customer metrics, it's already too late. If you were just logging the exception, you'd have said "huh that's funny" when you deployed the code, rather than when you have enough data to realise that you've already lost the customers.
Personally I'd rather just have the logs.
I'm assuming that you still have the local logs on the local server and you're looking at one instance as you deploy to find those exceptions.
I'm also assuming that we're talking about a system too complex to actually watch logs go by in real time.
If the log level is light enough that you can actually read the logs as you deploy, then by all means do so. But if you're aggregating them just to be able to search them because you can't read them as they go by, then I'm saying there are better things to be doing than storing all those logs.
One driver has been the lack of information I've had for debugging production issues (e.g. user X not able to log in against authentication provider Y) and so storing extra info is needed (vs the standard web server logs)
Not sure how to separate it out without data duplication (the boundaries aren't clear cut?) and for a small system it prob doesn't matter - I'm nowhere near Netflix, PayPal et al for scale.
There is still a lot of content to be written for languages like python, ruby, and more. We'd love to have more contributors, or even just suggestions for the site. I can also get in touch with any of the authors if you have questions for them.
Centralizing the all the knowledge that's been written about a given topic isn't scalable in terms of contributions. It's far better to focus on enabling contributions than it is to focus on building a solution in which to put the contributions. There's a reason products like StackExchange are more successful at providing features for encouraging contributions than topic specific sites.
> There is still a lot of content to be written for languages like python, ruby, and more.. We'd love to have more contributors,
There is already a ton of contributions to logging solutions on the Internet. Why not build a solution categorizing those, instead of trying to build new content in a centralized (and I might add commercially curated) solution?
Also, where's the source to this "open source" project? I'm not seeing any links to code.
I think the open source part of the title refers to the fact that we cover all different technologies from a variety of perspectives, not just a single vendor solution. Also, we allow contributions from anyone in the community. I couldn't find an easy way for WordPress to share or collaborate on source code, but if you think of a way, let me know. In the meantime, we're giving access to people who want to contribute.
Rackspace site on Github: https://github.com/rackerlabs/developer.rackspace.com
I agree with the readability part. It's important to curate from the stance of educating the multitudes quickly. Having disparate content can be an issue in starting the education, but variety is the spice of life - and use cases as time goes on. I'd recommend content + reference as a solution here. Cookbook style, if you will.
Keep in mind I've spent a fair amount of clock cycles thinking about this than most people. I've rejected things like logging standards as well. That doesn't mean I'm right, but it is, at the least, a data point for you.
0 results
https://www.loggly.com/ultimate-guide/?s=parse+python+logs&p...
Gives you a page on NodeJS
Ultimate ... not so much.
ruby - you guessed it - node.
I'm guessing from the specifics of your example queries that you already have some knowledge about WSGI and Python that you could contribute to the guide. Why not do that instead?
If it was titled "We want to build the Ultimate Logging Resource" I would be OK with it.
I would have no problem contributing, and may well do so over the next while, but titles like this get my goat
Edit: Also, while I know this will be downvoted to hell - Open Source? It is by a commercial company, with -all- most of the examples being their product.
This seems moderately suspect.