Do Not Log
sobolevn.me
sobolevn.me
If logs bring you some sort of value, then log away. If they don't, then don't. In some domains you absolutely crucially need logs, in others they're largely a waste of time.
Logging is probably the most foolproof way to extract debugging information from your application. The killer feature of logging is that standard out very rarely breaks, no matter how bad things get it works. It will work from the moment the application starts, it will work in an OOM situation, will work if your network switch is on fire, when the DNS records are wrong and your technicians can't even enter the building, it will work no matter how bad things are. If logging does somehow fail, then that is on you, in that scenario you have fucked up your architecture somehow.
Because the reason to log is about people, not about technology.
There's no technical reason to make logs: Just don't make bugs, and you won't need the logs to find out what went wrong because nothing will have gone wrong. Go to the beach!
This is hard though, and bugs happen for all sorts of reason that have nothing to do with the programmer making a mistake-- the customer can be unsatisfied because they didn't know what they wanted. Programmers learn very early in their career to distrust the requirements, because they know that saying but that's what the ticket says gets them nowhere besides an argument: They have to do what the business/customer meant.
Put another way, the user doesn't want logs: The user wants to know if everything is okay or not, and presumably they have some way to figure that out, but if it isn't working correctly (or they merely think it isn't), the programmer will probably want those logs to help them.
In many environments, the loop is larger than just those two people, there's QA support people, and internal experts/operators, and triage, and all of those people are going to have different needs from those logfiles. People might also have a second use for those logs involving anticipation of failures and tracking capacity.
So there are a lot of non-technical, non-functional needs for logs, and once you understand this, it may be a little easier to understand why people would have almost rabid opinions about what those logs should look like (or if they can get the cure they need another way, what they might prefer instead of logs)
> [standard out] will work in an OOM situation
If you use write() which doesn't allocate memory, that's almost true, but if you use printf() it might allocate memory, and it might not work in an OOM anyway depending on what "standard out" actually is.
> will work if your network switch is on fire, when the DNS records are wrong and your technicians can't even enter the building
And yes you can write logs that nobody will ever be able to read, but what's the point of that? When your manufacturing plant is a smoldering rubble, someone is going to want to launch an expedition for those hard drives and try and figure out what caused millions (or potentially billions) of dollars in damages, so if you just wrote everything to standard output and washed your hands of the rest, you might have to find another job anyway.
I've been programming for over three decades at this point, and I'm convinced there is no "easy one-size-fits-all answer" to whether to log or not, or if to log, how to log, so I find it useful to learn about this bit of theology so that I'm better equipped to talk to others about this, and maybe convince them of my opinions.
Reducing the number of points of failure is almost always a good idea. The network being down is one problem, not being able to ever know what your applications did when the network was down is another problem added to the first problem.
I think it is good to think about what happens when things fail: Things fail all the time and most of those failures are benign -- and even those that aren't can often be remedied by turning it off and on, but some of them aren't, so what happens then?
If you run out of disk space, writing a log line to the disk complaining that you're out of disk space is no help, but what else can you do? You could send a network alert, or print something to hardcopy (if there's a printer nearby), or ring a bell if you have one attached -- all of these things are also potential points of failure, but they can also help prevent a meltdown.
And if that sounds surprising, you might also be shocked the number of problems that can be solved with:
<img src="cdn1/img.jpg" onerror="(function(t){t.onerror=function(){};t.src='cdn2/img.jpg'})(this)">
It's for this reason (I think) that people usually qualify your statement as reducing the number of "single" points of failure, instead of the way you put it, because having two single-points of failure isn't exactly like having two weak links in a chain, and in many cases, it's nothing like it at all.In any event, there's another advantage to thinking about the failure-process instead of trying to merely avoid failures:
The biggest reason reducing points of failure is a bad idea is when it is more expensive to do so than to simply tolerate the failure: One thing that dominates the attention of startups looking to improve their availability beyond 99.999% uptime, is "removing the failure points" to get higher uptime, but if Amazon is down, what are Amazon's customers going to do? Are they going to buy from someplace else? Sometimes it is simply that more-established brands have the benefit of ignoring some problems, but you can be sure a large company like Amazon has someone who knows exactly how much more they can spend per second of uptime they may desire, and I think it is worth it for a small company to get good at that exercise as well, because sometimes the juice just isn't worth the squeeze.
As a specific example: My ad servers (HTTP event collectors) are critical, and I want 100% uptime for them because missing a click means missing money, but the reporting dashboard (so the publishers know how much money they made, and the advertiser knows how much they owe me) doesn't need that, so as long as the user got redirected to the landing page correctly, and I can generate line charts and excel spreadsheets once a day, nobody cares that the reporting UI might be down for a few minutes a day-- they will simply press refresh, not knowing if it is their internet or mine, and combined with aggressive javascript-side caching and automatic retries, I am not quite sure any of my customers even notice my downtime.
If instead, I focused on reducing the things that could go wrong (because as you say, it is almost always a good idea), I might spend time trying to improve the availability of my database simply because that's the thing that fails instead of whacking a setTimeout into the JavaScript since the impact is that the user would just retry on failure. That setTimeout saved me the cost of a whole server (and power and heat and electricity) and in the early days of my ad server, that would've been a lot of beer.
Reducing the single points of failure is usually a good idea.
You can easily reduce the number of points of failure by reducing redundancy and removing sanity checks from your process. Doing that does normally not improve your reliability.
Sending the logs to your SAN for it to write in a couple of disks arrays is much more reliable than writing to stderr (stdout is not a reasonable place to log failures) and redirecting it to your local disk.
That's the normal state of life; when a building burns down, forensic investigators pick over the evidence trying to work out what started a fire. They don't ask everyone to record every time two bits of metal touch, every switch switched, every use of electricity, every flammable material entering the building, just in case there's a fire one day. Use a teaspoon in a metal cup? Log it. Post-it note? Flammable, log it. Don't know if your shoes are flammable? Log them anyway just in case.
And yet the cause of the fire is not likely to be any of those things, it's likely to be the normalization-of-deviance fireworks next kept next to the open gas tank where the employees take smoke breaks which people were reporting for ages until they were told to shut up about it.
Beverley Hills Supper Club was known to be overcrowded, badly designed and ignoring fire department reports for years before it burned down. Cocoanut Grove nightclub had flammable decorations and an employee using flames for light. Rana Plaza had big cracks appear in the building before it collapsed, employees were told to go back to work or be sacked. The World Trade Center collapsed because a plane hit it. Titanic sank because it was going too fast and ignoring reports of icebergs and chose nice looking corridors over high bulkheads. Grenfell Tower burned while "everyone" knew the building was coated in flammable coating and nobody in power wanted to pay to do anything about it. Station Nightclub burned because someone accidentally brought outdoor fireworks indoors, and the sound damping foam was flammable not fireproof and it killed people because the emergency exit doors were locked.
We have a culture of log every order the captain makes, every response from the crew, every movement of the rudder, every radio signal sent and received, and at the same time "why did the front fall off?" "a wave hit it" "is that unusual?" "Yeah, at sea? Chance in a million."
> "you can write logs that nobody will ever be able to read, but what's the point of that?"
Cover your ass, cargo culting, habit, defaults are powerful, the cognitive bias effect like psychic powers where the one time people predict a bad thing happening and it happens is very memorable but the all the times people predict a bad happening and nothing happens are very forgettable, it's easy to remember one time logs helped and ignore all the times they didn't and the costs associated with them. For some possibilities.
It's industry specific. In my line of work, yes, they do ask you to make extensive records like that and it is why we log nearly everything. Though I've mostly been in safety-critical and aerospace work (space now, previously just aero). I imagine medical, pharmaceutical manufacturing, and others have similar requirements. If you're making systems that don't matter (in the sense of, if they fail no one cares except a few possibly angry users), then don't bother logging anything.
People treat logging as if it's free, when the reality is that it very often has a significant impact on performance, exacerbated by the fact that it's treated as free.
Add in the fact that today's distributed systems approaches add complexity to the logging paradigm and make it much harder (vs something like Jaeger/OTel) to understand the context of a single log message, and I would rather logging disappear almost entirely in favor of metrics (fast, coarse) and tracing (slower, fully contextualized), with some ability to fall back to just printing metrics and traces to stdout or other local debugging options as a fallback.
Visibility into programs running in environments you can’t access is unreasonably important, but any code that creates visibility will impact your performance, too. The “right” solution is to balance those two needs for your situation, which means there is no “right” solution for everybody.
No it doesn't. Distributed tracing can do a far better job. As do metrics.
There's your performance issue, logging doesn't need to be persisted or even processed by your main program thread or any thread/process actually doing the work. Who's logging straight to stdout or directly ::write'ing a file in a production application?
That said I’m not sure what context we are discussing. I work on nodejs apps now where blocking is a non-issue. I’d obviously think about it more in C’s fprintf() or a similar block-by-default environment.
Edit: also, my area is fintech with $0.1-$1000 range per request, not telecom. I log everything because my employer can ask about everything, and it’s always money.
If you’re logging so much that you’re saturating a dedicated thread that just reads from a queue in shared memory, does log formatting and writes to a file then you have bigger problems than figuring out how to log because there isn’t going to be a solution that lets you log with a deterministic latency without skipping events or wrapping around your queue.
Of course, the approach scales fine if you can have multiple log files per process - e.g. a normal log file and a message log or transaction log, because you can give each file its own queue and thread.
IMO, the advantage of logging is that it some degree of information about actual failure events - in contrast, monitoring only gives you aggregated information and tracing is too expensive to be applied everywhere at any time.
$thousands per transaction? Sample 100%. Cheaper to refund a customer than debug the problem every time? Then sample. You'll have plenty of sampled traces with the bug/error shown as soon as the problem is expensive enough to matter.
Metrics are great for demonstrating a problem exists (and alerting), and tracing is great for helping you understand how and why it happened once it's at a big enough scale to matter.
Beyond that, with a lot of what's coming out of Open Telemetry, you can get fancier like sampling a higher percent of traces to a localized processor middleware that then filters out any traces you doesn't care about (e.g. dropping non-errors or low-quantile latencies).
There is a lot of talk in this thread about the performance impact of logging. I don't see how sampling 100% wouldn't have a far greater performance impact.
> If logs bring you some sort of value, then log away.
I have been on the other side of the problem for years now. Logging has a cost. And the cost benefit is also not great. It requires a human to read it(either that or a regepx hellscape awaits you). This can be fine if you are working with a small service, but it quickly gets out of hand once scale increases.
I've had to deal with double digits(approaching triple) terabytes of logs daily. Most of the information logged was, like the author suggests, completely irrelevant. It makes it much harder to see what's actually relevant. And, for whatever was actually relevant, generally not enough context was provided (something the author also argues).
What you really have to have in production systems are metrics. That's what's going to tell you if shit is on fire and which ass is burning.
Has your app has hit a catastrophic failure and cannot continue? OOM? Sure, by all means log away. Did it have an error while trying to talk to a service? Are you retrying? Then _do not log_ until your retries (or whatever recovery mechanism) fail. Networks have become very reliable, but failures are expected and normal. Spit out metrics instead.
If you don't do this, some schmuck will have to spend their days trying to keep the logging infrastructure from falling over (Cthulhu helps you if you are on an ELK stack) while management will be complaining why noone seems to notice failures before customers do.
If you find yourself parsing logs to alert on issues, please revisit your architecture.
METRICS FIRST. Logs are a safety net and a tool for troubleshooters.
Logging is something that needs to be maintained (like all code), it can't be just some vestigial by-product of the development that's left to its own devices, at least not if it's to be of any use. The same is true for metrics. If your metrics are just some arbitrary selection of times and counts that seemed good at the time, they're probably not very useful. Which data to collect is something that needs to be deliberately chosen, and re-evaluated over the lifetime of the application. When you are launching a new product, then lots of logging is probably a good idea. When that product is mature, you probably should have dialed it back several notches, as you most likely already know the general areas where things go wrong.
The issue is -- some people find value in the logs but some other people don't.
I am in the strong camp that find logging provide negative value, mostly due to the extra complexity it brings and its obstruction on code reading. And I believe the reason the other camp finds value in logging is because they lack insight to pin-point issues and lack effort to gain that insight. But I can't really use that argument because it can easily be deemed offensive and solicit slash backs rather than constructive discussions and debates.
Logging is my most frustrating pet peeves.
IMO, it’s because the right logging solution for one situation is not the right solution for another situation. That means that one person is “right” and the other person is “wrong” (because the context required to identify why one person is right is complicated and hard to express, and is thus missing).
Observability (of which logging is a subset) has a performance impact, and software has performance requirements. This means that the right amount of observability is going to depend on the software, not a standard.
It is true however that looking at the individual replica, application logs should be very reliable.
And then you hope you are not logging to a logging service that'll try to execute something that was logged.
In the end, to integrate Sentry and get actually useful information for monitoring, you have to do the same work you’d do for logging. Except now you’re tied to Sentry. Want some other service? Good luck.
Do log. And then send the log messages to Sentry or whatever you’re using, too.
Also, do overlog. It can be hard or impossible to attach a debugger to the deployed application. But you can always turn the log level up to 11 and then reconstruct what happened from the log messages.
And then catch it at the highest level, including all the stack trace. And then maybe send it to sentry, or log it, put it on data dog.
But don't catch exceptions all over your application code.
I have a related mental model for this: When you crash, crash loudly.
But in that case you probably still want to know about these failures, so you can review them from time to time, or see if there are upticks in failure frequency.
If you know that an operation is unreliable, you need to add retry / recovery logic and perhaps log or increment a metric that you can track what's going on; a high number of failures often indicates an exceptional situation.
In some cases, your retry logic will be a "Please try again" message to the user. I think that's perfectly acceptable, but you'll still want to know about it.
I have a third party which returns errors ~20% of the time outside of working hours (yes, really).
1. Exceptions should be sent to sentry. 2. Exceptions should not be normal. 3. Do not log.
Solve for ability to review & debug after the fact.
Then presumably those errors shouldn't be exceptions. Of course, there's a lot of a lot of assumptions behind this:
1. You have control over the code that interfaces with the third party. Otherwise presumably you'd have to write a shim that catches the exception and turns it into an error value, and that's probably not worth the effort.
2. You're work in a language that doesn't type check exceptions (i.e. not Java with checked exceptions) and your language also doesn't have a nicer way of handling errors as a (discriminated) union.
3. Your co-workers will go complete bananas if you start using some sort of result type or [success, value|error] tuple instead of an exception because that's how they've done it for 20 years.
The reason for this is that exceptions in most languages escape the type system and are in general more awkward to catch.
Of course, I'd also disagree with "do not log" in this context, though I might prefer a logging mechanism that allowed me to write queries on it (like recording values in a database).
Even if you really really like error return codes and really really don't like exceptions, you'll likely agree that writing very un-idiomatic code that goes against the grain of the language and all the libraries you might be using is likely a much bigger loss than the value you might have added by avoiding exceptions in your part of the block.
Out of curiosity, in what languages are error return types non-idiomatic?
This isn't to say the idioms are never broken or that they shouldn't be broken. I agree that if you are doing something like parsing a string to an integer, that should probably be a result type and not just throw an exception.
In languages like Java an Python, you can return error values but because you can also always return nulls, you're just making more trouble for yourself than it's worth avoiding exceptions.
It is not easy to stay away from exceptions if your language supports them.
But even for exceptional exceptions, you may still want or need to retry or branch based on what was thrown.
It's not an either-or situation though. Even if you let the exception bubble up and sentry logs the exception, logging performed prior to the exception and at the time the exception is caught (prior to being bubbled up) as valuable context to how the exception happens, which can lower debugging time.
The exception just tells you that something went done, it doesn't always have the reason why something went wrong. And the context for why is not always available at the time of raising the exception (the context may be from a prior function call, or the consumer of the function that raised the exception).
Usually, but not always. That is, for the classic "exceptions should only be used in exceptional cases" scenario, this is usually what you want to do.
A good portion of the time, though, this scenario doesn't apply. You're dealing with a subsystem that is throwing an exception (for good reasons or bad) but no, you don't want to just the application crash. You want to ... catch it, and do something else.
Which is what try-catch was devised for in the first place, after all.
Ok, I did all that. And I still want to know why this bug report happened, and logs help me. Why should I Not Log?
Yes, but you should catch as soon as possible (or validate stuff as soon ass possible) and re-throw with a better error message if possible. You will thank yourself when some problem appears and instead of a random "errorcode 123" you get a detailed message like "Operation Xyz failed with errorcode 123";
If this is a concern, just create a class with a method that takes an exception, message, context, etc. as parameters and use that. Now you can change it in one place instead of throughout your codebase.
Agree on logging messages as well; is useful to see live what is happening in the system as well as it is useful to see logging broken up by different errors
Please don't overlog. Only ever do this if you can dynamically change your log level on demand, without impacting users.
If you do have to reconstruct a sequence of events and you do not have distributed tracing, fine. Log the sequence of events _when needed_.
But please don't log messages on a perfectly fine system for no reason. It's noise and health can be better served by looking at metrics.
Because of this overlogging nonsense our systems generate terabytes of logs daily. Most of it is never ever looked at.
>don’t use logging, use business monitoring
Okay, if revenues suddenly drop 50% - then what? Where does the problem lie and once it’s discovered how can we alert on a recurrence of the same issue and fix the issue? Presuming the underlying issue cannot be trivially permanently remediated.
The author seems to see logging as essentially making excessive alert noise inevitable when logs are meant to just be collected somewhere and scanned for exceptions known to cause incidents. Then when that exception happens at the same time revenues fall 50%… there you have it?
The authors observations are mostly valid but their conclusions just don’t make sense. The observation you shouldn’t ping an admin every time a log generates an error doesn’t mean “don’t log”
The "real" point seems to be that logging should be useful and consistent, which I agree with.
So one tells her husband he's about to die and is dying and finally dead and while he originally objects she is so persistent that he ends up alive but perfectly silent in his own funeral, that is - until his neighbor comes in in his birthday suit after his wife has made him a stylish suit of the same kind that the emperor got.
At that point the supposedly dead guy can't manage it anymore and cries out laughing at his neighbor coming naked in his funeral.
Sometimes I seriously wonder if some of the stuff that gets posted here is that kind of stuff.
Edit, a translation of the full story, "Stupid Men and Trolls for Wives": https://everything2.com/title/Stupid+Men+and+Trolls+for+Wive...
If the author is reading this, the Sentry-related link you have in the article [1] is currently giving the readthedocs.io 404 page. Same for the "docs on logging" link about 4 sentences earlier [2].
[1] https://stories.readthedocs.io/en/latest/contrib/sentry/
Logging each exception vs logging nothing is a false dichotomy. I'm not sure anyone whose suggested fudging a 'Result' type in a language without sum types has ever done so, or at least stuck around afterwards. I'd like to be equally condescending in my answer to 'is my team ready for logging side effects?'. The fact that a distributed log aggregation, storage, and visualisation framework has a docker-compose file no more complex than a well configured SQL database does not terrify me.
While ERROR level logging should immediately generate alerts, it's still useful to watch metrics on the WARNING level to see if something is potentially heading towards ERROR-level failure.
ERROR / FATAL: The application failed to execute the business operation.
WARNING: The application had an issue executing the business operation, but has a fallback -- for instance, retrying. Might still result in an ERROR later, if the fallback operation fails.
INFO: Document the successful exit of business operations and decisions. This typically helps you determine where in the process the failure occurred, and what might need to be repaired if you decide to manually rollback the operation.
DEBUG: Document the entrance of business operations and decisions, along with any additional technical operations, decisions, and values that might be relevant.
DEBUG - Only useful for developers and QA looking to understand exactly what happened and why on a per-request basis.
INFO - Provide information about the current state of the system. This was our production Log Level
WARNING - Something happened that will likely require action at some point in the near future, but not immediately. i.e. Approaching quota limits, temporary conditions we can work around, etc.
ERROR - Something unexpected happened and should be looked at soon if only a one-off or immediately if consistent. SEV 1 or 2 depending.
FATAL - Conditions prevent the system from being able to operate correctly. Generally a complete and catastrophic failure that will not resolve on its own, such as missing dependencies, invalid configuration, etc.
The reality I've far too often ran into is: something is logged as debug, production is set to info+, something rare happens, and we don't have the logging to diagnose it because it was being logged at debug. So we swap it out for an info log, or raise the global logging level of production, but we still can't diagnose it because the rare event happened in the past; we just hope it happens again, or we spend hours trying to blindly figure out the conditions which caused it to happen without the information to actually reproduce.
Eventually management gets fed up, and authorizes debug+ ingestion. The logging bill doubles, but at least we can "solve it" with money (the log stream is twice as noisy, and half as performant, but at least we have the data right?)
A lower-level problem surfacing here: As engineers, we decide on a log level when we write the code. Its in response to expected system behavior; not actual system behavior.
In comparison, a better solution would be: surface logs only when behavior betrays expectation. I don't want that debug line in 99% of requests, but for a 5xx request? I do. Logging systems are horribly unsuited for this, because it requires (1) a context (the request), (2) correlation between different execution points in the context (only log if later the request ends in 5xx), and (3) delaying log output (or ingestion) until that correlation is completed.
We have a solution which does this out-of-the-box: Tracing. Tracing is designed to be sampled, and trace ingestion systems all have a "sampling" function designed to identify and surface only "interesting" or "anomalous" traces (if you want).
Also, this [1] is doing logging right. Completely underutilized.
This has several benefits, the largest of which being that you only look at log files when a thing logically goes wrong, and every entry you see pertains to that exact activity.
The other benefit is that these logs get purged at the same grain as the work items being cleaned up. Successful items are purged quickly, exception items can be kept perpetually.
So, the implication for us is to log with a lot of verbosity with the expectation that we are probably never going to see any of it, but also with the knowledge that a ton of helpful debug lines will be available when it does all go wrong. Stack traces are not enough for us, we need to know exactly what communications occurred between us and external systems over time to figure out many issues.
We do maintain a global log file, but this is mostly just service start/stop entries and other administrative actions.
There have been uncountable cases where a GOOD log trail has been the difference between a 20 hour journey through a code base and a short stop to immediately discovering the problem. Common failure points should be logged, especially when they are deeply nested, or they involve complex hard-to-diagnose processes.
Overlogging is just as bad. But to say "do not log" and then suggest using services like sentry or datadog instead is really reaching around your nose to touch your face. You're doing yourself a disservice by lacking the final bit of correlation - the approximate location and nature of the failure. There have been times when I've been on a service so complex that just building a harness to reproduce the case takes hours or days. Not everything is a nice, clean, small and simple microservice where you only have N guesses at where the problem could originate. On a large enough, or old enough, service the upper bound could approach N! very quickly.
I'm not sure why this was posted to HN as it seems like a very low-quality ranty post by someone whom I'm not sure has ever worked on a large enough project where it becomes non-obvious what the origin of some error alerts are.
It really is not: you can trim overlogging, you can't add missing logs.
Now one thing I am very much in favor of, but which I didn't see anywhere in TFA, is the ability to dynamically instrument instances (add or remove log points).
It's not necessarily easy to know what to log or not to log. With "normal" logging adding new log points is a complicated process (they're static so you have to introduce them, ensure they're correct, and then go through whatever deployment process you have), so there's a natural tendency to overlog because if you underlog you may miss critical information.
In systems where the capability exists, instrumenting log points "from the outside" without having to update the program being observed is much better, not only does it let you refine your logging more easily, it also lets you brutally increase logging density (possibly locally) without downtime.
>>> l = [x for x in l if ...]
lines. I imagine there's a better way.In some systems / runtimes, yes. For instance every browser supports "logging breakpoints" which can be added and removed on the fly (create a breakpoint then right click, you should get an option or sub-option about logging). I would also expect that there's Java instrumentation things which do it, though I haven't done any java for a very long time.
It's pretty rare though, obviously so for runtime-less programs, somewhat less so for those with a runtime but while "external" profiling is common logging is much less so (e.g. Python supports updating the configuration of the existing static logging system on the fly, but not defining the log points).
> It sounds like an application restart would be warranted at the very least, which can be disruptive.
See above, if you can profile an application on the fly without restarting it you should be able to instrument for logging.
And tools like dtrace/ebpf should allow this, but I've yet to take the plunge and actually try and look at that, it's only been an idle thought so far.
I had situations where the difference was between having an intermittent problem for months or years that nobody can figure out and pouring over logs for days and finding the problem.
I do pretty extensive logging in my code and it has saved my ass many times. Before people complain about performance or storage problems I want to see hard profiling numbers. In most cases you it’s only a few places where the logging is causing problems and you can reduce it there.
Overlogging is almost never a real problem in my experience. You can always discard log data you have but you can’t create log data you may need but didn’t produce.
I would disagree, after managing a 100 node ELK monstrosity to try to process the terabytes of logs generated by our systems.
Overlogging is not a problem if log levels can be easily changed (preferably without an app restart). One app logging like crazy is not a problem. A thousand of microservices, is.
> what authors qualifications are
The author is definitely not very qualified, yes.edit: how are you going to communicate information as to why something crashed to your colleagues who do not understand certain parts of the system as well as you? A stack trace is not effective at expressing intent. Explaining yourself in logs helps people triage more efficiently.
I also agree that too many log statements can actually be counterproductive by providing a lot of noise.
But the idea that logs have no value is silly. Every developer writes bugs and will run into situations where the code is behaving differently than they thought in production. Being able to follow the logs to understand what happened is extremely valuable.
This reeks of functional programmer brain.
The world is not immutable. the world is weird. Your systems will interact with other systems that don't always quite work. Sometimes, their systems are wrong, or update without your knowledge, or theres other issues in routing, networking and whatever else.
Logging errors like "I can't resolve this domain" is more helpful than "my customers can't use this integration and business is at zero"
Logs are there for observability, so you can get those, "Oh wow! Cool!" moments when you see something that you never knew.
If you never make a mistake, or if you know your external systems and uses perfectly.... Don't log.
If you're a mortal that gets surprises... log.
I don’t quite understand the alternatives.
However, I do agree that logging is a side effect and should be done in specific ways. That does not mean that you should remove all of it. A little goes a long way.
That's usually enough argument for people that I get to not have a logger statement in the catch. Frankly it's a religious thing. No one really wants to have the argument, and the best you can hope for is a sigh of defeat.
Don't follow this article literally - as great as monitoring (which I take to mean metrics that you can use to plot graphs) is, you'll never have defined as many monitoring metrics as you want, and logging can contain a much greater level of detail for each event, which allows you to reconstruct the course of an incident after the fact (allowing you to solve the incident much faster than trying to figure out how to reproduce it). Also, in regulated environments, logging has the advantage that it can contain data that's too sensitive to be recorded long-term, even with ACL checks, since debug logs are generally retained for only a short time (few days or so). You generally want your graphs to be able to use data from farther back than that (up to a year).
The best article that encompasses the Google logging philosophy is https://apenwarr.ca/log/20190216
Other points:
- Don't log to a synchronized file if possible - log either to the network or at least to RAM (which can be emptied to the network in the background).
- For the overlogging problem, get a logging library that rate-limits your logging for DEBUG/INFO level events, ideally rate limiting by call site.
- For Java, it's very convenient to be able to instrument code at runtime to add and remove logpoints.
Not necessarily, e.g., you can buffer in memory, and write to storage later.
> loggers can fail and break your app
Anything can fail and break your app. Weak point.
A lot of devs think of logging as free; it's easy, it's almost always something that's either global-scope or standard library, and it's the first obvious solution to inspect internal program state when the program misbehaves.
But logs are a subsystem and adding any subsystem to your code increases your risk surface (for both crashes and privacy violations).
Took an operating systems class in college where the first thing the prof drilled into us was "Don't debug with printf. You're going to want to. It's the thing you're most accustom to having written application code. But you're writing an operating system now. If you have a memory overrun, is the first thing you want to do really allocating a kilobyte of stack to mangle strings and determine if you need localization support to tell yourself something's wrong with memory? By the time the output gets emitted to console, the memory you want to inspect will be a blasted wasteland."
As an example of how logging can break...
> you can buffer in memory, and write to storage later
You may as well not have logged then if the failure issue is a segfault; your program context will collapse before the buffer is output.
Once, when working on a shipping desktop application, we got to the point where we had to care about performance. When we turned on a CPU usage trace, the highest CPU consumer by a few orders of magnitude was the logger.
That surprised everyone, so we dug in a little deeper. Turns on log4net's default behavior is to calculate a complete stack trace for every logging statement, even if the stack trace isn't being logged.
Now, there may be good reasons to log a stack trace now and then, especially if you're trying to diagnose a difficult-to-reproduce bug. But, it's trivial to calculate the stack trace yourself.
When we reconfigured log4net to not calculate stack traces, our CPU usage lowered dramatically.
Lesson learned: Don't ship code with obscure features enabled that are performance pigs. There's no reason to calculate stack traces on each logging statement when the stack trace isn't being logged.
try:
do_something_complex(*args, **kwargs)
except MyException as ex:
logger.log(ex)
Basically this method is a giant @TODO, which does make sense in some cases. Either it's unknown at the present moment what a piece of data will look like (like scraping, or some poorly undocumented API); or the developer just doesn't have enough time and will deal with it later, which you can laugh at, but it can make a lot of sense when looking at the higher level goals of the project, our time is limited.I usually do this at the top level of a batch job, where you don't want it to halt the entire job. Or for a GUI app, but then also launch some sort of notification to the user / reload the app. Basically whatever makes the most sense in the context for accounting for something you didn't account for.
> Just log all uncaught exceptions in something like Sentry, or a file for that matter, via a global exception handler.
(And in before "just swallow it" - nope. I've frequently found bugs by monitoring non-fatal exceptions).
I agree though that catching, logging and then immediately rethrowing exceptions is redundant.
> What? Perhaps in some programming languages, but catching exceptions globally in C# works just fine.
I think what sesm means (as do I) is this situation:
try { a(); }
catch (ex) { log_warn(ex); }
return b();
Right now, the function will continue executing and return a value even if an exception occurred, which sometimes is desirable. If you removed the catch block, nothing after a() will be executed.There are some cases where you basically have to make your own top-level error handling - for example in (java) worker threads. An uncaught exception would either terminate the thread or would be swallowed without showing up anywhere.
The new meta in Rust is to lift out loggers and replace the concept with tracing. There is a `tracing` package where you can setup spans like in OpenTelemetry but it also supports the logger macros so you can attach log messages to the spans.
To consume this you add a subscriber configured at server startup to send structured logging to a Json logger, or ship them to a Jaeger endpoint. Or sentry, as suggested in the article.
https://www.lpalmieri.com/posts/2020-09-27-zero-to-productio...
My motto is that so long as the logs don't break your performance budget, log away. However, do take care to keep logs easily searchable and filterable. If needed, you can even have certain logs go to separate places so that you can have the most important logs in one place and more detailed logging in another. However, in that case it is usually best to have one set of logs be a superset of another to avoid needing to join across the logs.
Comes off a bit as xor thinking. Thank god for jazz music. The off notes and in-between notes create the most interesting sounds that expand the thinking.
PPSocialHighlightStorage: Constructed social attribution has missing data: (null)
What is that message doing in shipping software? It cannot possibly help anyone.
Correct. Those low level messages should not be present in 'production' environments _unless requested_. The ability to log extra info for debugging is fine. Doing that by default isn't.
(User-visible) Warnings and errors only please. Can always dig more and turn up verbosity if needed.
Perhaps I'm being too cautious (or too naive)? But it's saved my skin a few times.
The only thing that I don't like about this technique is that the 3 month period is arbitrary. But it works well enough for me.
But then says anything that is important enough to log should notify you.
You don't want to be notified for every spurious network connection, login failure, etc or you will be drowning in notifications. So you realize you want to queue up those notifications so you can filter them. Congratulations you have now just recreated logging :-/
Here is how I recommend to perform regular log reviews: https://elastisys.io/compliantkubernetes/ciso-guide/log-revi...
Most log collecting solutions have ways of producing metrics and alerts from logs.
It takes more discipline to be able to understand why a metric or alert happened without supporting logs.
Sourcing logs from multiple services, including dependencies, allow for creating more intelligent metrics.
Treat logs as a stream of events rather than noise created by developers and logs become more useful.
Imagine you build all sorts of dashboards and Workflows on top of logs or log based metrics and the dev team changes the structure of the logs or just stops logging. Suddenly all your analytics is out of whack; this is because you created a hidden dependency.
Communication and the decision to depend on the logs as a source of metrics would be key to keeping these systems connected and functioning.
Taking a dependence on logs based metrics without communicating that is the failing here.
The same breakage could happen by depending on the event being emitted and a developer changing the event structure.
Communication and understanding helps both scenarios, not duplicating the effort by emitting logging for debugging, events for auditing and metrics for business.
This article is so riddled with weird assumptions and wrong advice that I have to seriously question the author's practical experience.
The real solution is to make both automated and manual feedback extremely easy, for example send crashdumps without asking and have a forum connected to the users account where you redirect the user upon a crash.
Just like downgrading my GPU to a 1030, so I cannot play stupid games; this is a motivational trick to curb me into making things better!
Apparently I'm also somewhat of a masochist.. :S
-------
> I either want a quick notification about some critical error with everything inside, or I want nothing: a peaceful morning with tea and youtube videos. There’s nothing in between for logging.
1. No, there's lots in between.
2. You should be able set the log severity threshold, and hopefully the log verbosity threshold (e.g. with/without stack trace and other information).
3. We don't just log failures.
Ye just log everything. You mostly don't know in advance where the error will be anyway.
The author seem clueless. Don't log but use a Cloud $$$ SaaS for monitoring? Really?
And then he goes on to argue that the code should just not fail. Like, "write good code".
Logging is a byproduct of a past time; everything is a file, stdout is a file, lets persist that file, now we have multiple replicas, lets collect the file into a multi-terabyte searchable database.
The biggest downside: Its WILDLY expensive. Large orgs often have an entire team dedicated to maintaining logging (ELK) infrastructure. This price-tag inevitably leads to bikeshedding on backend teams about how to "reduce the amount we're logging" or "cleaning up the logs" or "structuring them to be more useful".
Outside of development, they are so rarely useful. Yet inevitably someone will say: "you're just not structuring your logs correctly." Maybe that's true; similarly, I don't find vim to be a highly productive editing experience. Maybe I just don't have the thousands of extensions it would take to make it so. Or maybe You're stuck in the past and ignoring two decades of tooling improvement. Both can be true.
I'm phrasing this as a false dichotomy, because in many teams: it is. Logging is easy; its built-in to most languages; so devs log. The information we need is in the system; its a needle in a haystack, but the needle is there. We log when a request comes in, when it hits pertinent functions, when its finished, how its finished, the manager says: "just look at the logs". Instead of "what better tooling can we make an investment in so future investigations of this nature don't take a full day."
For starters: Tracing. Tracing systems should be built-in to EVERY LANGUAGE, just like console.log. We have a standard [1] sponsored by the Linux Foundation and supported by every major trace ingestion system. This is not a problem of camps and proprietary systems; its a problem of culture. I should be able to call a nodejs stdlib function at startup, specify where I want traces to go, sampling rate, etc, and immediately get every single function call instrumented. Its literally just highly-structured-by-default logging! Dump the spans to stdout by default! Our log ingestion systems can read each line, determine if its a span, if so route to trace ingestion, otherwise route to log ingestion.
This is a critical step because it asserts that tracing is actually a very powerful tool that everyone needs to learn, like logging. Everyone knows about logging. Why? console.log. Its there, it gets used. Tracing right now is relegated to a subculture of "advanced diagnostics"; you gotta adopt a tracing provider, bring in dependencies, learn each implementation of OpenTracing, authenticate to send traces over HTTP... as a community, we should normalize just "dump traces like you dump logs, to stdout", have a formatter to make them nice to use in development, and now all that instrumentation work that any dev is capable of utilizing (just like console.log) "just works" in production.
It's like someone going to a hospital emergency room and the triage asks what's wrong hoping to get "slipped on ice, possible broken hip" but instead gets a five hundred thousand line report of "blood oxygen level nominal" "Ow! bumped shoulder into doorframe" "Counting body arms still 2 present" "PANIC: hasn't got a card for that birthday tomorrow" "HELP! heard on TV show earlier" "Counting body arms: still 2 present" "NOT FOUND error when looking in drawer for spoon used one from dishwasher instead" "Counting body arms still 2 present" "Result 0xDEADBEEF from procedure defined in D:\dev\2016\src\folder\inside_joke.h" "Counting body arms: arms not found. Retrying in 10uS" "Counting body arms: arms not found. Retrying in 10uS" "Counting body arms: arms not found. Retrying in 10uS" "Counting body arms: arms not found. Retrying in 10uS" "Counting body arms still 2 present".
Oh well, you say, that would help the developers divine the internal state of the system. ... but they're asleep right now because they're too good for on-call work. Isn't it kind of them to provide you with log messages which is easier for them than carefully making things which work well? Oh and they won't divine the internal state of the system until after you've determined if it's a bug, reported it, and their PM has prioritised it. Until then, hopefully the tech support department knows that counting body arms stops working properly when balance is lost.
I suggest that logs are mostly useful when they indicate some general flaw like "can't contact server" and less so in other situations, similar to survivorship bias: planes returning from missions with bulletholes need reinforcing in the places with no bulletholes, those are the ones which stop the plane flying when hit. Or similar to the principle that robust systems have weirder problems, because the easy problems have been designed away. Similarly, the places you've carefully built with logging are the ones which you've spend enough time on that they'll probably work. The places where faults will come from are the unloved unlogged rotting parts, or the unexpected interactions between systems.
[davnicwil's "make writing more fun" article yesterday said not to head off a bunch of tedious nitpicking, so I won't, just assume trivial dismissals have been headed off as necessary].
Logging is a side effect
Right -- inside of a pure function, logging is probably not a great idea.
But in other parts of your code (that are already generating "side effects" -- because you know, that's what software does) -- it can be not such a bad idea.