Great Programmers Write Debuggable Code
henrikwarne.com
henrikwarne.com
It's easy to swamp useful messages in tons of noise, so just logging whatever is not necessarily an improvement. Logging means code, and more code means more work, less readability, and more bugs. In particular, logging somewhere, sometime means I/O, which opens up a whole new class of bugs you'd rather avoid. For instance, I've seen several systems go down simply because the disk space ran out.
Logging is also slow - his example of finding a cheapest path is a good example of this. These are well-studied algorithms that explore huge state spaces. Adding logging as he suggests can easily mean the application's running time is spent almost entirely logging. This might not be an issue: maybe you don't need this very often, or maybe the path's aren't very complex, so you can get away with huge overhead because systems are so fast anyhow. But often as not, you can't. And even when you can, you might be adding a trivial DoS vector for your app, or you might be needlessly limiting its applicablity.
If you have an error, it's nice to log whatever you can to help understand it. But the benefit of understanding the odd error more easily needs to be weighed against the non-trivial costs in time, readability, maintainability, disk space and runtime, particularly if you're pre-emptively logging before an error ever occors.
My main point was this: if you don't get the results you expected, do you know where to start looking for the problem or not? If you have no idea, then there is not enough logging/tracing present in the program.
There's little to gain that way. If the logging was off when the problem happened, you have no logs. You can switch it on and try to reproduce. However, if you are able to reproduce, you can as well attach the process to the debugger and check what's happening. No logging needed.
Logging to disk means I/O.
Also, if you can turn logging/tracing on and off dynamically, the average volume of output can still be quite manageable.
Apparently when saying logging doesn't require I/O at runtime, you gotta talk it out a lot more before people get it.
And since it is a circular buffer the memory cost is fixed. Since the memory is shared and available through the filesystem the log stays around if the application crashes.
In my work, if people wrote more and better log messages I would be delighted. They almost never do and instead rely on the Visual Studio debugger to inspect program state. It works great when debugging your own code, but is pretty useless when debugging someone elses code (you dont know where to place the breakpoints) or when analyzing a crash on a production site.
However, what the article says is "One of the differences between a great programmer and a bad programmer is that a great programmer adds logging and tools that make it easy to debug the program when things fail." In other words, he's more or less explicitly saying people who do not use logging are bad programmers. So it's not at all surprising that there is vocal disagreement.
It's just the nature of HN. I've given up on bitching about it.
java.net.ConnectionException: Connection refused ...
Why not include at least the hostname and port? And what connection is this exactly that failed? You might be able to find out what it was trying to connect to based on the stack trace, but that requires knowledge of the internal structure of the software.
This seems to be a common problem in software written in Java, I don't know why exactly. Maybe something to do with the way exceptions are handled, so that the relevant information is not available when printing the error?
Who are these people writing code like that - they never had to debug a problem before? Include as much as you can! A stack trace is only part of the story.
The key point though: however the problem is detected, you need to have enough information to help you answer the question "Why did it fail?". You can add information in an exception, or in traces/logs, as long as you can tell why it happened, not just that it happened.
2. log everything, write tools to parse the logs. All this turning logging off / on / up / down. It's disk space, its cheap and trust me when you wanted -vvv debugging, the live server was set to -q
3. testable code - oh yes, yes, yes. Functions that return common structures that are then sent on not wrapped up for that one specific use you had in mind(my most common crime) . But make it simple. If you cannot throw in a config file and run it there on the command line, then its hopeless. Stubs, dependancy injection - I get it, I just want to avoid it.
4. Don't do as I do, do as I say :-)
"Debugging is twice as hard as writing the code in the first place. Therefore, if you write the code as cleverly as possible, you are, by definition, not smart enough to debug it." --Brian Kernighan
That said, sometimes cleverness is necessary, or actually serves the goals of making things testable and debuggable.
"2. log everything, write tools to parse the logs. All this turning logging off / on / up / down. It's disk space, its cheap and trust me when you wanted -vvv debugging, the live server was set to -q"
While I tend to agree, it's not (just) disk space - it's IO, which sometimes isn't cheap.
I am not against cleverness, just appropriate for the system - for example some part of the system needs to be very clever else its a pretty mind-numbing project. However the rest of the project should be less clever.
One of my colleagues is a fantastic programmer. Like anyone else he'll occasionally suggest stuff that won't work, but you can tell he's come up with something great when he starts out with "I understand this is trivial, and the stupidest implementation possible, but..." Those ideas have led to some of the most maintainable and debuggable code in our repo.
To me, easily debuggable code (whatever that means) is code written clearly with good documentation. Often well architected systems are modular (rather than a spaghetti mess) and problems are easy to isolate.
Write testable code. Each method should have a specific mandate, should specialize, and should delegate additional functions to other methods that follow the same pattern. It should be straightforward to verify that a particular method is working by exercising it according to the various branches it might take.
This is very hard to facilitate with huge methods that try to do too much, there are too many permutations.
If your software is a computer game, for example, you might have dozens of items, characters, and particle effects being tracked by the game engine at a rate of ~60 Hz. That's ~2000 log lines per second, or 160 KB / sec. Steam says I've played 200 hours of Civ5. Where are we going to put 115 GB of logs?
Not to mention the string formatting and I/O of that much logging potentially leading to unacceptable lag.
Ok, write lots of logs everywhere, but how and where should those go? Wont all those messages get in the way? Will they obfuscate my code? Most of the codebases I've worked with didn't have judicious logging, so where should I look for examples?
Spring's source code doesn't look too bad. Is that a good standard?
Maybe other languages?
No wonder most people just Google error messages.
Every week I come across at least 10 articles telling people about what makes a great programmer. Sometimes its test code, sometimes its some concoction of agile or xp etc, other times its about functional and oop.
Let it rest guys. Just go and hack.