Logs Are Streams, Not Files
adam.heroku.com
adam.heroku.com
Using non-blocking IO is a solution, but then when the fd is unwritable for a while, you buffer data and your machine runs out of memory. So also not good.
I'm a big fan of the strategy Varnish uses. It mmaps a block of memory, and then treats it as a ring buffer. A separate program can connect to this and read the logs to a file or a pipe, but if that program fails, you don't stop serving web pages. As long as allocated memory doesn't randomly disappear and the Universe doesn't implode, logging adds no overhead or risk of failure to your app.
Logs are essential, but I don't necessarily want a failure in the logging system to break my app. When you use files or pipes, you run this risk.
If it's an error that you're logging the moment before you crash, then you probably want it to block.
At Google LOG(ERROR) blocks but LOG(INFO) doesn't, for this reason.
If you were being pedantic one might say that logs are a repository for temporally tagged event notifications that have occurred in the past. You can store these temporally tagged event streams in files, in a memory cache, in an email inbox, or a round robin database (RRD).
Not particularly profound, but useful. Logs from disparate processes which share a common time basis however are the systems analyst's go to tool for trying to untangle secondary and tertiary properties of a loosely coupled system.
This also assumes that anyone is actually really reinventing the appenders as most frameworks already have a library of common appenders.
Not to mention, it's just less development work to send out messages with, say, different syslog levels and be done with it.
I certainly did, so many years ago...
Everything is an integer.
And of course there are different sorts of files... everybody knows that.
If everything was a file, Linux would have like 10 system calls total.
This is simply not true. I changed the logging method of a C program recently from timestamped text file to syslog in about 10 lines... Utterly trivial in any language.
A few months ago I published a suite of scripts that allow me to run complex queries using shell pipes. They work by communicating in YAML. The first stage, cat_logs, simply parses logs from a file and emits yaml for each line. cat_logs understands that log files are split across multiple files, and that they may be gzipped.
I've taken a look at statsd+graphite, but that seems to be much more of a realtime solution, and I want to mine the logs I have on hand.
It's fairly obvious that a log file is a file. But "the log" is clearly a never-ending stream of information.
The author makes other good points but the title is linkbait.
Considering the complementary term they use ("drains"), wouldn't it be more intuitive to call them "faucets" or "spouts"?
Just sayin'.
tail -f will continue reading a file for as long as possible, but if that file is renamed underneath it, then tail won't know and you will be wondering why you don't see anymore log lines being output to your terminal.
tail -F checks to see that the file we opened is still in the same location as before, so if the file is renamed underneath us, and a new file is created in its place (think log rotation) then it will open that new file and continue outputting data to its standard out.
http://git.savannah.gnu.org/cgit/coreutils.git/tree/src/tail...
Doesn't look very bad, but I am known for being an optimist.
I think it would be wonderul if everyone started logging to twitter.
Then you can just follow each machine to get a read of its logs... or, at least pipe the output of LogWatch or something like it to twitter.
Maybe for 4/1/2012 - we should organize all sys admins to output logs to twitter randomly choosing @jack or whomever else and flood them with @ log bursts...