Introducing LoggerFS (2013)
developer.rackspace.com
developer.rackspace.com
[1] http://www.bizjournals.com/pittsburgh/blog/techflash/2013/09...
I've spoken with Eric (the engineer who spearheaded this), and while code exists it needs some work for release. It hasn't been worked on since the closure of AutoRef.
Eric and I will be working on getting a release ready within a week. If you'd like to be notified when a release is ready, please sign up for this mailchimp list and we'll email you: http://eepurl.com/WlqXz
I think all the arguments against log shippers are pretty weak, workarounds are simple, especially if the alternative introduces any instability.
In sum, if fuse/loggerfs together can guarantee that every possible state will result in sane (error) state on writes -- this shouldn't be worse than a disk dying/fs corruption/disk-full type failure.
It will obviously be another point of failure.
On a side note, mounting this under eg /var/log, and then having a strong guarantee that failure will result in an unmount, revealing a writable /var/log seems like the best of both worlds. Would probably have to HUP all writers though...
These are two different cases though. If a disk is full the write fails and software usually handles that well. Depending on the state of FUSE, it will just block indefinitely.
A smart logger can have a writer thread and a buffer and decide to drop log lines if the logger is too slow. But LoggerFS is meant to be a drop in solution for legacy code, so the concern is perfectly valid.
I have been there, over-designed myself into a hole and then looking back at what was started as a simple 3 step idea now turning into exponentially growing number of branches and corner cases that have to be handled..
Sometimes it is easier to just say "ok this was not a good design" and just throw it away. I have done that and looking back it was a good decision.
The best I can figure is it's a shipper replacement. So apps write to a log file as usual, but it's actually buffered in memory and pushed to a central server via a FIFO queue.
Given it's all in-memory, it will be small and transient, so you're completely relying on the central server to store it reliably. (Not a criticism, just trying to understand it.)
Basically, it's a virtual filesystem that pretends to create files, but actually intercepts writes to log-style files in this virtual location and buffers those in memory, eventually shipping them off to a place where you can run central analytics on them (ie: Splunk, Logstash, etc).
The application doesn't know that it's writing to a virtual file: it just opens a descriptor, dumps a line or two every few seconds and continues along its merry way. The logs never touch the disk, which means that they don't content for limited disk I/O bandwidth.
I learned about FUSE while getting into Plan 9 and the "everything as a file" mentality. Lets make everything a file!
But this looks like a fun project. Is there an advantage to this over using NFS for logs?
Started in 2007, last update in 2013. Not sure if it's the same thing, but the description is similar:
> LoggerFS is a fuse-based virtual file system that allows you to store log files from apache, syslog and more directly in a database instead of a regular file.
> I actually wrote exactly the same thing years ago (2007), even had the same name :)
I LOVE the idea of turning the idea of logging on it's head!
> The log data is buffered in-memory (potentially journaled for reliability) and sent over a configurable transport.
What are the options for this 'configurable transport'? What is/are the endpoint(s)? Does LoggerFS have facilities for storing and reading-back logs, or does it rely on other services for this?
The post only seems to /hint/ at answers to these questions.
Backend/Aggregator agnostic (includes multiple log transports)
Supports any Syslog-based log manager
Loggly, Splunk, Logstash, Rsyslog/Syslog-ng
ZeroMQ
NSQ transport – used internally at AutoRef.com
Generic UDP/TCP
And soon: AMQP and Redis (and later: Scribe? Fluentd?)
It deals with the specific problem of collecting the logs from the applications on your servers and shipping them to an aggregator (of which there are already many, e.x. Loggly, Splunk, Logstash)As an example, nginx only recently gained the ability to log to syslog; Apache has a logging module but it's not exceptionally customizable if you wanted to log to, say, ActiveMQ, or to a custom service (unless you write a separate binary to accept logs on stdin).
So, basically, it's like systemd-journal except not actually like systemd-journal, significantly easier to use, and can use printf directly instead.
Log-centric (LoggerFS) is a filesystem around managing log files.
Log-structured is a way of structuring data such that it's written to sequentially and is always appending (while dropping from the head).
Log files are most similar to log-structured, but not a filesystem dedicated to shipping log files for centralization.