A log/event processing pipeline you can't have (2019)
apenwarr.ca
apenwarr.ca
From my experience if it comes to log shipping from hosts rsyslog + relp + disk assited in-memory asynchronous queues are preferred, most of the time you just only have network i/o as logs would not touch disk.
The idea is to ship logs off the device ASAP as well as destination acts as a sink server capable to handle most of the spikes withouth stressing local source. All done via rsyslog which also wraps actual logs into json format locally. The glue could be syslog-tag.
At the other end you could have ELK stack and logstash using json_lines codec input (pretty fast) structuring data further to your likings.
Just looking into metrics now the avg time for logs showing in ELK is 7-200ms (the latency comes mostly from specific reads happening against the ES cluster).
As ELK is always the slowest component, dropping logs compressed in-memory directly onto disk is also an option.
One thing to note is that RELP can produce extra duplicates which are easily handled by inserting into Elasticsearch using specific document ID (some performance penalty) which could be some unique hash computed on (log content, timestamp, host) etc. With this in place you can also easily "replay" stream of logs to fill potential gaps.
This type of setup scales really good as well.
Edit: typos
(These are micro-optimizations anyway.)
https://www.elastic.co/guide/en/logstash/current/plugins-fil....
Edit: typo
We make this stuff harder by rebuilding functionality on top of systems that already have the functionality.
If there is log storm coming from misconfigured app depending on your traffic levels the sink servers should give you a lot of space for an on-call engineer to be notified that something is off and adjust / rate limit / drop accordingly.
Edit: udated with more info
The json that syslog wraps data in is actually using the same syslog protocol - nothing is chaging here. It's encapsulated within the same protocol. The only reason to unwrap json is when you need to start structuring your data for BI views, your milage may vary. The key is that these two are de-coupled.
Because you work on a basic level networking / storage layers its very easy to reason and build reliable pipelines according to your needs. As you say — you can't beat that.
PS: The pain starts when you need to deal with multiline logs like java stack traces but then you just work with devs to log everything into json.
Edit: split comment, more info
syslog-ng, logstash or fluentd on the host to collect and aggregate logs. (logstash/fluentd can parse text messages with regex and handle a hundred different things like s3/kafka/http but they are much more resource intensive).
kibana or graylog to centralize logs and search, the storage is elasticsearch.
A simple syslog-ng on the devices could probably do the job. Little known fact about syslog, it can reliably forward messages over TCP, logs are numbered, have retries and syslog-ng can do DNS load balancing.
If they were storing everything in S3, it's possible to do something similar with ElasticSearch for the same order of costs (maybe three times?), the money going to EBS storage instead. ElasticSearch allows to query and visualize logs, which plain S3 storage doesn't, it's worth a bit more IMO.
hosts -> rsyslog collector -> kafka -> custom crunching -> elastic
which is a bit inefficient and we try to go hosts -> kafka when we can but a lot of stuff only supports rsyslog so it's there and simple enough.
The rsyslog collector and our crunchers are teeeny tiny compared to the rest of the pipeline and can chew through up to a week of backlog in a few hours. The bottleneck is the network for us and if we upped the pipe to 10G we could probably get away with a single host.
Edit: Sorry, I meant "Fluent bit". No idea how fluentd handle this scenario, but I was told it was too slow (being written in Ruby) so that's why the switch to fluentbit was made.
I'd personally recommend to not bother with processing multiline output into a single message. Lots of trouble for no benefits. It's just a stream of lines at the end of the day, it will look the same in tail and kibana.
Fluentd would not just be slow, but also run out of memory. I am no ruby-head, but a then-colleague of mine helped configure it. It still ran out of memory. I had no patience for a log system that did not work out of the box on one computer basically just logging failed ssh login attempts.
We ended up building our own general purpose log parser in Go to support our hosted log monitoring platform. Though it's still somewhat young, performance is great (beats Fluent Bit in most of our benchmarks), and it's nearly as flexible as Fluentd in terms of configuration. If you want to check it out, we recently open-sourced it and are always looking for more feedback: https://github.com/observIQ/stanza
...
> The kernel notices that a previous dmesg buffer is already in that spot in RAM (because of a valid signature or checksum or whatever) and decides to append to that buffer instead of starting fresh.
This sounds like it should be very unreliable. Perhaps it works in practice but I couldn't see myself relying on such a mechanism.
I don't know about anyone else, but I have this inherent hatred of company marketing material disguised as blog posts.
If you are going to write a decent blog post, then write a decent blog post. If people are curious about the author they can look them up (and their affiliation). Don't turn it into a sales pitch.
1. it appears to have been added 2 months after the initial posting rather than a cynical cash grab.
2. it appears the resulting company has pivoted to VPN stuff and is no longer in the log/event processing business anyways?
I like the writeup though I need to do a second pass to fully review everything. The biggest thing that just suddenly hit me was 5 TB/day is really 60MB/s.
It seemed like this is a career summary from his time at Google Fiber (which was killed by google).
It sucks when your baby gets killed for reasons outside of your control. You want those years of your life to mean something, for the work to live on in some form (hence the "Please, please, steal these ideas").
He probably realized after publishing that the post wasn't enough to give him closure so he started a company to keep working on it instead.
(I don't know why I am psychoanalyzing a random blog. I am probably projecting really hard.)
I've added a new update at the end to describe what happened, for anyone who is curious.