104 karma · joined July 13, 2011
https://www.elastic.co/guide/en/logstash/current/plugins-fil....
Edit: typo
If there is log storm coming from misconfigured app depending on your traffic levels the sink servers should give you a lot of space for an on-call engineer to be notified that something is off and adjust / rate limit / drop accordingly.
Edit: udated with more info
The json that syslog wraps data in is actually using the same syslog protocol - nothing is chaging here. It's encapsulated within the same protocol. The only reason to unwrap json is when you need to start structuring your data for BI views, your milage may vary. The key is that these two are de-coupled.
Because you work on a basic level networking / storage layers its very easy to reason and build reliable pipelines according to your needs. As you say — you can't beat that.
PS: The pain starts when you need to deal with multiline logs like java stack traces but then you just work with devs to log everything into json.
Edit: split comment, more info
From my experience if it comes to log shipping from hosts rsyslog + relp + disk assited in-memory asynchronous queues are preferred, most of the time you just only have network i/o as logs would not touch disk.
The idea is to ship logs off the device ASAP as well as destination acts as a sink server capable to handle most of the spikes withouth stressing local source. All done via rsyslog which also wraps actual logs into json format locally. The glue could be syslog-tag.
At the other end you could have ELK stack and logstash using json_lines codec input (pretty fast) structuring data further to your likings.
Just looking into metrics now the avg time for logs showing in ELK is 7-200ms (the latency comes mostly from specific reads happening against the ES cluster).
As ELK is always the slowest component, dropping logs compressed in-memory directly onto disk is also an option.
One thing to note is that RELP can produce extra duplicates which are easily handled by inserting into Elasticsearch using specific document ID (some performance penalty) which could be some unique hash computed on (log content, timestamp, host) etc. With this in place you can also easily "replay" stream of logs to fill potential gaps.
This type of setup scales really good as well.
Edit: typos
I’ve managed to get all refunded 3 x chargeback via credit cards and 2 x paypal. Provided basic proof of purchase. (UK)
I’ve bulk exported generated srt/vtt files from my fav podcasts and using tinysearch that was posted here recently with ableplayer to provide audio full text search of my Jekyll published podcasts posts and with clickable timestamps to audio play of search phrases.
Whenever I want to know what podcaster has to say on specific subject a quick search makes such a difference!
Disabling third party cookies works much better, and for mobile safari using free ka-block.
Any thoughts? Comments?