http://docs.fluentd.org/articles/free-alternative-to-splunk-...
1. Easier to extend that syslog-ng if you have a modest knowledge of Ruby
2. Easy to configure file- and memory- based buffering and failover.
3. Advanced filtering out of the box.
4. Rich plugin ecosystem with 300+ plugins.
At least that's what I've heard from the users who switched from syslog-ng to Fluentd. I am happy to learn more about what makes syslog-ng great since I've never used it seriously myself =)
We've toyed with and pretty much failed using Graylog2. Although it has been coming along steadily in features and stability, we just found that the interface-although pretty-was not intuitive to us: lots of links and multi-click scenarios to get to what you want; and creating filters and streams was difficult and prone to failure.
After watching a couple of very compelling presentations by Jordan Sissel (Logstash founder), we decided to test it out. Once I realized that creating a filter (Grok rocks!) that searched for a term and reorganized the log to my liking only took a couple hours, I was sold.
Another selling point for us was that Logstash has over 2 dozen ways to suck logs in, including the usual suspects - syslog, files, tcp, udp and *mq. You can also perform a bunch of log parsing on the client (i.e. the servers with the logs) before sending them to your central ELK server/cluster.
At the end of the day, there is nothing magical about any of these systems. You alone know your logs best and have to figure out how to read/parse/search them. Our switch to Logstash from Graylog2 was our failing, not Graylog2's.
http://secopsmonkey.com/migrating-graylog2-servers.html http://secopsmonkey.com/migrating-graylog2-servers-part-2.ht... http://secopsmonkey.com/migrating-graylog2-servers-part-3.ht... http://secopsmonkey.com/migrating-graylog2-servers-part-4.ht... http://secopsmonkey.com/migrating-graylog2-servers-part-5.ht... http://secopsmonkey.com/migrating-graylog2-servers-part-6-le...
That's for personal projects, the CloudFlare setup is a little more complex and perhaps one of the data team would be best answering that... if you're interested then I can ping them to see if there's a volunteer for a blog post describing how we do logs at scale.
Would you mind providing more details about why you felt that Logstash was lacking?
Here are some comparisons out in the wild (Note: neither of them is a Logstash or Fluentd maintainer afaik)
- http://jasonwilder.com/blog/2013/11/19/fluentd-vs-logstash/
- https://blog.deimos.fr/2014/05/13/logstash-vs-fluentd/
Also, if you have any question or doubt about Fluentd, please feel free to email me at kiyoto@treasure-data.com
- ELK is great (in fact, we use E+L under the hood) however you really need to know what you're doing with it and you need to spend some time configuring things while putting it together. Do you have that kind of time? Maybe, maybe not...
- Every available tool lets you search and create alerts for monitoring. So, the analysis is always on you. This still takes a lot of manual search time during troubleshooting.
- What if pattern and behaviour detection on your logs can be done automatically? Well, it can be. And that saves you some good amount of time instead of you creating regexs and following the trails to find the root-cause.
I would love to hear your thoughts on automated analysis and anomaly detection on logs. If you're giving a keyword and a specific time frame to search for anomalies, is it real anomaly detection? or a just an improved search?
EDIT: The most fascinating aspect for me is that echofish is more geared towards the actual log entries, rather than statistical analysis, in order to automatically detect anomalies in your logs activity.
Sounds cool.
Echofish is a purpose-built solution for filtering & monitoring of syslog activity. By whitelisting regular messages through the web UI, the administrator can instruct the log processing mechanism to create alerts only for anomalies (irregular messages).
...and actually, it can do lots more once you read the built-in help (such as distribution (using BGP) of IP blacklists, consisting of IP addresses collected through syslog activity).
TLDR; It's gearred towards filtering noise from logs. This also means you can possibly have another daemon reporting network activity through syslog, while echofish can act as your noise-filter.
As much as we absolutely love ElasticSearch for our other indexing needs, we find it quite hard to get the LK-part of the stack to deliver as promised. Kibana may serve up nice graphs and charts, but when you need to drill down into a large amount of log data, we often feel like loosing both overview _and_ detail.
It might very well be that we are to blame, and that we are just doing it wrong (tm) - but I would love to hear how other people are leveraging the ELK stack in production environments?
We log every request (everything but the body usually) and response. If an error occurs, its logged as part of the request. We can practically replay actions taken by users and easily drill down to the exact requests pertaining to an error.
A great little tool is 'since' - a stateful tail. http://welz.org.za/projects/since
since is a unix utility similar to tail.
Unlike tail, since only shows the lines
appended since the last time. It is useful
to monitor growing log files.
It's in the usual yum/apt repos as well as homebrew for Mac.It's a bit hard to websearch for because 'since' is a common word.
Once I get the logs I'm interested in, it's usually a straightforward combination of jq, grep, sed, awk, cut, sort, uniq, xargs, etc. If I need to do some fancy queries on the data, I have another script that will parse the logs and load them into a SQLite db.
1. It requires very little hacking and setup to work (cough ELK cough Splunk) since Sumo is completely cloud-based. It literally takes 2 minutes to sign up, 5 minutes to download and configure some log collectors, then voila you're ready to send data and search. Also I believe you can tell Sumo Logic to grab logs directly from S3 for example if you're running everything on AWS.
2. SumoLogic is pretty easy to use (it's basically cloud-based grep/awk) and has some really cool features that makes Splunk feel clunky in comparison. Parsing and transposing data for graphing is really simple. Also little things like auto-suggesting sources/hosts while you're typing a query makes the experience much smoother than jumping around tabs copy/pasting shit.
3. If you start to generate a lot of logs, and I mean a metric fuckton of logs from 1000s of servers, Splunk Storm will most definitely not be able to help you. In-house Splunk / ELK clusters will need to be carefully sized (just google ELK sizing).
As software developers, we have enough on our plates that it really pays to use tools that help rather than make you wanna throw up your fists and curse. KISS
This comment seemed incredibly positive for a neutral comment, but didn't disclose a relationship to the company (here or in the HN profile, at least as of this writing). The comment seemed odd as the first-ever comment from a 5+ year old HN account.
Meaning, I can send: 2015-02-02 01:00:00 event="Product sold" price=5
And with zero configuration in Splunk, I can now query: event="Product *" price>2 | stats sum(price)
And in the next iteration of my app, I could add 30 more key/value pairs to that message and could query on the rest of them just the same, no configuration. It makes development incredibly rapid to be able to instantly report on any metric anyone on my team logs out, debug-related or otherwise, without having to maintain some master list of every key in every log message in every service we write.
I was floored in my Sumo call today when I was told that wasn't the case in that product yet. It seems like such a basic feature-- and is why many products have switched entirely to JSON-based logs. Have you discovered a workaround, or find that to be as cumbersome as I'm anticipating?
| kv infer "event","price" | sum(price) by event | where price >2
The kv operator refers to key value pairs. There is also a json operator which functions the same way.
Also Notepad++ has a Analyze plugin which i recommend for complex stuff and if you dont like emacs.
Currently handle +100gb/day of heavy document (more or less 100 items per document), on our current setup. And probably designed to handle way mooore.
Dashboard are constantly opened on +50 screens and we use them also to track MySQL, mailing, internal stats...
Moved to EL + Kibana, but not liking the interface yet and it doesn't seem to have 'tail -f' kindof functionality.
- no logs coloration
- a lot of bugs / weird refresh behavior
- no auth
The list of files was saved on their service (rather than in a text-file on the server), and the name of our servers was also guessed by their servers which made it hard for us to add and maintain servers.
I think they should build a better agent that embraces UNIX more, and can be configured through a local configuration. Their platform seems nice, but we weren't able to use it, sadly.
If you have suggestions how to improve it let me know, happy to look at it. Cilent-side configuration an metrics will be available in a week or so.
So the logs being followed are actually configured in a text file on the servers. This makes it super simple for deploying via chef/puppet in large scale environments.
Works pretty damn well.
We use found.on and qbox.io for managed hosting of ElasticSearch clusters.
Splunk (good, but ridiculously expensive)
[1] http://docs.splunk.com/Documentation/Splunk/6.2.1/Admin/More...