Working with the ELK stack
engineering.skroutz.gr
engineering.skroutz.gr
But that didn't work. Kibana was unusable. With 75,000 records it never loaded. Not in anyone's browser. So I cut that in half, to 36,000 records. And still, it never loaded. So I kept cutting the amount. And eventually I got down to 10,000 records. Then it loaded, but was so slow no one could use it. So finally I cut it down to 7,000 records. Now it loaded, and it was fast enough that we could use it.
To do the real analysis, I ended writing another script that dumped out the 75,000 records as a CSV file, then I uploaded it to a spreadsheet on Google Docs. This worked fine.
I am curious why Google spreadsheets can render 75,000 records, but Kibana can not? I am also curious what the real use case is for Kibana? If it can't handle large datasets, then its ability to make pretty charts seems useless -- we could never get the data in there to make the chart. I assume that other people will do what I did, and use a spreadsheet instead.
We (at skroutz.gr) use a moderately small cluster (four nodes) and easily process > 4 million rows a day and view them on Kibana dashboards in real-time. We also have success running aggregations on data going back as far as 3 months.
How big was the record for each of the customers? The only thing I can think of was that the records were much bigger than typical.
Just now I pulled up 711,117 records in Kibana (in a jiffy). And the instance this query ran on isn't particularly fast, I should add.
I'm assisting with a ~500GB/day cluster right now (with that number expected to quadruple in the next year or so), and ELK has proven to be an amazingly resilient and flexible tool.
I'm kind of pointing out the obvious, but it's my two cents for whatever they're worth.
Is anybody here successfully using ELK in a low ram environment with low message flow?
Scalyr provides especially powerful features for log parsing and analysis, as well as integrating system and application metrics, and it's wicked fast -- most searches run in well under 1 second. (Disclosure: I am the founder of Scalyr.)
If you're interested, the respective web sites are easy to find, or drop me a line (my email address is in my profile).
I've helped a bunch of financial and security companies implement and deploy elk as a critical infra component. Generally I'm seeing it used to complement Splunk, Alienvault and other log/network/security monitoring solutions.
Happy to help with any questions about deploying and managing ELK in production.
We're running a similar configuration and I'd like to know the limitations before we'd need to start using a clustered setup.
I assume it would reduce the disk usage as well.
I've got a t2.small server on AWS that's taking in a low/moderate logging from three servers (syslog x3, nginx x3, zero-low traffic redis) and it's usually in the 750MB of RAM with absolutely no notable CPU load.
Granted Kibana/ElasticSearch gets pretty much no traffic other than checking it once or twice a day, just to get a glance at 4xx/5xx errors.
So yeah, you just defined my current use-case down to the letter -- centralized logging, fancy interface for filtering/searching/visualizing said log data.
I haven't tried aggregating a years worth of data into a single search, but for reasonable log troubleshooting everything is going well.
Here it is: https://github.com/skroutz/elasticsearch-analysis-turkishste...
There is a ruby version also here: https://github.com/skroutz/turkish_stemmer
It seems like displaying multiple histogram panels may be quicker in real time, given the shorter timeframe, but it would be nice to be able work with something like a month's worth of data without major performance hits.
Not sure if this is something specific to Kibana design, Elasticsearch indexing/search configuration or certain JavaScript engine behaviour.
We had a dashboard that would graph the number of errors grouped by the component over a given span of time. If we saw the chart grow beyond our tolerance, that's when we would break out elasticsearch-head or grep or $tool and try to dig into the details.
Do you have a tool in mind that does what you are describing and thus would supplement Kibana?
I don't understand why usage metrics that are used to calculate business analytics would be logged in a data store that is separate from business transactions and reporting related data. How would you run cohort studies or track funnels?
(Sulks off pining for ELK.)