Time series database Graphite seems to be falling into disfavor
vividcortex.com
vividcortex.com
http://obfuscurity.com/2015/11/Everybody-Loves-Graphite
Personal response:
I've used Graphite, OpenTSDB, Ganglia, Cacti, and a bunch more solutions. Recently, work transitioned from OpenTSDB to a hosted solution from a startup called Wavefront.
This has been a smash hit.
Scale matters. If you are a small shop, and can workaround the atrocious UI shortcomings of Graphite by all means, go for it. As you start to get larger, OpenTSDB looks attractive. We went too far down that path but were able to quickly (< 3 months) transition over to Wavefront and haven't looked back.
The proprietary vs open argument is less important then the argument over where the data is physically held and ownership of it.
There are finally some self-hosted commercial options in this space, but we've already built our solution and it works well.
https://github.com/akumuli/Akumuli
It's a standalone solution without dependencies on other services or databases. It can handle more than a million data points per second and uses fixed amount of memory to store data on disk but without Graphite's shortcomings (compression is used for everything and timestamps is stored with nanosecond precision).
I'm focused mostly on realtime time-series analysis. Akumuli has built-in anomaly detector: EWMA, Holt-Winters and sketch based methods. It can generate different time-series representations, like SAX and PAA. DTW and correlation search is in progress now.
My issues with the present state of TS isn't the volume. I was looking for using TS outside of the DevOps world. Everyday things like your heart rate over time, price of gas at the nearest station, number of people in line at your coffee shop, etc. All these are interesting, and Graphite/RRDTool/InfluxDB/etc did not seem like appropriate storage, because to use it with your other data (which is most likely in a relational DB of some kind) you need to export/import it and who wants that.
I call this problem "data seclusion". When data exists in some kind of an incompatible format (e.g. Whisper files), it will end up ignored because of that extra conversion step necessary to link it with your other data. Data in Graphite and such is mostly good for generating charts, but TS analysis is so much more than that, even at its simplest.
I think that the good old relational database is fine storage for TS and we gave up on it way too early, especially given what's new in PostgreSQL. Making it horizontally scalable, distributed, using consensus protocols, etc - these are not time series problems, these are database problems and we do not yet have a good solution for these. (We have many that "kind of" work, support some features but not others e.g. Cassandra). Projects like InfluxDB are mired in solving the wrong problem which will eventually get solved at the DB level.
More thoughts on the subject: http://grisha.org/blog/2015/03/28/on-time-series/ http://grisha.org/blog/2015/09/23/storing-time-series-in-pos... and http://grisha.org/blog/2015/05/04/recording-time-series/
I recall that Kibana is a plugin for Elastic Search but don't know about the others. Isn't it possible to connect a hypothetical standalone visualization frontend to any TSDB?
The TSDBs I tested when I was building this setup (Graphite, OpenTSDB and InfluxDB) all had some sort of basic plotting facilities, but they were extremely minimal and all pointed towards Grafana if you wanted something more substantial.
Have you documented your setup anywhere?
I'm DIY-ing something similar and I'm always curious how others have approached things.
While it is used for machine learning, it actually makes a lot of sense to follow a similar approach for more general time series applications.
That's making some presumptions about the time-series database use case. SAX is convenient for storing and retrieving the data as well as identifying trends or recurring behaviour. What more do you need?
Yep. Each SAX word should be mapped to the list of seriesid:timestamp pairs. This list is often referred as postings list in information retrieval. The resulting data-structure is an inverted index. SAX and iSAX papers describes inverted index variant (really bad one) based on folder structure but one can use convenient IR tools for this.
The thing is, particularly with time series data, a lot of times it is sufficient to at least start with the summary data in the index.
We're also in the process now of finding a good storage solution for time series data at my dayjob. We're not storing any server metrics, but more personal user health data and related metrics. So we want the flexibility of being able to write new metrics on your personal timeline without altering a schema. Most reads would be querying single users timeline to fetch their data and deliver through APIs for rendering in clients. Also of course to do analysis on all users timelines, find correlations etc. but those queries are less frequent and not so time critical. I'm starting working on a prototype now with MongoDB. Other DBs that have come up has been InfluxDB, Riak TS, Amazon DynamoDB and probably some others I don't remember. Haven't actually thought that much about Postgres, but thanks for your links to your blog posts I will read up on how Postgres might work. Most other non-time series data would still be in our MySQL setup.
You know, that sounds like exactly the kind of thing you'd want to do in a DevOps world (leverage your data and new ideas to predict problems).
So, you're not just right that we need to look at timeseries because there are many interesting new problems that we might solve with them, but also because there are still many problems we want to solve in DevOps.
It's kind of now accessible to the masses (previously you'd have to pay SAP a few million for BusinessObjects then EMC another million and a half for enough storage to setup a DW), but open source tools have been available for at least a decade. Now it just has a fancy new marketable name "predictive analysis" that Oracle et al can charge an extra couple million for with their price-gouging RAC licenses.
For that reason we still use graphite's API, but not the UI or datastore. Whisper nor whatever the newer one is would survive the load we put on it.
Our setup looks like:
grafana -> graphite-api -> graphite-influxdb -> InfluxDB <- statsite (C impl of statsd) <- metrics
I have Grafana 1.9 pointed directly at InfluxDB 0.8 and here's one thing you _can't_ do: group a series by X, then plot only the top Y groups. Is that the sort of thing that using graphite-influxdb provides?
Not so much for the design (oh look it's white instead of black, who cares), but the way they've handled it subsequently.
The 3.x -> 4.x transition for Kibana left a product that was missing really basic features (like, the ability to set graph colors for one - a bug/feature request that's been open since last january).
As it stands, upgrading from Kibana 3 to Kibana 4 is a step backwards. You lose functionality rather than gaining it.
They also decided to require a major version bump to Elasticsearch (to 2.x) with a point release of Kibana.
I used to be super optimistic about the Elastic guys, but some of these decisions are just head-scratchingly awful.
This might be derailing the thread a bit, but is there any log management platform like ELK or Splunk that has an expressive and versatile query language like Splunk's? My biggest issue with ELK is that analytics is mostly expected to be done through Kibana's GUI, while with Splunk you can craft terse queries to do almost any sort of transformation and visualization imaginable. I don't like how ELK is so GUI-oriented.
https://www.elastic.co/guide/en/elasticsearch/reference/curr...
There is a Python implementation that makes creating complex queries pretty easy:
https://github.com/elastic/elasticsearch-dsl-py
I agree with the criticisms of Kibana, but I have had no problems querying Elasticsearch directly. It also supports scripted queries if the built-in aggregations aren't enough.
Of course, then you have to build your own visualizations with the results...
Using one of their examples, this:
{
"query": {
"filtered": {
"query": {
"bool": {
"must": [{"match": {"title": "python"}}],
"must_not": [{"match": {"description": "beta"}}]
}
},
"filter": {"term": {"category": "search"}}
}
},
"aggs" : {
"per_tag": {
"terms": {"field": "tags"},
"aggs": {
"max_lines": {"max": {"field": "lines"}}
}
}
}
}
would be the following Splunk query: title=python description!=beta | stats max(lines) by tags
It would be nice if there was some kind of query compiler that could generate ES JSON from an expressive query language. s = Search(using=client, index="my-index") \
.filter("term", category="search") \
.query("match", title="python") \
.query(~Q("match", description="beta"))does exactly that with a reasonably good subset of SQL and ES query language mixed in.
It operates in 2 modes; in one it runs the query, in another it spits back out what the equivalent ES JSON query is. We use this as a quick prototyping tool and modify the ES query as needed, as most of us here still "think" in SQL for a lot of things.
I have had some decent luck with sending syslog into mtail and counting generic word events like 'error' and 'warning' as a way to do "something might be wrong" alerting.
- OpenTSDB
- InfluxDB
- Graphite
- Elastic (Expects to be populated by logstash)
OpenTSDB was the original time series database so that has some extra UI features compared to others for graphing. My hope is that InfluxDB will mature to the point in availability that we can use it and ditch the hbase dependency. But Currently OpenTSDB is the best option for us.
It has a dimensional data model, a powerful query language to go with it, and covers aspects from instrumentation to storing data, all the way to alerting and dashboarding. The latest version of Grafana has native Prometheus support now too. Many tools (like Kubernetes or etcd) already export Prometheus metrics natively, so you can monitor them with Prometheus right out of the box. Support for many kinds of service discovery (Kubernetes, Marathon, EC2, Consul, ...) make it work very well to monitor dynamically scheduled services as well. Disclaimer: Prometheus author.
edit: yup, the prometheus docs indicate InfluxDB is better suited to this use case.
http://prometheus.io/docs/introduction/comparison/
Aside, I really appreciate this comparison page and the prometheus docs in general are well done.
1) it is simple to set up!
2) nice to have a fixed-size flat file as database, although performance degrades very quickly. Cache misses is too frequent.
3) over raw socket is also quite attractive
No other solution can compete with Graphite for its simplicity. But there is so much more than just sending data to graphite from collectd, or your custom Python program....
Whenever I see OpenTSDB I sigh. Do we really need H technology here? Yeah I work for a small shop but do I really want to maintain OpenTSDB....when I already have so many databases? I am choosing between Cassandra and PostgreSQL for my TSD.
The Graphite URL API was visionary, and still makes the cut in most common use-cases.
It is kind of amazing how hard it is for existing time series database systems to track the changing needs of the marketplace.
Btw, this article was heavily flagged. It's not really legit to flag a story just because people don't like the title. Plenty of good stories have problematic titles. Depriving others of a chance to read the content, especially when there's a good discussion going on in the thread, is a bad use of flagging power.