InfluxDB is now InfluxData, a platform for time series data
influxdata.com
influxdata.com
I really, really wish they'd hire 2-3 more database engineers to focus on the actual timeseries database, not nice looking tooling around it.
Simple things like incremental backups are still not working as expected in recent releases. Useful features from the 0.8 series still haven't been ported. And tsm1 is a neverending effort it seems.
We've went back to Postgres for storing timeseries data since we didn't feel confident to use InfluxDB in production.
I have also not been able to get incremental backups working, have not been able to restore any snapshots, but I have had data survive reboots.
Unless I can get fairly predictable behavior out of Influx the proceess on my systems, it's going to remain in the "unpredictable-so-undependable" category, and far away from "I should pay monthly for this service".
Unfortunately, we were passing that through to the DB which was causing problems. The current release 0.9.6 will return an error.
You can drop an entire series, but you can't yet drop a specific time range. We'll hopefully get that in soon, but we need to get TSM out in production before we start looking at those additional features.
Which features from 0.8 aren't there?
TSM is an effort, but not never ending. We just refuse to put it out there until we've done enough testing on it to be certain about its robustness. We think the community will be happier if we wait to release it rather than rushing something out that may or may not be ready.
That being said, please test out TSM on the 0.9.6 release. I think you'll find that its performance is better than anything we've released previously (including 0.8) and so far we've been able to write far more data into it than in previous 0.9 storage engines.
I'm sorry to hear that you had to go back to Postgres for your time series storage. Hopefully our upcoming releases will give you a reason to take another look.
For me, the lack of HISTOGRAM and DIFFERENCE have actually stopped me doing things I wanted to do.
Obviously, JOIN isn't there either, although that hasn't been an issue for me. :)
We'll also be writing up instructions for how to contribute new functions. We're trying to make the system simple enough for anyone to come in and add them in without having to know all the database internals.
I'll certainly revisit Influx once tsm1 is the default and backups are in :)
Re 0.8: As the other poster noted, I think it was related to histograms and/or derivatives - would have to go back to look.
derivative and non_negative_derivative are in the current release.
Hopefully 0.10.0 will convince you to come back over :)
I had no idea InfluxData forked Grafana and developed their own (currently closed-source) visualization tool named Chronograf. It looks like a InfluxData branded version of Grafana. Also interesting is that it's the only closed-source element of their TICK stack.
Seems like there could easily be a conflict of interest here, such as new features of InfluxData only working with Chronograf and not Grafana.
grafana - http://grafana.org/
chronograf - https://influxdata.com/get-started/#visualizing-data-with-ch...
We're fans of Grafana and want InfluxDB to work well with it. However, we also want the ability to change things over time and possibly take a different focus. Over time we'll be pushing Chronograf to be usable for people that fall less and less on the development side of things.
It's still early stage with Chronograf so we're still iterating and testing different ideas.
Just curious, what are some reasons why one would choose Chronograf over Grafana right now?
Our longer term goal with Chronograf is that it'll help with doing ad-hoc data exploration and, more importantly, give developers an easy tool to quickly build visualizations for their applications.
We're currently filling in the basics of what we believe to be a strong dashboard tool, that it makes it easy to build queries, create graphs and put them in a dashboard in less than 15 min.
We're at 0.3 and still implementing basic features. Where we differ in product is our approach. We see dashboarding, and building graphs as a design problem. What are the issues we're trying to solve? How can we solve them better?
Main differentiation right now is the query builder which makes it easy to prototype influxql queries and see data. Also, our dashboard user flow is pretty quick to go from no graphs to full dash. We will add more visualization options later. Right now, it's a really good tool for mvping a dashboard.
We're still figuring out the next steps, like more features and functionality to make it a more complete tool. My main concern is if we're solving the right problem for the right user.
If you'd like to give us your thoughts as a user of dashboards we'd love to hear from you.
Email us at Chronograf@influxdata.com
I've been experimenting with collectd+influxdb+grafana, and have been pretty impressed with how easy it's been to build a bunch of metrics.
The addition of Kapacitor seems promising- alerting was a piece of the puzzle I was missing. Some people were using Sensu, but the grapevine was giving it less than stellar reviews.
(just curious, I've tried both collectd and telegraf and telegraf has been much easier to setup and control imho)
In the long run, the full ecosystem development here will pay off.
Basho, the maker of Riak has also moved from being a db company to being a data platform company. And they've also launched a time series version of Riak.
How are they moving to be a data platform company?
There's a lot of mention about the ephemeral nature of time series data, but how is InfluxDB for storing permanent time series data? Are there any performance limitations I should be thinking about (aside from the limits of the hardware itself)?
https://influxdb.com/blog/2015/10/07/the_new_influxdb_storag...
We use them for storing "permanent" data. They slice up the data into weekly files, so if you don't access it often, it's no biggie to store a lot of data. If you constantly access data across time, you may run into performance degradation from a caching point of view (having to pull in the files from past, blowing caches of new stuff constantly).
The ephemeral nature of time series is more about the fact that metadata and the raw data are constantly changing. For example if you use a round robin database (like Graphite or RRD) you have to pre-allocate the file based on how much of the series data you're going to have (e.g. 10s samples for 3 months).
Having series that are ephemeral and may exist for an hour or a day make this much harder to deal with. Our storage engine, the Time Structured Merge Tree, is designed with this ephemeral nature in mind. So it's able to store this data efficiently.
The 0.9.6 release that came out today has that storage engine in it for testing. We'll release it for production use at the end of January after more testing.
I'm particularly interested in Telegraf. We're using Prometheus for data collection and monitoring right now, and collection system is definitely the weakest part.
Telegraf can both read and produce Prometheus metrics, which is something I've worked on as I believe metrics shouldn't be locked into any one ecosystem.
First, Prometheus prefers that you create an HTTP API for metrics. This means having to manage running daemons, set aside ports, and so on.
The fact that it's not pluggable leads to an ecosystem where every "exporter" is HTTP-based and needs to be run and maintained this way. It's not lightweight or ops-friendly.
So for our own stuff, we built a single pluggable collector that simply emits text files (scheduled via cron), and uses prometheus-node-exporter's "textfile directory" support to export the data. It's not ideal.
And since it's a kind of mini-framework itself, that doesn't follow any community-defined plugin system, we can't easily open-source the bits we think would be useful to others (we have some nice metrics for RabbitMQ, ElasticSearch, PostgreSQL, etc.) because people would have to buy into our little plugin system.
I'd much rather that the exporter was pluggable, and was capable of spawning subprocesses.
Prometheus also works very poorly/not at all with existing web frameworks such as Ruby's Unicorn and Node.js's cluster module, where there isn't a single process that can track metrics. So far, the attempts to accomplish this with Unicorn, for example, seem unsatisfactory.
We're fans of Bosun and actually contributed code to the project to get it integrated with InfluxDB. However, we wanted something that was purpose built with the InfluxDB schema in mind, which doesn't map to any existing time series solution.