InfluxDB Clustering Beta and Data Explorer
influxdata.com
influxdata.com
Like many other users here, we were disappointed about paid clustering, but when the original press release said $400, we were willing to wait it out and see. However, we ultimately decided to go a different direction after seeing they wanted $20k+ to run clustering on a 256GB node. We ingest 10s of millions of data points per day from IoT sensors, and expect our data size to far exceed that capacity.
That said, we plan to run our own Cassandra cluster w/ KairosDB (http://kairosdb.github.io/) acting as a read / write abstraction layer. It'll cost us about $11k to run the cluster for the year, with 3 nodes @300GB/ea., leveraging Cassandra's (free) and open source clustering, HA, and replication technology.
Also, Brian (project leader of KairosDB) agreed to a small consulting project for us to help us configure and tune Kairos. He's a very nice guy and was very helpful.
Though I wouldn't say it was a smooth transition. I started with 0.8, IIRC, and while it worked ok it used an amazing amount of storage. 4GB for a year worth of graphite data blew up to 100GB for a month of InfluxDB.
I gave up on InfluxDB a few times during the process, but at 0.11 I tried it again and is has been pretty good. We are only putting the Telegraf data and one small service statistic in it, but the storage is pretty reasonable at 12GB for a few months of data. Querying and graphing the data with Grafana is great.
If you have tried it before 0.11, definitely try it again. The guys giving a Prometheus talk at PyCon were really down on InfluxDB, but they hadn't tried it for 6 months. I was like "Yeah, it was unusable then". I wanted to like Prometheus, but I just couldn't figure out how to feed my data into it.
Really curious as I am looking forward to set up InfluxDB (moving from graphite too). I was going to use collectd, but your comment makes me wonder if it's the right choice.
It's really hard to compare and say better or worse than graphite, because it isn't an apples/apples comparison. I, so far, haven't figured out how to do roll up like graphite has built in, so I think I'm holding onto the high precision data. So that makes it not even close to a fair comparison. Still though, 12GB for our current data set feels reasonable. Back in 0.8 when it was more like 150GB for a month, that was not gonna work.
For a while I was feeding collectd into InfluxDB. It didn't really produce anything I could use immediately, and I haven't gone back to revisit it.
I don't blame them for going down this route (we all have to eat), but I think had that descision been made earlier it would have looked better on them.
Hosted services is a viable business model. The risk here is to ensure that your product, if successful, cannot simply be deployed by Amazon. Dual licensing has been one response to the Amazon risk.
Admin tools alone are also thought to be difficult to monetize at sufficient value to grow a large business.
While DataStax has a nice UI, they have built a data platform with Cassandra the core, much like Cloudera has built a data platform with HDFS/HBase at the core. Personally, I view the data platforms as prepackaged and productized SI (which is very valuable..
https://docs.influxdata.com/influxdb/v0.10/guides/hardware_s...
Show stopper if you're in the same boat that I am.
I compared a dataset to postgres and the new mongo storage engine is very good -- dunno if wiredTiger is something you guys are looking at, but it's very compact and fast to query with the proper indexes.
Dumped some stats into a gist as well: https://gist.github.com/mattbillenstein/89969980025414e2bca8...