InfluxDB v0.10 GA with hundreds of thousands writes/sec, 98% better compression
influxdata.com
influxdata.com
Keep Rocking !!
Looking at less than a month of issues:
* https://github.com/influxdata/influxdb/issues/5440
* https://github.com/influxdata/influxdb/issues/5482
* https://github.com/influxdata/influxdb/issues/5534
* https://github.com/influxdata/influxdb/issues/5540
If you need something that won't fall over, it's not glamorous and a bear to setup, but OpenTSDB will sail with massive load once you've gotten it running.
One of the issues you linked to was for 0.9.6. Others were there because they had super high tag cardinality and not enough memory to actually run.
When people post comments like "I have 50 million series and it crashes on my box with 2GB of memory!!!", they're not relevant. You shouldn't expect miracles...
There's something different with what you're doing than with what we're testing. If you can give more detail about what your actual data looks like, I may be able to help.
Are you doing this on v0.10 (beta1 or greater) and you're seeing this problem?
Did you folks fix the "oops, I accidentally selected too much and OOMed the machine" crash that was supposed to be resolved by the new query planner in .10? That was my immediate showstopper for InfluxDB for a very high-volume infrastructure, because I didn't feel like proxying InfluxDB just to enforce chunking, and the mere existence of the bug gave me pause. I hit that in 24 hours and shelved the system in 48.
Even pre-1.0, I'm not encouraged when I have to "work around" pretty obvious oversights in reliability, so I put you back on the back burner to give you time to mature and transitioned back to a Kappa-style architecture for my needs. I might still use InfluxDB for low-volume aggregations out of my stream processing, but it's for sure out of the hot path for the foreseeable future.
I think it's common in databases that if you throw a massive query at it that the server doesn't have the resources to handle, things will go wrong. It'll thrash, or crash, or generally have poor performance. We'll be working on improving the failure conditions, but if you put a query to a database that is too big for the server to handle, some sort of failure scenario will be encountered. Just a question of how it's handled.
Also, there's no new query planner in 0.10. We have a bunch of work getting merged in for the query engine for 0.11. But if you throw a huge query at the DB that the server can't handle, you'll still have great sadness.
Which DB ended up solving your big query problems? Maybe we're just a poor fit? Or what kinds of queries should we be working on optimizing?
That means any ad-hoc query can trivially crash your server. That's pretty serious. It's even worse if the ingest is using statsd-influxdb-backend with udp since during the crash you'll lose data.
Do you really think it's ok for a novice user to be able to crash the server from your web gui if they make a minor mistake (e.g something like
SELECT value FROM /series.*/
not realizing how many series their regex actually applies to until it's too late and the has gui stopped responding)?I think it's common in databases that if you throw a massive query at it that the server doesn't have the resources to handle, things will go wrong.
This thinking is worse then wrong. Not only is it wrong - a simple unqualified select won't crash any "common" database - it makes anyone with any "common" database experience doubt your other claims.
This has never ever happened to me in the last 20 years I've used SQL databases. Sure it might take a long time but that's it.
> The current version still has the problem that if you do a massive SELECT then it will fill up memory until the process gets killed. In the future we'll give controls to limit how much memory a query can take.
...
It locks on the query until its killed manually and/or kills it automatically.
It doesn't crash.
We have people that have given us useful troubleshooting information that we've helped (and have helped us improve InfluxDB). Constructive criticism with an offer to help us improve is the best.
While I'd love to trace down every problem for everyone on the internet, my time on this planet is limited.
As for your concerns about if it scales, we'll try to put out more benchmarks and test code over the coming weeks. For reference, we built a tool for stress testing that you can see here:
https://github.com/influxdata/influxdb/tree/master/cmd/influ...
You'll have to make your own decisions.
As a random problem on the internet, this is the workload you should expect from any midsize business and probably something your company would like to target for revenue. Good luck.
Maybe they can sink a big user/customer that's willing to gut through all the pain of helping to make an operation data store.. someone in this thread mentioned mongo and it sounds very familiar.
I'm wondering if I'm being too negative but the flip side of this is that it caused me and others a lot of stress, lost face, and time. Most developers and operators are overworked already so I am trying to save some pain since the blog post makes it seem like everything is fine and awesome.
I'm in no way involved in influx, we're just evaluating it, but this entire thread reads like complete FUD at the moment, so it's odd that you're calling out a maintainer or whoever pauldix is for "disappearing".
FoundationDB was in stealth mode for at least three years before they launched their closed Alpha release, so perhaps this isn't a fair comparison.
(I was an intern at FoundationDB during summer 2012.)
In my case I had a 0.9.4 influxdb setup fed by statsd for about a week. The server hard crashed the first time I tried to back it up.
AFAIK you're using float64 compression scheme from "Gorilla" paper. 2.2 bytes per data point is possible with it but only on data that doesn't utilize full double precision (example: small integers converted to float64). You should compare compression algorithm used by TSM storage engine with zlib or any other general purpose compression algorithm, otherwise this number will be meaningless.
What kinds of queries are you seeing poor performance with? That would help us troubleshoot and improve, thanks.
I do love the simplicity of influx+telegraf, kapacitor also looks cool. Chronograf seems like a bad version of grafana still... maybe in the future if its FOSS and somehow manages to be better than grafana I'd use it.
They (InfluxDB) made huge progress, they work hard on an open source DB (and ecosystem) and people insult them and question Paul Dix's >30 minute response time to your support questions in a HN thread.
Last time I tried it out I vaguely recall that configuring rollups was kind of painful -- lots of nearly duplicate CS queries, even for a relatively small number of series.
And then add synchronized disk commits (like postgresql has always done) and performance goes puff (meaning lower than pg) (example: elasticsearch, mongodb).
We haven't run any comparisons vs. other solutions, but we will do that soon and I expect that we'll be competitive and better in some cases.
Does Todd Persen talk about durability testing in the video or is more about performance benchmarking? I haven't watched yet.
I use postgresql as comparison because antirez did in http://oldblog.antirez.com/post/redis-persistence-demystifie... (and ~none else has done something similar).
The next release will focus on improving query performance, but this one still works for many cases.
1) Does "DELETE * FROM foo" still cause the system to lockup, freeze, and require a restart to free memory? Or are there other conditions/queries that cause the system to become unstable?
2) There's no README/CHANGELOG/dependencies info on the download page. Which is the preferred version of Go to install on my servers for Influxdb - 1.4 or 1.5?
Introduction from last October: https://influxdata.com/blog/new-storage-engine-time-structur...
Is writing less than 20MB/sec of data something to brag about?
Switching to 0.10.0-nightly-614a37c in combination with switching to the TSM engine resulted in a very stable InfluxDB instance. So far my only gripe has been that some queries can get pretty slow (e.g. counting a value in a large measurement can take ages) but work is being done on improving the query engine (https://github.com/influxdb/influxdb/pull/5196).
To give you an idea of the data:
* Our default retention policy is currently 30 days
* 24 measurements, 11975 series. Our largest measurement (which tracks the number of Rails/Rack requests) has a total of 28 539 279 points
* Roughly 2.3 out of the 8 GB of RAM is being used
* Roughly 4 GB of data is stored on disk
This whole setup is used to monitor GitLab.com as well as aid in making things faster (see https://gitlab.com/gitlab-com/operations/issues/42 for more info on the ongoing work).
Unfortunately, I need 2+ instances with Active/Active or failover to seriously consider anything for production which is why I've not touched InfluxDB beyond some light testing.
The input is actually much higher than that. Data points over the network look like this:
cpu_idle,host=serverA,region=uswest value=23.0 1454617920
That's actually a toy example. Most real data would probably have more tags and longer measurement names. Obviously that's much more than 3 bytes.
We persist that to disk in a write ahead log (WAL) and then later we can do compression and compactions on the data to squeeze it down to 3 bytes per point. However, that takes more than a single write against the disk to get to.
Run a load test against it. See how much network bandwidth you can use. See what your HD utilization looks like. My guess is you'll be surprised by what you see.
There are way too many haters in HN. You venomous minority who shit on every bit of good news that isn't yours -- keep your negativity to yourself.
You fucking monkeys infected with rage.
On the other hand, your use of the word 'minority' there was both astute and thoughtful.