A Review of Time Series Databases
blog.dataloop.io
blog.dataloop.io
I understand the idea for, say, physics experiments. Lots of parameters sampled thousands of times per second, gotta store that stuff somewhere. But this is HN, not an experimental physics subreddit, so I must be missing something.
What do people on HN who care about this stuff use time series databases for? Why are there 20+ competitors?
If you want to identify issues with good precision, you will record _hundreds_ of counters per each server.
Sticking this kind of data to relational database is not good. With couple of servers you might have millions of _writes_ per second. While the data is only read when you refresh some charts. This makes the time series db unique - they need to handle write-heavy loads, and perform well when asked to retrieve couple of parameters within defined time window.
So two things:
- time series db's are needed for anyone running live systems and measuring them on-the-fly
- time series db's have unique access patterns, not usually optimized by more conventional databases
See example supported inputs by Telegraf https://github.com/influxdata/telegraf/tree/master/plugins/i... from InfluxData
Take, for example, fleet management data streaming in. Thousands of vehicles with time, GPS and other metrics streaming in. There's nothing terribly relational about this data, so there can be better-designed systems to handle ingress and analytics.
And one of the key parts of it is building machine learning models which can do things like look at a customer's previous behaviour in order to predict what their future behaviour might be. So this is one scenario where a time series database is infinitely better than say a RDBMS or Columnar database.
Likewise performing analytics on IoT devices e.g. sensors in trucks or oil/gas equipment requires events to be captured, stored and later mined rapidly. And many time series database being schemaless allow you to manage data from disparate sources in the one table.
I still don't get it.
Most of my background is in "small data" and database programs tend either to store time series in "objects" (other kinds of objects are models and graphs), like Eviews, or as ultimately-isomorphic-to-spreadsheets tables with some added syntactic sugar for time structure that's only understood by some functions.
I'm beginning to do some "not so small data" (text analysis from news websites now; but the essential thing is the time sequencing of information cascades) with sqlite and pandas (pandas is just horrible, but it's already there) and basically the only problem I have is that raw data is unevenly sampled and I have to make some choices when downsampling to a fixed grid so statistical analysis proper can be performed.
That said: I understand, as the grandparent poster, that in some cases (high-end physics experiments) time-structured data is incrementally produced in enormous quantities, and just storing a thousand variables at ten thousand samples per second is this whole challenge.
But server logs? Sensor readings? At one point I had an Arduino hobby where I had a robot moving "of his own will" based on thermistors and did the necessary downsampling and filtering in the Arduino sketch itself for later analysis in Matlab. I mean, who's getting raw data from IoT devices into servers? Even my news scraping thing does some pre-selection before committing to the db.
People underestimate how quickly the signal-to-noise ratio decays at high frequencies; and it seems to me that Cargo Cult IT is massively overestimating its need for CERN-level compute.
Now instead of a robotic hobbyist, imagine if you were a Tier 1 ISP or large banking institution.
What it's not clear to me is what TSDBs (such as those listed by Wikipedia; I looked a bit into InfluxDB particularly) do better than RDBMSes.
This isn't "negative" skepticism, it's an earnest lack of knowledge.
My company deploys wireless sensor networks. Each sensor periodically reports its battery status, which we store in a database so we can see which sensors need their battery replaced.
Recently we found out that some batteries run empty much quicker than expected. The hardware vendor asked us for 'all battery updates for the past year or so'. We didn't save those; each update was just replacing the current value in the database. So it becomes very hard to diagnose this problem because the historic data is lost.
Basically the record response times, error rates, and internal health metrics from all kinds of services and service instances all as time series, and then define rules on top of these to define when to generate alerts.
Other things you typically want to record are all kinds of resource usage metrics, like disk usage, RAM usage, number of processes, network traffic and saturation etc.
I wonder what other extra-dimensional (totally made up term on the fly) data we might care about. Power per query. Temperature per query. Water consumed per query. Water consumed per current activity level of trending topics.
They have however been excellent open source citizens and contributed back improvements, bug reports and suggestions.
> Only free and open source time series databases and their features have been compared. Therefore if someone asks “have you tried Kdb+ or Informix?” the answer will be no. They are probably awesome though.
Article could have been titled "Review of the Top10 FOSS Time Series Databases" since that's what it is.
I agree that any comparison skipping it is meaningless, it is the "standard" in this space, to the extent there is one.
Only for non-commercial use: https://kx.com/2015/09/19/32-bit-kdb-for-non-commercial-use-...
Requiring a schema makes it ill-suited to many use cases since in many situations you don't know the schema ahead of time. And the biggest for me is poor Hadoop/Spark integration. JDBC is not a great choice since you have poor predicate pushdown support and it prevents say Spark from going specifically to the node with the data.
I really appreciate the work Brian Hawkins has done with KairosDB. We settled on KairosDB at work after having a bunch of trouble with InfluxDB. At the same time, it's very much stuck in the Cassandra 2.1 world* . There have been some significant advancements with Cassandra for this particular use case (new storage engine, TWCS) and KairosDB is missing out.
*Yes, it can be run on 2.2.x by enabling the now-disabled-by-default Thrift, but that feels like a last gasp.
"magnum opus (noun) a large and important work of art, music, or literature, especially one regarded as the most important work of an artist or writer."
Considering that this review has a word count of 3708, describing it as a magnum opus (something I associate with a work that takes close to a lifetime to complete) rankles.
Did you include the comments? I didn't myself, but thats sure to bump up the count. We should investigate what the cutoff limit is, maybe 10,000 words would qualify?
I'm struggling weather to classify the work as art or literature, maybe both? Its definitely a seminal piece.
Because those are probably the 10 most popular and well known.
It's really just a "top ten list according to my arbitrary thoughts", which is fine. But it's not really a useful comparison at all.
Does it make sense to use a time-series database if you're not too sure that you can trust the "time" in the series?
For example, let's say you're logging user's locations on a jogging app and using the system's clock as the time record (I understand this may not be ideal). Someone could log a run 2 years from now, for example.
http://www.aymerick.com/2015/10/07/influxdb-telegraf-grafana...
Does he say what the architecture of the platform he was testing this on?
https://gist.github.com/sacreman/b77eb561270e19ca973dd505527...
1 x DalmatinerBD server GCE n1-standard-16 (16 cpu, 60GB memory, 1 x 375G local SSD disk)
I would assume that is the same setup they used for all the other tests, though not clear about that yet
Anyone know the network bandwidth on these instance types?
The other benchmarks are linked in the spreadsheet to their respective details. It's not an absolute, direct, fair comparison. However, we wanted to start somewhere with information available right now and try to collect better results over time as the respective interested parties benchmarked and blogged about their databases.
I'd like to spend more time benchmarking every database in the spreadsheet but it feels like something the project owners should do themselves. I'd probably only get the setup wrong.
but I asked that because I've been using redis a lot lately. I built a real-time analytics using bitmaps and I could use sorted sets to store time series data.
I really wanted to know what benefits do they offer to justify adding something else to the stack...