Time-Series Database Requirements
xaprb.com
xaprb.com
Just thought I'd throw this out there since its a specialised area that not many people know about. I've done some work with them in terms of writing adaptors to a time series data visualisation product.
http://en.wikipedia.org/wiki/Operational_historian
http://en.wikipedia.org/wiki/OSIsoft
https://www.honeywellprocess.com/en-US/training/programs/adv...
http://www.geautomation.com/products/proficy-historian
On the topic of Historians vs Relational Databases, theres a blog post here about it ...
https://osipi.wordpress.com/2010/07/05/relational-database-v...
... admittedly this is by the developer OSISoft so it may be biased, but their points seem valid. Especially the swinging door algorithm reference and the fact they are far more efficient in storage.
Writes are faster than read, it's an AP, and you shouldn't really update frequently it unless you want tombstone hell. There's also TTL too.
Is there any cons of using Cassandra as a Time Series Database? I'd like to hear it.
The biggest thing for Cassandra is you should know your queries before hand before you data model.
It is what we use, and we use spark streaming for the rollups. We had evaluated Influx, OpenTSDB, and Druid also. So long as you know the exact read patterns for your client I think Cassandra is definitely the best fit for most things.
From what I read (so not confirmed) one problem is that it uses space inefficiently. Since I predefine the columns anyway, they might as well use an efficient storage instead of mongodb style kv-pairs.
I am also not convinced of the partition key necessity (although it doesn't hurt either once you got it).
Finally, since my application runs on the JVM, I'd actually like to see a direct integration / an API that allows me to skip the socket overhead and launch cassandra directly on start-up. The advantage is mainly memory (just need half of it) and latency although I agree this is more difficult to maintain in a distributed scenario.
This may be my experience as a Googler talking, but I also somewhat disagree with the notion that the data can't fit in memory. I operate a X0,000 task service, and our monitoring data could fit into memory in a large server, if need be.
Of course, we don't keep per-task data permanently, that would be prohibitive at a 5s monitoring interval like the one we use, even if it were put on disk. Instead, we accomplish what I described by aggregating away some dimensions, mostly task number, and then holding the aggregated series in memory, for fast queries. There are some nuances, particularly around having foresight over the cases where you do want to see individual tasks.
But suffice it to say that this person's experience does not match mine in terms of what I need from a TSDB. Perhaps his ops background comes from a different set of needs than mine, but if you're building a TSDB for many customers, I wouldn't take this list as gospel.
As the little girl says in the GIF: why don't we have both? Write to a write-optimized store of limited size that requires full access during reads, and re-write that into a read-optimized format hourly or daily. Because it's limited in size, you won't care that the most recent data isn't very efficient for reading, or isn't particularly compact.
There's also OpenTSDB (http://opentsdb.net) that's been around for a while.
Or, in short, if you do need scaling you probably can afford to implement it.
The license is commercial and there's a free CE version which can be scaled vertically without any throughput constraints. Tags are supported for series as well as for entities and metrics to avoid storing long-term metadata such as location, type, category etc. along with data itself.
I wouldn't be surprised if functional differences between TSDBs and historians will disappear in just a few years. Right now the historians are good at compressing repetitive data at source and on disk which makes sense given their heritage in archiving data from SCADA systems.
Using 1KB chunks of datapoints as simple K/V?
Or using Large Data Types [1] like Large Stack [2]?
Large Stack
A Large Stack collection is naturally aligned with time series data
because the stack preserves the insert order.
Stacks provide Last In First Out (LIFO) order,
so it's a convenient mechanism for tracking "recent" data.
Usage examples include:
Show me the last N events
Search the last N actions with a filter
Show me all documents in reverse insert order
[1] https://www.aerospike.com/docs/guide/ldt.html[2] https://www.aerospike.com/docs/client/java/usage/ldt/ldt.htm...
[0] http://youtu.be/EoUfkkrIbPg
[1] http://www.slideshare.net/vividcortex/scaling-vividortexs-bi...
MySQL by itself is plenty fast for us. As Baron mentioned, we don't even run the database servers at full capacity.
host.queries.type=GET.result=200.tput
and host.queries.result=200.type=GET.tput[HDF5](http://www.hdfgroup.org/)
For example, only one process at a time can open an HDF5 file for writing.
[2] http://martinfowler.com/eaaDev/EventSourcing.html
[3] http://www.se-radio.net/2015/02/episode-219-apache-kafka-wit...
Also, as another poster said, it has no indices, just an offset pointer.