A Fast Lightweight Time-Series Store for IoT Data
arxiv.org
arxiv.org
It's not a bad architecture, and the lockless queuing system and shared memory is cute, but:
- It only supports time indexing. Querying by other fields (e.g. if you wanted to build a time histogram of when you saw a particular event) requires reading the entire dataset.
- It doesn't address replication/distributed indexing, which seems like a must.
- The use of calendar time over micro-timestamps also needlessly complicates things, though presumably it makes their queries more efficient (assuming people query by discreet time buckets like "yesterday").
Few observations here:
* using udp means you don't care if you lose data and that's bad. if your buffers are full your data will be silently dropped, aside from network glitches and metric sizes that can easily exceed 512 bytes. Your buffers can get full when you wait on disk io etc..
Optimizing your network stack when your bottleneck is disk io doesn't buy you much. But the real issue that they haven't tackled is data distribution. If you can horizontally scale your data optimally, then when you hit a bottleneck (whatever it is) you can always increase your cluster size and solve your bottlenecks. In the current solution, If the machine dies, you will lose tons of data.
cow pipelining exist for a very long time now. There's nothing new in what they are introducing.
The persistent data they hold in memory is not guaranteed to be recovered if they crash before an fsync/msync.
Secondary index on time column is not enough. Queries often need particular keys only and you'll end up doing full table scans.
> Optimizing your network stack when your bottleneck is disk io doesn't buy you much.
Agree but also the famous "depends." If you are collecting A LOT of data over the network, first makes sure your network card is capable of handling the incoming traffic. An anology would be a 100M vs 1G port.
Also, if you are running on embedded system device like RPie or Arundio, I think disk will definitely be the first bottleneck before you even have a chance to respond to your network saturation.
After problems with a node, often the management tools cause additional problems.
I suggest anybody follow the cassandra-users list for a while before committing to it.
I am just interested in this kind of thing; I used to work with time series for a German telco; that was DB2 on heavy IBM metal for massive amounts of money and financial work for a startup. I wonder what the state of the art here is.
They also have good compression and buffering support, including catchup functionality.
They are by no means lightweight - usually they rely on PLC's or SCADA system for inputs. They may not offer a complete solution for IOT but I'm sure there are lessons which can be adapted.
It seems it's pretty crowded
For example if you want to store timeseries data long term and already use HBase then OpenTSDB is a good choice.
On the other hand if you want to do monitoring that's simple and dependable in an emergency with querying, graphing and alerting over short/medium-term data, then Prometheus would a good choice.
This in particular, but your entire post, would be a great addition to the Prometheus front page, or top of the documentation section. This wasn't clear to me initially, speaking as someone who evaluated prometheus a few months ago.
Ignoring that as it's planned work, the typical considerations are more around availability vs. consistency and that we're a metrics system focused on operations rather than a event store. Most of that's already covered in our docs. See https://prometheus.io/docs/introduction/overview/#when-does-... and https://prometheus.io/docs/introduction/faq/
> It turns out you can accomplish quite a lot with 4,709 lines of Go code! How about a full time-series database implementation, robust enough to be run in production for a year where it stored 2.1 trillion data points, and supporting 119M queries per second (53M inserts per second) in a four-node cluster? Statistical queries over the data complete in 100-250 ms while summarizing up to 4 billion points. It’s pretty space-efficient too, with a 2.9x compression ratio. At the heart of these impressive results, is a data structure supporting a novel abstraction for time-series data: a time partitioning, copy-on-write, version-annotated, k-ary tree.
That is why I am always looking forward to new things such as this or BTrDB. There is still a lot of room for improvement everywhere.
https://groups.google.com/forum/m/#!topic/druid-user/nqqb5RI...
Druid was built from the ground up to be distributed. It means there is some work to set it up, but once you do, it scales horizontally very easily. Bonus points if you run it on something like mesos making it quite easy to deal with (we do).
https://influxdata.com/blog/update-on-influxdb-clustering-hi...
https://influxdata.com/blog/influxdb-clustering-design-neith...
http://www.refactorium.com/distributed_systems/InfluxDB-and-...
etc. They redid their clustering a few times and the 3rd time around made it closed source. Note that $employer also sues influx for data that we are ok with losing ie: metrics. If the data has a low amount of dimensions, influx is faster than druids. However, if you have 3+ dimensions, druid spanks the pants off of influx by nature of distributing the computation better.