Gorilla: A Fast, Scalable, In-Memory Time Series Database (2015) [pdf]
vldb.org
vldb.org
- when storing time series keys, you can save a lot of space by encoding them as a first timestamp followed by timestamp offsets (deltas)
- when storing time series values, you can save a lot of space by realizing sequential data points tend not to be volatile... e.g. a "writes per minute" series is more likely to be 100, 99, 101, etc. than it is to be 100, 999999, 1, etc. This means you can encode the current value by using XOR tricks with prior values to save space
- the authors suggest using in-memory caching for recent data, but doing eventual persistence to HBase or similar distributed filesystems; thanks to this, you can get real-time operational metrics that are fast and space efficient, yet still have a system that scales out horizontally for historical storage
The Morning Paper also did a good analysis here: https://blog.acolyer.org/2016/05/03/gorilla-a-fast-scalable-...
It's actually a delta of deltas, which means for regular time series that a huge portion of rows will contain a 0 (more than 96% according to the paper) because all events arrive at a fixed interval; if there's a slight delay in either direction, the stored value will be small in most cases.
If you're operating on FB-scale, then sure that's what you have to do, but in most cases your database (especially Postgres) is a far superior option.
Time series is "unusual" in the sense that most people don't know first thing about it, event folks with degrees in math/statistics. I think this is why there is the prevailing misconception that it requires a specialized database.
Incidently I've just written a blog about storing TS in PostgreSQL: https://grisha.org/blog/2016/12/16/storing-time-series-in-po...
So a simple round robin scheme just doesn't work.
Good resolution on the timestamps helps avoid graph artifacts and other weirdness, even if you're only collecting them once a minute. Millisecond accuracy seems to be sufficient in my experience, there's other races of that magnitude anyway (network jitter, kernel scheduling etc.).
Yeah, that's a feature provided by specialized TSDB, you don't need to implement yourself.
One day is too short, I regularly have to look at the metrics from yesterday to debug an issue that we noticed today and they're worthless because they've been averaged.
SELECT ...
FROM tv
JOIN customers ON ...
JOIN purchases ON ...Disclaimer: I am a co-maintainer of graphite and try to write code for it when I have free time (so not much right now or for the past year or so)
https://github.com/graphite-project/whisper/commits?author=S...
BTW getting the total point count is as easy as:
select sum(array_length(dp,1)) from ts;
If you adjust the width (i.e. points per table row) to a higher number than the 768 default, as soon as you get above the PG pages size, TOAST (https://www.postgresql.org/docs/current/static/storage-toast...) kicks in, which PG compresses. Not sure how well this works, but at least in theory it should make it even more compact (though not sure about it being as performant).> Time series is "unusual" in the sense that most people don't know first thing about it, event folks with degrees in math/statistics. I think this is why there is the prevailing misconception that it requires a specialized database.
I'm one of the developers of Prometheus, and while I'm going to point appropriate use cases towards Postgres (and regularly do), even relatively small monitoring loads require careful handling such that a traditional database isn't suitable.
An example from previous company is that we had only ~50 machines in one datacenter and yet were up to 30,000 samples per second. This is not considered to be a large setup, which would be in the hundreds of thousands of samples per second.
All the buffering/batching required to make such loads practical requires special design, as naively making each new data point into a disk write (with fdatasync) will not work out - even with SSDs.
And that's just writes. You also need to be able to efficiently read and process back the data efficiently when it is queried.
I'd suggest https://www.youtube.com/watch?v=HbnGSNEjhUc to give a look into how Prometheus does it (which also covers our use of Gorilla). There was also a good post on the InfluxDB blog about the evolution of their design and why the problem is hard that I can't seem find right now.
The "apples" use case is when you want for lack of a better description a Grafana chart out of it. (E.g. the cluster cpu load, "we're getting X page views per second", etc).
The "oranges" case is when every data point matters, and I think that this particular case is not so much about time series, but logging, where a log is a "time series" technically speaking. Here of course you're talking massive amounts of data which you want to do your best to optimize and it's a hard problem. It's not a _time series_ problem, it's a problem of storage. Splunk (sorry, can't think of a better example ATM) doesn't describe itself as a "times series database", and yet it's addressing that very problem of horizontally scalable write-intensive storage. So is hbase, cassandra, accumulo, bigquery and all their friends.
In my experience the best solution is to use two separate systems one for each case. Because there is absolutely no value in a "disk used" measurement every millisecond - aggregated to once a minute (or few) is perfectly fine, and retaining the original data is not needed and you'd be a fool to build a Cassndra cluster to handle it. On the other hand if I'm recording equity bids and asks, well then I better record every point forever.
One is metrics vs. event logging.
The other is consistency vs. availability.
There's many types of logs, and you don't need the consistency that'd be required for billing-related logs as for debug logs. Similarly you may have billing-related metrics (e.g. bandwidth usage) where consistency is important.
> It's not a _time series_ problem, it's a problem of storage.
Technically anything with a time dimension is a time series, so Splunk is a time series database.
> In my experience the best solution is to use two separate systems for each case.
I agree completely. There's at least two general problems here that need different approaches once you get beyond trivial scale.
/disclaimer/ I work for Basho.
5 years ago i was getting that amount of records written on one Oracle node with 2 CPU and a bunch of regular HDDs.
>naively making each new data point into a disk write (with fdatasync) will not work out - even with SSDs.
seems like you're trying to say that this situation requires 30K IOPS. That would be really naive :)
You have 50 writers sending 30K/sec total, ie. 600/sec from one writer. A typical RDBMS naturally batches parallel writes, so a bunch of 3-4 HDDs providing totally 600+ IOPS would easily serve your situation.
If you want fast reads, it'd take at least that many IOPS for a naieve solution where each timeseries had its own block you were appending to.
> A typical RDBMS naturally batches parallel writes, so a bunch of 3-4 HDDs providing totally 600+ IOPS would easily serve your situation.
You could do that if you only cared about writes. There's two problems with such an approach.
First reads would unbearably slow as to read a single timeseries for an hour at say 10s resolution would take 360 operations (one per batch), or around 3.6 seconds presuming 10ms seek time. Wanting to read a hundred timeseries at once over a day is not unusual, which would take 8640 seconds.
Secondly it's not likely to be good disk space wise. As each point is individually written you'd not be able to do intra-timeseries compression, so you're probably talking at least 16 bytes per sample. Probably nearer 100-200 bytes, if you write the metric name each time.
Contrast that with the approach taken by something like Prometheus. We build up 1kB chunks, which hold ~780 samples on average as Gorilla gets us ~1.3B/sample. We batch these up per time series across several hours, so accessing 6 hours is only one disk operation. With the above example of reading 100 timeseries for a day that's only 4 seconds, which is 3 orders of magnitude faster than what simple write batching allows for.
Obviously this is ignoring many details like caching, but the general point holds.
Specialized TS stores exist for decades (process historians, tickerplants, and so on).
Less impressed with the Influx org and their SaaS, but I definitely want to find a usecase for getting stuck into Influx for time series collection
For example Prometheus (which I work on) is great at reliable monitoring and powerful processing of metrics at high volumes, but it'd be unwise to use it for event logging or customer billing.
If you're doing IoT or event logging then InfluxDB might be a good choice for you, though if you're doing more text-based logging then Elasticsearch is nearer to what you're looking for.
https://docs.google.com/spreadsheets/d/1sMQe9oOKhMhIVw9WmuCE... is one comparison of the various open source options.
We initially used Influx, but it could not perform well at the time (0.8). Our events are also heavily label-based. Basically, we do ETL at the time of write, collecting multiple documents into one mega-event, which is a complex, nested JSON document. It may have perhaps 150-200 fields. A single event may be something like "clicked button X". By storing the original document, we can aggregate based on any field value, including text and scalar fields, without having to think about a schema or about planning ahead of time what fields should be indexed or not. ES handles the rest pretty well.
To do the same thing with Influx or Prometheus I suspect we'd have to reverse this and store the document as the labels, along with a single count (1) as the "metric". I don't know how well Influx etc. scale with number of unique label values, though I'd love to find out. The last time I read about this, I think they recommended not going overboard with them.
What's different with business analytics is that the end product is typically multidimensional rollup reports over large time windows (number of page views per customer per web property per month, comparing by 2015 vs 2016, for example), and it's almost all "group by count", sometimes "count distinct" or averages. Whereas "rate per second"-type metrics aren't used anywhere in our app, for example.
Apologies for the email gateway for the video, but you can also see my slides here: https://www.elastic.co/elasticon/conf/2016/sf/web-content-an...
We found that as we scaled it up, we couldn't really keep the data in raw form, so we had to build rollup documents that cover 5-minute and 1-day buckets. Do you use the same trick, or is the number of pageview events for you manageable enough that you just keep it all raw?
My ideal solution would be one that rotated the dataset into historical rollups on a daily basis, so that we only stored the raw data for today, and gradually merged earlier entries at lower granularities. However, I haven't thought much about how to do that with Elasticsearch. I can see a way of doing it by embedding the value in the field label, and using the field value as a count, but Elasticsearch really doesn't like lots of unique fields; you shouldn't be using more than a few hundred at most in a single installation (across all indexes).
/disclaimer/ I work for Basho
My guess though is that Basho would opt to write their own backend off of influxdb's notes..
On a side note I find influxdb's use of nanosecond time precision spot on and Riak TS's millisecond resolution disappointing.
(I work for this company so I'm a bit biased)
The thing is, yes we have a time series database, but a lot of value comes from giving people the tools to analyze, distribute, and act on the data stored (why store data if you can't do anything useful with it?)
We're also working on something with Apache Phoenix over HBase to store TS so that we can query the data in more flexible ways.
https://blog.acolyer.org/2016/05/04/btrdb-optimizing-storage...