Top Ten Time Series DBs
blog.outlyer.com
blog.outlyer.com
Or you can throw in something like [0], I guess. (This thing is still in my todo list though, so I can't tell anything beside the fact that this thing also exists.)
Timescale, Citus, PipelineDB = postgres based but no columnstores. MemSQL, MariaDB, ClickHouse = with columnstores.
* Event sourcing collects events but also often has a notion of a “current” record, meaning inserting a new event requires updating a previous event to invalidate it.
* Financial transactions may involve many append-only tables, but which are linked with data in heavily updated tables, often requiring transactional mvcc.
* Logging events asynchronously or across several nodes often produces events that are emitted out of order which can sometimes span several time partitions, making insertion costly for column stores.
SQL already has window/analytical functions. Relational databases with columnstores also have their traditional rowstores which support all the OLTP features you need, along with easy joins across both table types. Many also now pair a rowstore or in-memory segment with each columnstore for background merges to handle rapid ingest and easy updates.
If you want a polished system, use MemSQL or SQL Server Columnstore Indexes, or more manual work with clickhouse and others. We have 28 billion rows of classic "time series" monitoring data in memsql compressed to less than 50gb and complex aggregations return in milliseconds.
Even if the data took twice as much space, it's still worth it to have everything in a single data warehouse with easy joins and the full expressiveness of SQL.
I feed analytics from my webpage into InfluxDB and it is impossible to compute a histogram of pageload times of the last 10k hits.
Also, check out the grafana histogram plugin. Works great for these scenarios.
In our case, influx+grafana+alert notifications work well.
Yes, the query language needs a lot of work. It doesn't support anything beyond simple queries.
It does seem like an interesting tsdb though, love that it's built on riak.
> Performing queries across billions of metrics looking for labels that only match a few of them (a common scenario with time series data at scale) is really slow in Cassandra. This is because of the way it stores data in columns. This extends to any columnar database including Google's BigQuery which all have a natural disadvantage with time series data.
I've pretty much only heard "columnar database" used as opposed to row store database, and it seems like storing time series data in columns makes much more sense. Could someone clear up exactly how "labels" (which I probably don't understand) are so much harder for column stores to deal with?
Storing labels in a row based system (like SQL) allows querying by value, not column name which takes advantage of all optimizations and indexes making it a lot faster.
That said there is nothing forbidding someone to do both, DalmatinerDB, for example, uses a column-based format for metric values but a row-based format (PostgreSQL) for dimensions.
[1]https://docs.google.com/spreadsheets/d/1sMQe9oOKhMhIVw9WmuCE...
This article is pure clickbait from someone who isn't a serious practitioner in the field.
> Only free and open source time series databases and their features have been compared. Therefore if someone asks “have you tried Kdb+ or Informix?” the answer will be no. They are probably awesome though.
would be nice to see how KDB compares though
While I like all of these databases they don't cover the same spaces.
This blog post is useless
My killer tool for this would be able to talk to PostgreSQL, have a charting backend based on Canvas/WebGL (so able to handle thousands of points rather than dozens), and be easily pluggable to add in new kinds of visualizations.
It has, like, 10x the per machine performance of the others. (See also: COST)