TimeSeriesDB
github.com
github.com
As a consequence of being two different things, storing and analysing them efficiently need different solutions. I've seen people stuff everything in Elastic Search and it's certainly possible if you really want to.
I think the thing is that the primary use-case for a lot of people combines the two - eg they want to log a whole key:value map where values can be numeric or string, and then dynamically generate time series. We can do this with SQL, like if I want a graph of traffic per domain name, "SELECT time_bucket, domain_name, sum(bytes) FROM access_log GROUP BY time_bucket, domain_name" - but all SQL servers are totally non-optimised for this use case.
Surely there must be some software out there which can do this efficiently?
The more pertinent difference is event logging vs. metrics.
For the former the ELK stack is popular, and Prometheus.io is suitable for the latter.
I found that by trying to see if it could be used as a metrics DB but stumbled upon a few issues like configuring retention times per target. See for example https://github.com/prometheus/prometheus/issues/1381
Prometheus has lots of potential but it's not a metrics DB at this point or in the near future. I just wish they'd made that a bit clearer on the page.
"We make design decisions that presume that Promtheus data is ephemeral, and can be lost/blown away with no impact."
That viewpoint pretty much limits them to be an ops tool. I wish they'd reconsider this point.
Any query language, network access, non-binary storage format will slow the data system down, so we had to create our own embedded timeseries system to get this speed.
The only downside with the jumps is that they will have to load worst case 9999 data points to reach the desired 10000 index of the file chunk. But with most strategies this is no problem and can be fixed by increasing the in-memory lookback window on demand. Having smaller files for smaller chunks would have decreased read performance because of having to switch files too often. So in that regard the read amplification is not a real problem.
And yes, currently the inmemory AHistoricalCache and file based ATimeSeriesDB is only for date keys, but could be made generic if desired (simple pull request).
Anyway a full db server definitely has more things to account for than this custom solution for this specific problem. While the point here is to show that a custom made solution can be a lot faster than general purpose timeseries databases.
Or do you have different opinions here?
Though the original title still was true about the approach in the link being more than 30 times faster than a pure LevelDB solution (which is also often used for this sort of storage).