45 karma · joined April 11, 2016
The only downside with the jumps is that they will have to load worst case 9999 data points to reach the desired 10000 index of the file chunk. But with most strategies this is no problem and can be fixed by increasing the in-memory lookback window on demand. Having smaller files for smaller chunks would have decreased read performance because of having to switch files too often. So in that regard the read amplification is not a real problem.
And yes, currently the inmemory AHistoricalCache and file based ATimeSeriesDB is only for date keys, but could be made generic if desired (simple pull request).
Anyway a full db server definitely has more things to account for than this custom solution for this specific problem. While the point here is to show that a custom made solution can be a lot faster than general purpose timeseries databases.
Or do you have different opinions here?
Though the original title still was true about the approach in the link being more than 30 times faster than a pure LevelDB solution (which is also often used for this sort of storage).
Any query language, network access, non-binary storage format will slow the data system down, so we had to create our own embedded timeseries system to get this speed.