Time-series data: Why (and how) to use a relational database instead of NoSQL
blog.timescale.com
blog.timescale.com
Any thoughts on Influx's upcoming release and its reduced memory footprint? https://www.influxdata.com/path-1-billion-time-series-influx...
But once that's settled, for migrating your data, probably the most straight-forward manner would involve outputting to CSV and then using that to re-import to TimescaleDB. Then TimescaleDB would just fit into place where your Cassandra instance used to be.
Do you all discuss event sourcing and how it fits within an overall TimeScaleDB strategy? Anything more than a peripheral concept?
https://docs.microsoft.com/en-us/azure/architecture/patterns...
When talking time series, there certainly seems to be less of a use case for it in the sense that data is mostly immutable with few updates to old records, just new entries as data streams in. That being said, I wonder if there are use cases where the event sourcing concept may bring value to time series DBs - maybe I ingested some bad data that I need to go back and clean up, maybe the structure of my data changed requiring a change in my DB schema, etc. Was just curious if this is something you all have put any thought into at a high-level.
Around if there are any questions.
(Yeah, I know, I'm asking the Internet to do my job for me... thank you, Internet!)
Maybe others can chime in on the first part of your question though.
(Edited to add: This was answered by another reply -- Prometheus has to do a full table scan for things like this.)
Not sure if this is part of the question but if you are querying the event data by something other than time (like a property of the event itself -- maybe the program return code in your example) than prometheus requires a full table scan since it does not have secondary indexes. With TimescaleDB you could add indexes on properties of the event data itself.
Yeah, so on-disk compression is one area where we aren't as competitive with NoSQL column stores.
However, two things to note:
1) Often many of those column-oriented DBs, based on LSM trees, actually need to consume a lot more memory to index all of their disjoint SSTables. So it's a tradeoff of memory vs. disk.
2) There are various things we have on our TODO to test, like just running Postgres on ZFS. We'll write up the results when we do.