With TimescaleDB, would I use the recorded_at or inserted_at column for the hypertable?
Does this change if data for an individual sensor can sometimes arrive out of order? If the sensor malfunctions and the data contains timestamps in the far past or the far future does this cause issues with TimescaleDB?
What we've done in postgres so far is have the tables with data generally structured around the recorded_at column because most analysis wants to look at the data "in order" . to generate reports, graphs, etc. Each data row also contains a "payload_id" relating it to a "payloads" table which helps group data by when it actually hit the system. Data processing has generally been built around the payloads and then query any additional data in recorded_at order on the main data tables if we need to look back or forward in time.
Out of order data should be handled fine by TimescaleDB -- if you do have data that is far in the future or in the past, you may get stray chunks to hold those, but it's not going to create all the intermediate chunks or anything that might be undesirable. You can later correct those fields by deleting and reinserting the record with a corrected timestamp.
I've looked into timescaledb for this, and it doesn't support them.
However, the intersection of bitemporal indexes and columnar time-series queries seems important and yet I haven't seen anything that looks like it might offer both, possibly asides from kdb+ and SAP HANA.
Disclosure: I work on https://github.com/juxt/crux (which is optimised for bitemporal graph joins and doesn't currently employ columnar indexes)
Question: how does the new compression compare to TokuDB, and is it tuneable for performance/size tradeoff?
https://www.percona.com/doc/percona-server/5.7/tokudb/using_...
[1] https://www.oracle.com/technetwork/database/exadata/ehcc-twp...
Quick scan of the Oracle paper couldn't find specifics, other than something like this:
"Warehouse Compression provides two levels of compression: LOW and HIGH. Warehouse Compression HIGH typically provides a 10x reduction in storage, while Warehouse Compression LOW typically provides a 6x reduction"
That would at least suggest that they aren't doing anything type-specific like we are.
This also leads to significant query performance settings if you common filter by device_id, for example. Which are super common in time-series workloads for IT monitoring / devops / IOT / etc.
* What is a plan for PG12 PLUGGABLE STORAGE[1]? this is for PG12?
* Can you compare with Zedstore[2]?
One of the interesting things of our technique is that it doesn't require low-level changes to Postgres, and actually then works with any version of PG that TimescaleDB supports (currently PG10, 11...PG12 coming soon).
That said, we're excited by the work Postgres has been doing with pluggable storage, particularly how PG13 will further open up possibilities such as Zedstore, and look to see how we can then marry some of these ideas. In terms of feature-by-feature comparison, haven't yet dug enough into the details.
Aside, one interesting aspects of our approach is discussed in the article: Having a "hybrid" row/column, rather than purely columnar, can actually be beneficial for many time-series workloads that constantly query very recent data (e.g., for dashboarding) as well as to improve ingest rates (although some column stores do build a temporary in-memory row-based cache before batch writing a column).
If you are interested in primarily showing the recent data -- which you see in many monitoring examples in IT/devops or IoT -- you often want the raw data or an aggregate per time period.
As an aside, TimescaleDB introduced continuous aggregations in v1.3. Here's a nice example of using it with Grafana for dashboarding: https://blog.timescale.com/blog/how-to-quickly-build-dashboa...
How well could be apply the same ideas for in-memory processing?
If I understand correctly, you have something alike:
- Store each column on a array of N=1000 - Store the group of columns in pages, with metadata of ranges of keys to locate rows in the adequate page