In particular, I'd love to know if theres anything major that generic RDBMS's could do better here.
In particular, I'd love to know if theres anything major that generic RDBMS's could do better here.
What "time-series" databases offer is more functionality around time being a primary component of the data. For example, automatic partitoning/sharding on time, various date handling techniques, better bucketing/gap filling/smoothing functions, data retention policies based on time, automatic rollups and aggregations, etc.
Some have custom data stores, some use key/value stores like Hbase or Cassandra, and others use relational databases. Using relational foundations offers more flexibility (like Timescale on top of Postgres) than the others like InfluxDB or OpenTSDB.
For example, double delta compression, or Gorilla for floating point numbers. For more, take a look at our open source VictoriaMetrics database, which uses all of such tricks.
- Availability preferred over consistency (you want your metrics when bad things are happening, even if they may not be 100% accurate)
- Extremely write heavy, with most data never being read (they mentioned that only 2% of data is ever read. in my experience at another large company, it was way less than 2%)
- For the data that is read, most of reads are for most recent points (mostly alarms, but some dashboards as well)
- Different SLAs for queries based on usage - ie. queries for alarms must be fast. Dashboards, and trend analysis - not so much.
If I didn't want to use MySQL or Postgres, I'd rank Prometheus #1 and Clickhouse #2.
The killer thing for Clickhouse is that Percona supports it, so if you want to outsource the installation, mgmt. and support, you can just write a check and get good results.
Also, Clickhouse is a column store with SQL, so you could use an instance for monitoring and another to replace Vertica or Greenplum or whatever so long as it has the client libraries you need.
I'm concerned with there being a support issue, as well as a smaller user base also affecting support in the long term.
Time-series databases make specific architectural decisions and introduce advanced capabilities that enable orders of magnitude better insert/query performance, while also reducing storage cost via compression.
For example, TimescaleDB (where I work), is a relational time-series database built on top of Postgres, which means it includes all of the goodness within Postgres, but also achieves:
* 96%+ compression (ie only uses ~4% of the storage of Postgres) [1]
* 100x-1000x faster queries than Mongo [2]
* 10x higher inserts and 50-1000x faster queries than Cassandra [3]
If your RDBMS is good enough - then please keep using it. :-) But if query latency, insert performance, or cost are becoming concerns, then I'd suggest looking at a relational time-series database.
[1] https://blog.timescale.com/blog/building-columnar-compressio...
[2] https://blog.timescale.com/blog/how-to-store-time-series-dat...
[3] https://blog.timescale.com/blog/time-series-data-cassandra-v...
Time series DBs (and OLAP dbs in general) have very different trade-offs/needs than transactional DBs.
Not sure what you mean by this. TimescaleDB outperforms InfluxDB, another purpose-built time-series DB, on most workloads [0], especially on ones with high-cardinality [1].
In particular our key insight, which some may still find heretical and hard-to-believe, is that it is quite possible to produce best-in-class performance characteristics for time-series using a relational database. In particular, we have been able to add columnar compression to our row-oriented format resulting in 96%+ compression rates [2], and multi-node scale-out resulting in 10M+ inserts per second [3].
TimescaleDB is not designed for workloads where most queries touch all data points (ie full table scans). But that's not what time-series workloads look like.
[0] https://blog.timescale.com/blog/timescaledb-vs-influxdb-for-...
[1] https://blog.timescale.com/blog/what-is-high-cardinality-how...
[2] https://blog.timescale.com/blog/building-columnar-compressio...
[3] https://blog.timescale.com/blog/building-a-distributed-time-...
Well, everybody with experience outsources monitoring now since it's a non-core cost center, unless there's a compelling scale or secrecy issue.
If RAM and CPU were free, I'd use MySQL or Postgres w/partitions because of their mgmt. features, tested replication and SQL.
But Prometheus or Clickhouse are 10-25x more efficient in terms of space, and often have much faster queries. The tradeoffs are bizarre HA gaps, lack of trained people, and ops groups are stuck supporting it.
I would never recommend monitoring with anything based on HDFS (OpenTSDB), written in Java (Cassandra), or in-memory for large clusters (InfluxDB.)
For monitoring under 200 nodes, anything will work.
If you only have a day to do something, just install Nagios and you'll get 99% of what you really need.
Source: DBA.
That has not been my experience. Quite the opposite every place I’ve been that outsourced monitoring ended up bringing back significant portions of their observability stack either for needing more control, different feature sets or because the outsourced solution was cost prohibitive.
I think it is true that lots of teams continue to outsource storage of metrics data but the outsourced vendors are not incentivized to make it easy to do retention/filtering/aggregation well.
Source: Have run observability stacks across a variety of domains.
In this scenario I find it odd that there are so many opensource projects with traction and success (one name above all, Prometheus) if they are addressed only to specialized companies. What you say applies to small companies or very big ones with lot of money to spend. Mid-sized companies in my experience prefer to spend money on in-house solution because at their volumes outsourcing is really costly. But that's just my small experience.