Not sure what's the best solution though. Having a "stable" but fundamentally limited product (I guess influxdb v1) or breaking stuff in hopes of ending up with a way better technical foundation.
Not sure what's the best solution though. Having a "stable" but fundamentally limited product (I guess influxdb v1) or breaking stuff in hopes of ending up with a way better technical foundation.
Sorry, but at that point, we've decided to rebuild the entire metric visualization once on TimescaleDB, since we're running postgres a lot anyhow.
If not, I would suggest looking at a proper OLAP DB. VictoriaMetrics has been great and was easy to set up.
Are managed "proper OLAP DB" solutions competitive with managed RDBMS from a price and ease of use standpoint?
This link has a comparison of features[1].
[1] https://docs.timescale.com/about/latest/timescaledb-editions...
And beyond that, TimescaleDB works with a few things we have already. We could migrate Zabbix to use TimescaleDB for a large performance boost. Also 1-2 teams are building reporting solutions for the product platform, and they are generating some significant timeseries data in a Postgres database as well.
My bad experience with TimescaleDB 3 years ago was that enabling compression required disabling the "dynamic labels" feature, which was a total nonstarter for us. A proper timeseries DB is designed to achieve great compression while also allowing flexibility of series. Hopefully Timescale will/has fixed that without adding another drastic perf tradeoff, but given how Postgres is architected for OLTP I would be surprised.
Just solve compression on the block level, why are you so specific about it happening in the database? It’s probably one of the least interesting feature comparisons when betting on which database to trust.
What is the "dynamic labels" feature? Is it a part of Postgres or Timescale?
I assume it's doing an automatic ALTER TABLE when necessary, which modifies each row and somehow breaks compression across the sharded tables. Or at least an automatic re-compression would cause massive latency on insert that they wanted to avoid.
The most common approach here is just to store the step of "dynamic" labels in JSON, which can be evolved arbitrarily.
And we've found that this type of data actually compresses quite well in practice.
Also regarding compression, Timescale supports transparent mutability on compressed data, so you can directly INSERT/UPDATE/UPSERT/DELETE into compressed data. Under the covers, it's doing smart optimizations to manage how it asynchronously maps individual mutations into segment level operations to decompress/recompress.
(Timescale cofounder)
Solutions like Grafana Mimir, Victoria Metrics, Clickhouse, or yes, the new Influx implementation, are much more scalable and will give you much fewer headaches.
ClickhouseDB is realy brilliant, btw, it's a powerhouse. Especially with the fairly recent additions that enable hybrid local + S3 option, pushing older metrics to S3 for cheap long-term storage.
It’s fantastic for workloads that neatly fit in the hypertable pattern though.
https://news.ycombinator.com/newsguidelines.html
Are you expecting a real answer of how I was hoping timescale's internal watermark system would help me roll up a total count or are you just implying I'm an idiot?
My issue was that for a grand total I didn’t have a time column, so I couldn’t define my query as a continuous aggregate and the query had to start counting from the start of my underlying series each time.
mike (at) timescale or DM on twitter?
But there are lots of approaches, depending on your needs.
You can (should) define a "cache disk" for S3, which will cache up to X Gb locally to avoid trashing.
Another option is is to move data into separate (purely S3 backed) tables after a certain time to avoid accidentally fetching large amounts of data from S3. You can still easily join the data together if needed.
https://www.timescale.com/blog/expanding-the-boundaries-of-p...
But now with these VC-funded tech products that have spawned over the last 5-7 years, who have a move-fast-and-break-things attitude, I’m seeing the benefits of the old approach.
I suppose it’s all a matter of trade offs, as with all things, and there’s no silver bullet.
Victoriametrics so far works very well.
As well, I am one of those folks that happens to find the Flux query language powerful, but it's not easy enough for folks to just make that jump from SQL. Flux is much closer to Splunk's search language. It is good at what it does. FluxQL doesn't even have date parsing (which is really odd for a time series query language), but FlightSQL in 3.x seems to be more complete.
I like that they are converging towards SQL, but at the same time it's a bit like going back to square one. They seem more convinced about going full SQL this time though, but yeah
Just searching for this, I stumbled on this documentation page that illustrates the point very well:
https://docs.influxdata.com/influxdb/v1/query_language/
In the same page (about the original influxql in v1), there is a depecration notice for v1 stating that v2 is the stable version, implying that InfluxQL is not recommended. And a pop up notice stating that v2 (flux) is basically deprecated and just in maintenance mode, and that you should use InfluxQL. But as I said in my earlier comment, I guess in some ways that's better than being too rigid and sticking with bad or less ideal technical decisions.
We really wanted to bring Flux along too, but found that it was too difficult in the near term to have it work well with v3. We spent a bunch of time building a gRPC API that Flux uses to talk to v3 (the same thing we have in our Cloud v2 product), but that API was designed with the previous storage engine in mind. It ended up being brittle and performed very poorly.
So at this point the long term supported languages for InfluxQL and SQL, but we're continuing to support Flux for our customers.
Hopefully it's improved, but last time I tried upgrading I found the UX in grafana to be subpar on the newer versions, as I recall you lost the autocomplete/UI to build your queries. Obviously grafana is it's own project but feels like they (influx) should invest more resource in areas like this to encourage people to upgrade - if you're going to do major upgrades make sure they have feature parity
I looked at TimescaleDB but at the time there was no easy way to get data from Telegraf to TimescaleDB. Telegraf finally merged code that allows writes to Postgres databases, but it took like 3 years to do that.
Ultimately, I still stuck with InfluxDB v1 because sending data to it via the InfluxDB line protocol is so simple. I have a couple of bash scripts that use awk to transform command output to Influx line protocol and send it to InfluxDB. It's just so simple. I love it.
I love learning about new things, but the InfluxDB v1 keeps working fine so I may not switch from it until something forces me to do it.
I ended up trying VictoriaMetrics by near accident as infuxdb didn't like something on my raspberry pi, and honestly it has been pretty painless. It is Prometheus-like stack which means you can use any PromQL-compatible things with it. There is "all in one binary", and version split by functions.
VM have tools to migrate from InfluxDB v1. I ended up just sticking old influxdb data in one database, as I wanted to change the format of what I write to it along with the migration.
> Ultimately, I still stuck with InfluxDB v1 because sending data to it via the InfluxDB line protocol is so simple. I have a couple of bash scripts that use awk to transform command output to Influx line protocol and send it to InfluxDB. It's just so simple. I love it.
It also have agent that's job is to convert from various protocols, and do the scraping, that includes influxdb, and few other popular protocols.
Thanks for sharing your experience.
Next time around I'm going to give TimescaleDB a look.
They are always… in flux * sun glasses on*
If we had dedicated personell to manage our monitoring we might have stuck with it.
We'll have data migration tools for v1 and v2 into v3 later this year/early next.
But there are a few things that aren't there. Continuous queries, SELECT INTO, and anything that modifies data isn't there.