Influxdb made the switch from Go to Rust
old.reddit.com
old.reddit.com
Not sure what's the best solution though. Having a "stable" but fundamentally limited product (I guess influxdb v1) or breaking stuff in hopes of ending up with a way better technical foundation.
Hopefully it's improved, but last time I tried upgrading I found the UX in grafana to be subpar on the newer versions, as I recall you lost the autocomplete/UI to build your queries. Obviously grafana is it's own project but feels like they (influx) should invest more resource in areas like this to encourage people to upgrade - if you're going to do major upgrades make sure they have feature parity
I looked at TimescaleDB but at the time there was no easy way to get data from Telegraf to TimescaleDB. Telegraf finally merged code that allows writes to Postgres databases, but it took like 3 years to do that.
Ultimately, I still stuck with InfluxDB v1 because sending data to it via the InfluxDB line protocol is so simple. I have a couple of bash scripts that use awk to transform command output to Influx line protocol and send it to InfluxDB. It's just so simple. I love it.
I love learning about new things, but the InfluxDB v1 keeps working fine so I may not switch from it until something forces me to do it.
I ended up trying VictoriaMetrics by near accident as infuxdb didn't like something on my raspberry pi, and honestly it has been pretty painless. It is Prometheus-like stack which means you can use any PromQL-compatible things with it. There is "all in one binary", and version split by functions.
VM have tools to migrate from InfluxDB v1. I ended up just sticking old influxdb data in one database, as I wanted to change the format of what I write to it along with the migration.
> Ultimately, I still stuck with InfluxDB v1 because sending data to it via the InfluxDB line protocol is so simple. I have a couple of bash scripts that use awk to transform command output to Influx line protocol and send it to InfluxDB. It's just so simple. I love it.
It also have agent that's job is to convert from various protocols, and do the scraping, that includes influxdb, and few other popular protocols.
Thanks for sharing your experience.
Next time around I'm going to give TimescaleDB a look.
Sorry, but at that point, we've decided to rebuild the entire metric visualization once on TimescaleDB, since we're running postgres a lot anyhow.
If not, I would suggest looking at a proper OLAP DB. VictoriaMetrics has been great and was easy to set up.
Are managed "proper OLAP DB" solutions competitive with managed RDBMS from a price and ease of use standpoint?
This link has a comparison of features[1].
[1] https://docs.timescale.com/about/latest/timescaledb-editions...
And beyond that, TimescaleDB works with a few things we have already. We could migrate Zabbix to use TimescaleDB for a large performance boost. Also 1-2 teams are building reporting solutions for the product platform, and they are generating some significant timeseries data in a Postgres database as well.
My bad experience with TimescaleDB 3 years ago was that enabling compression required disabling the "dynamic labels" feature, which was a total nonstarter for us. A proper timeseries DB is designed to achieve great compression while also allowing flexibility of series. Hopefully Timescale will/has fixed that without adding another drastic perf tradeoff, but given how Postgres is architected for OLTP I would be surprised.
Just solve compression on the block level, why are you so specific about it happening in the database? It’s probably one of the least interesting feature comparisons when betting on which database to trust.
What is the "dynamic labels" feature? Is it a part of Postgres or Timescale?
I assume it's doing an automatic ALTER TABLE when necessary, which modifies each row and somehow breaks compression across the sharded tables. Or at least an automatic re-compression would cause massive latency on insert that they wanted to avoid.
The most common approach here is just to store the step of "dynamic" labels in JSON, which can be evolved arbitrarily.
And we've found that this type of data actually compresses quite well in practice.
Also regarding compression, Timescale supports transparent mutability on compressed data, so you can directly INSERT/UPDATE/UPSERT/DELETE into compressed data. Under the covers, it's doing smart optimizations to manage how it asynchronously maps individual mutations into segment level operations to decompress/recompress.
(Timescale cofounder)
Solutions like Grafana Mimir, Victoria Metrics, Clickhouse, or yes, the new Influx implementation, are much more scalable and will give you much fewer headaches.
ClickhouseDB is realy brilliant, btw, it's a powerhouse. Especially with the fairly recent additions that enable hybrid local + S3 option, pushing older metrics to S3 for cheap long-term storage.
It’s fantastic for workloads that neatly fit in the hypertable pattern though.
https://news.ycombinator.com/newsguidelines.html
Are you expecting a real answer of how I was hoping timescale's internal watermark system would help me roll up a total count or are you just implying I'm an idiot?
My issue was that for a grand total I didn’t have a time column, so I couldn’t define my query as a continuous aggregate and the query had to start counting from the start of my underlying series each time.
mike (at) timescale or DM on twitter?
But there are lots of approaches, depending on your needs.
You can (should) define a "cache disk" for S3, which will cache up to X Gb locally to avoid trashing.
Another option is is to move data into separate (purely S3 backed) tables after a certain time to avoid accidentally fetching large amounts of data from S3. You can still easily join the data together if needed.
https://www.timescale.com/blog/expanding-the-boundaries-of-p...
Victoriametrics so far works very well.
As well, I am one of those folks that happens to find the Flux query language powerful, but it's not easy enough for folks to just make that jump from SQL. Flux is much closer to Splunk's search language. It is good at what it does. FluxQL doesn't even have date parsing (which is really odd for a time series query language), but FlightSQL in 3.x seems to be more complete.
I like that they are converging towards SQL, but at the same time it's a bit like going back to square one. They seem more convinced about going full SQL this time though, but yeah
Just searching for this, I stumbled on this documentation page that illustrates the point very well:
https://docs.influxdata.com/influxdb/v1/query_language/
In the same page (about the original influxql in v1), there is a depecration notice for v1 stating that v2 is the stable version, implying that InfluxQL is not recommended. And a pop up notice stating that v2 (flux) is basically deprecated and just in maintenance mode, and that you should use InfluxQL. But as I said in my earlier comment, I guess in some ways that's better than being too rigid and sticking with bad or less ideal technical decisions.
We really wanted to bring Flux along too, but found that it was too difficult in the near term to have it work well with v3. We spent a bunch of time building a gRPC API that Flux uses to talk to v3 (the same thing we have in our Cloud v2 product), but that API was designed with the previous storage engine in mind. It ended up being brittle and performed very poorly.
So at this point the long term supported languages for InfluxQL and SQL, but we're continuing to support Flux for our customers.
But now with these VC-funded tech products that have spawned over the last 5-7 years, who have a move-fast-and-break-things attitude, I’m seeing the benefits of the old approach.
I suppose it’s all a matter of trade offs, as with all things, and there’s no silver bullet.
We'll have data migration tools for v1 and v2 into v3 later this year/early next.
But there are a few things that aren't there. Continuous queries, SELECT INTO, and anything that modifies data isn't there.
They are always… in flux * sun glasses on*
If we had dedicated personell to manage our monitoring we might have stuck with it.
https://news.ycombinator.com/item?id=25049253
At some point HN is going to have to decide if it's the Rust subreddit or a news site.
After that nobody wants to hear about Java.
Funny enough, in contrast to when I joined, the pendulum seems to have swung, and comments disparaging rust seem to be en vogue.
Opinions could be different if first they implemented a complete compatibility layer, Flux included, prior to making the migration.
I ask because ClickHouse is quite hot at the moment from my experience in consulting and that seems to be reflected in Google Trends [1].
And there are some startups relying on ClickHouse for their log/monitoring products like https://signoz.io and https://hyperdx.io.
[1] https://trends.google.com/trends/explore?date=all&q=ClickHou...
The non-reddit link target
They moved their entire stack from Go to Rust, rewrote the system from the ground, and spent a lot of time on it, I guess this is a big cost.
Is it worth it?
[0] https://www.influxdata.com/blog/the-plan-for-influxdb-3-0-op...
For one database that receives 100M 600 byte JSON records/day, a single node AWS PostgreSQL RDS instance is handling it effortlessly, and the DBA work is very part-time. We keep year+ of summaries and 48 hours of detail, unloading the rest to S3 as parquet files, queryable by Athena if we need. AFAICT, we're spending <$4K/mon all in, including backups.
p.s. a buddy at a top-3 TV streaming service is also doing this for logging all viewing activity, but with Aurora.
how are you dealing with failover?
how are you dealing with interactive reporting?
how large are your records?
is the s3 data in cold storage? how much is stored in S3 right now? is it warm-enough to query via (e.g.) athena?
Someone needs to pickup the original ideas of 1.x since they can't seem to stay focused, as their marketshare is ripe for grabbing.
With v3, I prefer to think of it as us doubling down on core database performance and functionality. With v2 we tried to create this whole development platform. V3 brings our focus back to the core database, which I think will yield better results for everyone.
If we speak about metrics, Prometheus just win.
Let's put this way: is there any killer feature that can ditch most of the Prometheus installations in their favor?
If they on top of that provided some compatibility layer for Prometheus/PromQL, like few other competitors did, then the prospective enterprise client have warm fuzzies that if they don't like it they don't need to rewrite entirety of their stack to work with something else. People could also use the existing ecosystem and "just plug it in", even replacing Prometheus instances they might have.
In the end people want to ingest the metrics and display it in Grafana. They don't need another visualisation solution that has less support and documentation. They don't want to learn new weird query language that is simulatenously more verbose and less readable than PromQL or even influxQL from v1.
Plus, in the last years with the rise of Kubernetes, most corporates have at least one cluster with Prometheus monitoring it.
Always, base on my experience, Prometheus is very popular, but the adoption is not so wide due to its `oss nature`: CTOs want someone to blame when things go wrong.
I like Prometheus and think this is a great piece of software. But even if we won't go into actual details, Prometheus is baked in into Kubernetes monitoring [0]. That's the first monitoring system young engineers will meet with when learning k8s. Although, k8s and Prometheus are both CNCF projects which means both of them will be promoted in synergy with each other.
> It is more likely that the younger engineer is going to learn that companies don't care about what is popular on HN.
This is not what I think younger engineers do :)
[0] https://kubernetes.io/docs/tasks/debug/debug-cluster/resourc...
I imagine similar dashboard services that don't necessarily work well in Prometheus are a good market for these types of databases. Prometheus is nice, but I don't think it's suitable for all Influx use cases (and vice versa).
Prometheus-compatible interface for query, a bunch of ingest protocols, smaller memory usage than InfluxDB (v1, haven't tested v2 coz new language have less grafana support). Options to scale too
Basically either you manage it yourself, or you pay them to do Serverless/Dedicated/Clustered hosted setup for you.
https://www.theregister.com/2023/07/11/influxdata_apologizes...
Telegraf and InfluxDB are solid, but Chronograf was behind Grafana in usability and features.
Kapacitor was pretty rough. The language it used was hard to write, the docs were barebones, somewhat confusing, and sometimes inaccurate. When I switched off of Kapacitor, the CPU usage on the server dropped significantly too. So I'm guessing Kapacitor wasn't too CPU friendly either.
Prometheus is not just the db itself, it’s the ecosystem around it. You’ve got service-discovery, alertmanager and basically every application in existence having a /metrics endpoint and some pre-made Grafana dashboard.
interface_if_octets host=router,instance=eth0,rx=123584,tx=213956
while in prometheus it would be interface_if_octets host=router,instance=eth0,type=rx 123584
interface_if_octets host=router,instance=eth0,type=tx 213956
which in theory yes it is more compact but it gave me more annoyances than advantages during querying>While Prometheus is often good enough for standard metrics, it is just things it can't handle.
My experience is that just anything made to ingest and analyze logs ends up mediocre for metrics and vice versa. I don't think I've seen single product that did both well or efficiently. So I'd rather have good metrics and just use ELK/Graylog/whatever else for logs.
With influx you can save every event. So you know exactly when it happened, the unique labels for that event etc. It's a completely different paradigm.
So in Prometheus you have
myCounter,pod=1,time=20:23,value=1000
myCounter,pod=2,time=20:23,value=500
myCounter,pod=1,time=20:24,value=1100
myCounter,pod=2,time=20:24,value=700
so all you know is that some event happened 100 times on pod1 and 200 times on pod2 the last minute. But with influx you could have a row for every single event. Of course that explodes the query time in comparison, but allows you to do much more with the data if needed.That's not true. You're referring to pull-based approach for metrics collection. It has its tradeoffs (like fixed interval scraping), but has a lot of benefits too (like higher reliability). Check the following link [0] from VictoriaMetrics docs, which supports both push and pull approaches. Prometheus also gained push support this year, though.
However, the main difference between Prometheus-like systems (Thanos, Mimir, VictoriaMetrics) and more traditional DBs for time series like InfluxDB or TimescaleDB is that first are designed to reflect system's state, and last are designed to reflect system's events. That's the main difference in paradigm, data model, and query languages. There is a reason why PromQL is so easy in 99% of cases, and so complex and annoying when users want to express what they get used to in traditional databases.
I'm saying this because I went through creating a Grafana datasource for ClickHouse [1] and I felt how complicated it is to express a most straightforward PromQL query in SQL, and vice versa.
If you'd like to learn more about differences between common queries for plotting time series in PromQL and SQL see my talk here [2].
[0] https://docs.victoriametrics.com/keyConcepts.html#write-data
[1] https://grafana.com/grafana/plugins/vertamedia-clickhouse-da...
Metrics is certainly one use case that people pay us for. With v3, we expect that real-time analytics and some more data warehousing types of use cases will become interesting. We always envisioned InfluxDB as a store for observational data of all kinds, not just metrics.
On how we make money, we sell our products. We have at this time:
- InfluxDB v1 Enterprise (a self-managed, clustered implementation of InfluxDB)
- InfluxDB v1 Cloud (Enterprise, but as as single-tenant managed service. We still run this for hundreds of customers)
- InfluxDB v2 Cloud (multi-tenant, usage based, we're running this for thousands of customers)
- InfluxDB v3 Serverless (multi-tenant, usage based v3)
- InfluxDB v3 Cloud Dedicated (single-tenant, resource based pricing)
- InfluxDB v3 Clustered (self-managed v3, clustered database)
We'll have single server versions in the future, but we're a bit off from that. Right now our focus is on continuing support for our v1 and v2 customers, and further developing our v3 products for new customers and customers that want to migrate over.
Maybe you have some context or background info I don't have?
Unfortunately I was stupid enough to build my dashboard on flux, which I’m really sorry to say I dislike quite a bit, while still wanting to be respectful for the people who build the stuff.
All that said, I think that influx is a great tool although I’m mostly using it for personal projects and haven’t run anything at scale.
Like, it looks super powerful for complex queries I will never need to make...
With C/C++ (and CMake + Ninja) it seemed we were finally getting to a point where incremental builds would complete before hitting the 400ms attention span "Doherty Threshold", and now it seems we are going back to the days of having to spend our time sword-fighting (xkcd) during slow builds.
People got the memo that slapping Serialize on everything has a cost an unless you're doing huge native dependencies which require long compilation time, it's pretty snappy.
A company can make good money for quite a while as long as the performance is "good enough", and usually focuses on adding features instead of worrying about any rewriting/re-architect until there is a bottleneck or issues start to significantly slow down development. How "successful" such a rewrite still needs to be seen, but this is risky for most companies/products.
>Then there's the question of why we did a rewrite at all. We wanted to get at some important requirements: >Unlimited cardinality >Analytics queries against time series at the performance of a columnar DB >Use object store as the durability layer for historical data (i.e. separate compute from storage) >SQL and broader ecosystem compatibility >All of that stuff taken together meant that we'd be rewriting most of the core of the database. ...
If you are a fan of Rust and want to rewrite all your just because of that and you are the CTO, whatever, what can I say. But the title here is a bit clickbait-y and I really don't see much reflection on the languages themselves and how to balance the business needs for such a transition.
I totally agree that a rewrite is risky. It's not something I'd choose to do again, but at the time we didn't really see any way around rewriting the bulk of the database (even if we kept it implemented in Go).
Using Rust and the Arrow ecosystem of projects (Parquet, DataFusion, Flight) meant that there were a ton of things we didn't have to do from scratch. One of our staff engineers, Andrew Lamb, has called it a toolkit for building databases. Thanks in part to his contributions, I think he's right.
When people talk about separating compute from storage, they mean pulling compute heavy tasks like query, ingest, indexing, and compaction apart and using a shared storage tier that many systems can talk to. Usually this is object storage paired with some sort of catalog (kept in either object storage or some other store or API).
Snowflake popularized this approach in the data warehousing and OLAP space with great success. Their papers submitted to VLDB are great reads on the topoic.
wow, the from scratch rewrite. I can't even imagine that for a major piece of software
https://www.influxdata.com/blog/rust-can-be-difficult-to-lea...
The rewrite started in 2020... they rationalize now, but it's pretty clear they just really wanted Rust and found the reasons they list later as a post-decision justification. Nothing wrong with that, if you don't mind risking the future of your business on a risky rewrite... though if I had a job working on the Go code base and my employer suddenly announced we should drop everything and start a rewrite, so go learn Rust, which has a much smaller job pool in my area as far as I can see, I guess I would've been really pissed off and would leave as fast as possible.
It wasn't until late last year that we made the decision to go all in on the rewrite and made that the focus of everyone in engineering. And we did that because we had 4 years of experience trying to get v2 and Flux to be successful, with modest results.
Most of the time we were developing this version, we were spending massively more engineering effort on developing v2 or maintaining v1 for our customers.
Using rust for any other reason than its the best option technically, is just fad chasing.
How does learning Rust reduce your chances of getting a new job? Learning new stuff increases your chances, not decreasea them.
- No garbage collector
- Fearless concurrency (thanks Rust compiler)
- Performance
- Error handling
- Crates
- they thought they were gonna use C++ and wanted interop (ended up not using C++?)
- ecosystem: Apache Arrow DataFusion
- "I thought that if we're going to rewrite most of the database anyway, we might as well do it in the best language choice in 2020"
But the real reason might be: "Rust good, Go bad" /s
To be honest, most of the "we started in Go and switched to Rust" stories read to me as "you should have always known that you should have started in Rust" (or C++ or something, though I'd choose Rust out of the viable set of languages here too). It's IMHO always been obvious that Go was not really viable, I mean, sure, you can get farther than you could with starting with Python, but, it's not something Go was even trying to solve.
Not enough, anyway, to make a complete rewrite more profitable over adding features.
After all, it's not as easy if influx was plagued by concurrency bugs, was it?
A move to a safer language than go just doesn't seem worth it, and there's little to no performance gain when you can just throw more hardware for that tiny performance difference.
Admittedly I was at a shop that wasn't likely to shift to an enterprise version so not a direct economic loss on their part but definitely stopped thinking of it as a solution to consider.
</cope>