CERN swaps out databases to feed its petabyte-a-day habit
theregister.com
theregister.com
I was using Promscale (TimescaleDB) but they EOL'd Promscale which forced us to Victoria. But either way both of these are much faster than Influx
Don't get fooled into the latest InfluxDB rewrite. I think the latest is cloud hosted only too? So stupid
I also prefer the golang-esque simplicity of the Prometheus ecosystem. Monitoring is the last place I want unnecessary abstraction layers and complicated configuration files.
I also use Telegraf with TimescaleDB.
Telegraf actually made interop with their competitors super easy
I can say from what I've seen that the amount of data they have to deal with is still the #1 problem. Before the multi layered filters they generate a petabyte of data per second.
vmagent takes care of all the pesky edge things like emulating prometheus config parsing and various scraping bits. It also does buffering in case you lose network connection for a while, and accept vast spread of different protocols
vminsert/vmselect scale separately from eachother and your queries don't bother your ingest all that much.
vmstorage does just that, storage. Only thing that bothers me (compared to say, Elasticsearch), is that data can't migrate between nodes so you can't "just" start a new one and drain an old one, but a tiny bit ops work in rare cases is IMO price worth paying for straightforwardness of the stack..
PromQL compatibility is also great, tools like Grafana "just work" without anyone having to write support for it.
We started migrating from InfluxDB at work, and on my private stuff I already did. Soo much less memory usage too.
Also frankly Prometheus support is a massive positive. For better or worse industry standarized on apps using Prometheus as ingest for metrics, and also most of the materials related to that will of course give examples in PromQL
Flux is frankly hieroglyphs for people using it 20 minutes a month like our developers
This is given example on how to raise value in Flux to power of two
|> map(fn: (r) => ({ r with _value: r._value * r._value }))
This is example of that in prometheus value ^ 2
This is example of calculating percentage in Flux (from their webpage) data
|> pivot(rowKey:["_time"], columnKey: ["_field"], valueColumn: "_value")
|> map(
fn: (r) => ({
_time: r._time,
_field: "used_percent",
_value: float(v: r.used) / float(v: r.total) * 100.0,
}),
)
This is how you do it in PromQL space_used / space_total * 100
Flux is atrocious for "normal users".I probably wouldn't default to it, but I've had to solve some issues that were complicated in PromQL that would probably be easier here. E.g. I had to work on a Prometheus monitoring/alerting setup where metrics were all recorded in UTC, but alerts should only fire during local business hours. We ended up with something like a thousand character PromQL query that was utterly illegible (have pity on whatever poor soul has to update it when DST goes away).
Flux looks like it would have a more legible solution to that, though at the cost of making simple analysis way more complicated than it probably needs to be.
I've personally found InfluxDB more user-friendly than Prometheus for metrics work, since I've used both at scale and compared my own experiences. The industry leans towards Prometheus nowadays, so I'm used to dealing with Prometheus, but I've found PromQL particularly unpleasant for any kind of complex aggregations. You may disagree, and that's okay. Regular SQL seems far nicer to me in comparison, so it's nice to see Influx v3 focusing on that. I wish Prometheus and VM would develop a standard SQL interface.
VictoriaMetrics doesn't fully support PromQL either; its MetricsQL is only about 73% compatible according to this article[0] that the VM docs link to. I certainly hope that VM took this incompatibility and used it as an opportunity to make MetricsQL more enjoyable to use than PromQL.
[0]: https://medium.com/@romanhavronenko/victoriametrics-promql-c...
[0] https://docs.victoriametrics.com/vmalert.html [1] https://docs.victoriametrics.com/vmalert.html#rules-backfill... [2] https://docs.victoriametrics.com/vmalert.html#never-firing-a...
That is, VictoriaMetrics has not really built a true time series DB that handles reasonable cardinalities.
What do you consider reasonable cardinalities or a true TSDB?
If I understand this correctly, it deals with high cardinality by dropping data. The operators need to monitor for this and adjust their data to lower the cardinality.
> By default VictoriaMetrics doesn't limit the number of stored time series.
They have put out some benchmarks showing VictoriaMetrics ingesting 40M time series: https://valyala.medium.com/high-cardinality-tsdb-benchmarks-...
It's possible to smooth this loss out (by assuming a normal distribution of lost data) if it's noticed and limited. Though I do not think there's any commercial TSDB that does that automatically.
Speaking to The Register, Roman Khavronenko, co-founder of VictoriaMetrics, said the previous system had experienced problems with high cardinality, which refers to the level of repeated values – and high churn data – where applications can be redeployed multiple times over new instances.
Implementing VictoriaMetrics as backend storage for Prometheus, the CMS monitoring team progressed to using the solution as front-end storage to replace InfluxDB and Prometheus, helping remove cardinality issues, the company said in a statement.
"InfluxDB said in March this year it had solved the cardinality issue with a new IOx storage engine."
Does this mean that in the end it wasn't really necessary to switch to VictoriaMetrics' offering?
InfluxDB said in March this year it had solved the cardinality issue with a new IOx storage engine for it's hosting customers only
There isn't an influxdb that you can download with cardinality solved.
https://www.influxdata.com/blog/the-plan-for-influxdb-3-0-op...
> But Brij Kishor Jashal, a scientist in the CMS collaboration, told The Register that his team were currently aggregating 30 terabytes over a 30-day period to monitor their computing infrastructure performance.
So 1 TB / day, that's about 10 MB/s.
I looked at some of the alternatives to victoriametrics for Prometheus and they all seem… much much worse…
https://github.com/VictoriaMetrics/VictoriaMetrics/blob/mast...
/tumbleweed...