Also sometimes you'll have some sort of a batch job, or a cron that can't be scrapped the regular way, checkout Prometheus Pushgateway https://prometheus.io/docs/practices/pushing/
Also sometimes you'll have some sort of a batch job, or a cron that can't be scrapped the regular way, checkout Prometheus Pushgateway https://prometheus.io/docs/practices/pushing/
> If you get serious about Prometheus, eventually you will want longer data retention, checkout https://thanos.io/
Any idea how it compares with https://victoriametrics.com/ ?
We're slowly looking for a replacement for InfluxDB (as 1.8 is essentially on life support), the low disk footprint is pretty big advantage here.
They have a blog post explaining their reasoning:
https://prometheus.io/blog/2016/07/23/pull-does-not-scale-or...
Prometheus remote_write is what people graduate to, this gets you the rest of the way, and you are correct it's less PITA at scale.
If you're looking for retention your choices are large, there's Cortex (CNCF), Mimir (most Cortex work moved here), Thanos, VictoriaMetrics, TimeScale, Chronosphere, and many others.
All seek to do a similar thing from a distance, they all store metrics (likely from Prometheus) and allow retention and some variety of how to query it (if you want SQL you got it, if you want non-standard functions you go it, if your reads are more important than your writes you got it, if you need a billion active series you got it, etc).
If what you want is "Prometheus but bigger" then the Prometheus project provides a compliance suite that you can run to help you evaluate your options: https://github.com/prometheus/compliance
I work for Grafana Labs, and we have maintainers working for us who have touched Prometheus, Thanos, Cortex and Mimir. Mimir is currently the largest investment we have https://github.com/grafana/mimir and it is 100% compliant with Prometheus (though that is about to be temporarily untrue as Native Histograms is landing in Prometheus soon https://github.com/prometheus/prometheus/milestone/10 and we'll need to add a perfectly compliant support to Mimir to get back to being compliant).
Pull does work like 90% of the time, + you get a handy uptime check built-in as well, as if the pull didn't work, the service is probably offline and you'll notice it immediately. With Push you still need a way of checking if the service is online/offline.
"Sure, Bob, just write a promql query against this interface"
https://news.ycombinator.com/item?id=1636348
(in case my point is too subtle: I believe that the value of long term metric retention is to support business decisions more than supporting availability related monitoring.
The tools "business deciders" use to look at "the things they call metrics" are simply unable to look at data lodged in a prometheus silo. So the data you have for long term retention is simply unusable for the thing it's most valuable)
(edited for clarity)
Do I point the biz at it? Only when I'm trying to show them availability related metrics.
Further, the vast majority of workaday devs are stressed out and under tremendous pressure to thing all the things in their ticket queue. They're probably pretty fluent in their thinging language and likely also conversant in SQL. Asking them "hey learn this other query language, because" is just adding stress to them. Sure, some are going to rise above and promql things too, but for lots, it's telling them to "write a distributed map reduce query in erlang."
However, Prometheus is not a tool that specializes in reporting availability to people who live in Excel. Prometheus scrapes metric from a lot of systems in a standard format, and enables you to see what's happening and write alerts to fix things when they break, without requiring complex frameworks. Reporting availability to people is just a small, tiny thing that Prometheus is able to do.
It doesn't rely on reading the Prometheus WAL (rather, it uses Prometheus's built-in remote-write functionality). This is both less fragile and has the benefit of being about to run Prometheus in stateless mode (or replace it entirely with something like Grafana Agent, vmagent, etc.
Grafana Labs has done a proof of concept showing Mimir scaling to 1 billion active timeseries. There are also "real" users serving 700+ million metrics in production.