A Prometheus fork for cloud scale anomaly detection across metrics and logs
zebrium.com
zebrium.com
"Get up to speed with Prometheus, the metrics-based monitoring system used by tens of thousands of organizations in production. This practical guide provides application developers, sysadmins, and DevOps practitioners with a hands-on introduction to the most important aspects of Prometheus, including dashboarding and alerting, direct code instrumentation, and metric collection from third-party systems with exporters."
Also, accurate reporting and backfill are not related.
Curious to know why you don't think accurate reporting and backfilling aren't related. In my experience, they absolutely are -- mostly during disaster scenarios, where that information is more critical than any other time.
I strongly disagree. If you are unable to backfill data then report accuracy will be affected during multiple scenarios, including migrations from other products, unplanned downtime etc. Not having backfill is accepting that reports will suffer as soon as historical data is inaccurate and basically saying that you don't care.
There's also a ticket open about it[1].
[0] https://prometheus.io/docs/introduction/roadmap/#backfill-ti...
https://github.com/VictoriaMetrics/VictoriaMetrics/wiki/FAQ
(Not affiliated, just impressed with the author's blog posts explaining how it's better than competitors.)
We've had a great experience with prometheus - both the community and the project. Building what we did would have been immeasurably harder without them. That said, this has been a rapid process for us. We are starting to engage and see what parts of our fork they would be interested for upstream. (In particular our storage needs are a bit different, so we'll see how that fits.)
And yet again, it would be far better to just export the data from Prometheus into a distributed columnstore data warehouse like Clickhouse, MemSQL, Vertica (or other alternatives). This gives you fast SQL analysis across massive datasets, real-time updates regardless of ordering, and unlimited metadata, cardinality and overall flexibility.
Prometheus is good at scraping and metrics collection, but it's a terrible storage system.
The original design inspiration, borgmon, also was a terrible storage system and had an external long-term store layered on top of it.
This isn't a design flaw, it's an intentional trade off to make the core use case as bulletproof as it can be. Having seen "monitoring systems" based on something like Cassandra, aka distributed storage, is cringe inducing. The first thing to crash a the first sign of network trouble is distributed storage.
Wouldn't adding that let you switch back to the main project and lower the local storage buffer to as small as possible?
I would be interested in a demo of the logs AI stuff using real data. Something like, https://www.honeycomb.io/play/ would do.
Do you have such?
The ability to export from prometheus to a full-featured SQL database is definitely one aspect. In addition, for scaling, efficient transport from the scraper to the data store becomes pretty important.
Beyond just compression, having a transport protocol that takes advantage of the fact that much of the data does not change across successive scrapes, and the data that does change is often incremental (e.g. counters) makes a big difference.
hence we did not use remote storage adapter. We introduced a new interface that plugs into the scraper directly.
Having scanned this article, read the wikipedia entry and hit the landing page, I am none the wiser.
https://prometheus.io/docs/prometheus/latest/querying/functi...
Nothing stopping folks from adding in more.