Granted, you don't need to touch this confirmation very often, but anyone who's going to operate the cluster will need to understand it thoroughly.
The fact that it's ad-hoc (prometheus.io/probe etc. aren't built in) means everyone's config is probably going to be unique and not portable. For example, we found the current config example to be insufficient, since blackbox-exporter needs information about whether its endpoint is HTTP or HTTPS.
Kubernetes' template system, combined with variable expansion, seems like it would be a better model for what you're currently trying to do with service discovery.
Also: I'm setting this up right now, but it seems there's no exporter for the Kubernetes API proper, just Kubelet?
The way I would describe Prometheus in relation to those other tools:
- Prometheus is open-source and self-hosted.
- Prometheus is about dimensional numeric time series metrics only (no log-based analysis, no per-request tracing, etc.).
- Prometheus has a strong focus on systems and service monitoring, not so much on business metrics.
- Prometheus is more of a Swiss army knife of monitoring rather than a ready-to-drop-in package that starts monitoring everything automatically.
- Prometheus is very much about whitebox monitoring and manually defining any metrics that could be useful for you (although we support blackbox exporting and bridging metrics from existing systems as well).
- We don't do machine-learning-style anomaly detection, but we do alerting based on manually defined rules.
- For a purely metrics-based solution, the insight we deliver is one of the best in the field (via the dimensional data model and the query language to go with it).
- Many open-source projects are starting to expose native Prometheus metrics (like k8s, etcd, ...), which gives Prometheus an advantage when being used together with those.
EDIT: Also try the "Getting Started" tutorial - that should only take a couple of minutes to try it out: https://prometheus.io/docs/introduction/getting_started/
Any thoughts on a companion (open source) system that focuses more on logs and such?
https://prometheus.io/blog/2016/03/23/interview-with-life360... and https://prometheus.io/blog/2016/05/01/interview-with-showmax... look at two of our users.
Read more here: https://prometheus.io/docs/introduction/overview/ and here: https://prometheus.io/community/
A lot of open-source projects (Kubernetes, Etcd, ...) are also exposing native Prometheus metrics now, making it easier to integrate with those.
EDIT: also check the PromCon schedule for other user companies giving talks: https://promcon.io/schedule/
And the sponsors are also users: https://promcon.io/#our-sponsors
So it's a question of whether you can squeeze your data into that model and whether you need per-event details or whether aggregated time series are ok.