Infrastructure Monitoring with Prometheus at Zerodha
zerodha.tech
zerodha.tech
https://medium.com/@valyala/high-cardinality-tsdb-benchmarks...
Good to see it getting more attention.
They have basically self hosted almost everything with a team of 30
I would be curious is this syncing being done by something custom they wrote or is this something built into Alertmanager now? I had look at the latest Alertmanager docs and nothing jumped out at me regarding this.
Alertmanager doesn't have any solution for this, although if you decide to use the Alertmanager using Prometheus Operator, then you get automatic config reloads. We decided to keep our AM cluster independent and out of K8s, as we pointed out we have a hybrid environment.
Hope that answers your question.
One other thing I wanted to ask was how many nodes are in your Victoria cluster?
Victoria Metrics right now runs as a single node setup. That works out for now, because as mentioned Prometheus maintains a WAL. So even if the Remote storage is down for sometime, the core functionality of alerts isn't affected. And we've setup alerts on VM health, so if it goes down, we've to act on it and get it back up. Once that happens, all the previous data also is ingested automatically.
This is something that we would surely revisit on at a later point of time, and setup a proper HA on it if needed. :)
[1]: https://github.com/prometheus/alertmanager/blob/master/READM...
Although yes you're right, it was a bit difficult to get it running out of the box with native Federation capabilities. A secondary long term storage like Victoria/Cortex/Thanos is something you can check out.