Good in-depth walkthrough of Prometheus internals.
The CloudFlare specific part seems to be a couple notable extra limits they've placed on their instances.
They also said those patches with extra limits aren't going to be accepted to Prometheus repo. That's sad, people with the same issues would need to look for alternatives.
Interesting they are using tsdb for storing the time series. I thought it was out of favor and fashion. Clearly it is serviceable.
What alternatives do folks use?
Influx Db, Victoria metrics, clickhouse, some postgres add ons?
InfluxDB do not provide cluster version in open source and I know user who has been reported that InfluxDB droped its data for the last year - that's why he to migrating to VictoriaMetrics.