Right now I use graphite but would love something that also handles replication/redundancy with a good query language & enough performance to also use it to fetch the underlying data used to render front-end graphs for users.
Right now I use graphite but would love something that also handles replication/redundancy with a good query language & enough performance to also use it to fetch the underlying data used to render front-end graphs for users.
The big questions to ask are: - what data is of interest - what are the requested ingest patterns (how is data getting into this system? How frequently? What rate of ingest is expected?) - what are the requested query patterns (who is doing queries? What do those queries look like specifically? Are people querying over unbounded time ranges or are people doing more focused queries? Do queries regularly involve aggregations or not?) - what is the requested SLA (i.e. how much partial down time is okay? How much full downtime is okay? What kind of query response time do you need to target?) - what resources are/will be available (money, man power, compute, storage, tech on hand)
It's possible that a time series data store might not be the correct system choice once these questions are answered. It's not unusual to see data split or copied into multiple systems to answer all requirements.
So I would say, the system is 5 minute ingest of about 1 million metrics. This is spread out over a half dozen locations, each which currently records in their own silo.
And aggregate metric is calculated with a 5 minute lag, which reads all the just-written data points and aggregates into sum-totals which are themselves stored and cached in one place. This is another million metrics basically stored separate from the rest.
But it doesn't really change the character of the system. In the end I'm trying to; Write batches of mostly numeric data Queries over time against those numbers Aggregate data over different blocks of time; 5 minute, hourly, daily, monthly Store it efficiently Ensure redundancy, integrity
Seems like a simple and common enough problem to have been reasonably "solved" for orders of magnitude higher scale than I'm operating at.
At reasonable scale, I've seen people get really far with pure graphite setups by utilizing tools like [carbon-c-relay](https://github.com/grobian/carbon-c-relay), [carbonate](https://github.com/graphite-project/carbonate), and high integrity filesystems like zfs underneath. It's a very hands on operation though. Things like growing the cluster are hard to do without downtime.
If constant growth and uptime is a concern, something like openTSDB might be a great choice. The complexity of setting up a Hadoop + HBase cluster is a pretty big upfront cost, but man is this thing the cockroach of time series data stores. Adding storage is just growing the HBase cluster. Querying across years of data is pretty simple and straightforward. For the complexity involved, openTSDB is worth it.
Prometheus is not recommended for uses where 100% accuracy is required, see https://prometheus.io/docs/introduction/overview/#when-does-...?
I'd also be wary of using Graphite for such a use case. For billing a more traditional database is probably best.