InfluxDB now supports Prometheus remote reads and writes natively
influxdata.com
influxdata.com
Prometheus comes along every so often and scrapes metrics your program exposes via a very simple API. This has the advantage that your code doesn't need to know about some endpoint of some cluster, it just needs to buffer up and expose some info it knows about itself from the recent past.
You don't need to maintain a cluster of Prometheus either, you can just run more than one for redundancy (kind of like an active-active) - it is meant for relatively ephemeral information, it is efficient, and one big Prometheus node will probably do you fine.
Where to store compacted historical metrics (less coarse resolution, still interesting data that you might not want to throw away) and how has been an open question. Sinking it into influx could be a good answer, so this is welcome news.
1. UDP packets are lossy, we had the UDP buffers of our Influx server fill up, and it took us a long time to detect we were dropping packets.
2. Many people want to detect when they are not getting data from an endpoint. Polling is a great way to quickly detect a endpoint is down.
1. My understanding of UDP being lossy typically refers to it happening via transit but your example is an endpoint failure.
2. Since the whole point of metrics is to keep track of operations then the monitoring of the metrics themselves should be alerting to anomalies?
Additionally, I would argue some sort of discovery/registry mechanism is in order anyway.
For example, Prometheus has very solid integration with Kubernetes. In this universe you have one central control plane thingy (the scheduler) responsible for bringing resources up and down, it updates the collector (Prometheus) about all the devices (pods, containers, damn we have too many terms for things) accordingly, which in turn scrapes on a best effort basis.
Once everything is hooked up like this you're basically guaranteed that applications that provide information about themselves will get scraped eventually. If they are reachable you'll know what is going on inside them, and if they are not, you'll know that too.
The advantages of this sort of decoupling are subtle and difficult to get right, but you'll be thankful down the line for having done so correctly from the get go.
(Which is also why I'm super excited to use timescale [http://www.timescale.com] once grafana gets support for postgres!)
https://www.influxdata.com/blog/the-open-source-database-bus...
Why is this useful? Prometheus is wonderful for ephemeral application state and monitoring, but isn't really meant for storing metrics longer term. Sometimes you want to look at the same metrics over a year or so. This is what Influx is built for. So you have prometheus for collecting metrics and monitoring your cluster, then you have it reading and writing to Influx. This is literally the best of both worlds. This will be a big deal going forward for our team.
Does this make sense?
Influx is a time series database where the way the bits on disk are stored is written in "time series" order, so you can do certain types of operations literally orders of magnitude faster than you can on a more generic datastore (such as cassandra, mongo, mysql, postgresql, etc). The clustering bits in Influx are enterprise only, but influx (non-clustered) is entirely open source.
Our biggest goal was writes & uptime. We sometimes do over 150k writes/sec. We also needed it to be up and accepting writes even if one node goes down.
We regularly take nodes offline for updates/etc and cassandra never misses a beat.
We ~really~ wanted to use influxdb, but as a startup we couldn't justify the cost/benefit over Cassandra since we have 8 nodes for the DB. I just went to the influx site to try to find the pricing again and it seems to be hidden now :/
EDIT: As a PS, just remember every one of the influxdb benchmarks ( that I've come across ) are single node. Cassandra is meant to be horizontally scalable. Testing a single node Cassandra is like testing a racecar on your driveway...
For an excellent academic example, see this paper of Facebook's gorilla in memory TSDB: