A simple way to get more value from metrics
danluu.com
danluu.com
I've done this in fintech a few times already and the best stack that worked from my experience was telegraf + influxdb + grafana. There are many things you can get wrong with metrics collection (what you collect, how, units, how the data is stored, aggregated and eventually presented) and I learned about most of them the hard way.
However, when done right and covering all layers of your software execution stack this can be a game changer, both in terms of capacity planing/picking low hanging perf fruit and day to day operations.
Highly recommend giving it a try as the tools are free, mature and cover a wide spectrum of platforms and services.
This might be obvious to some but I thought it worth mentioning.
Can you share the important things that are overlooked?
I'm in the middle of doing a test-run of vmagent + victoria-metrics as a mostly-replacement (see: not replacing the cases where it's known exactly what's wanted, and used explicitly for that), and victoria-metrics fully support telegraf pushes which is a big bonus.
Also they're apparently going to be dropping InfluxQL entirely, for Flux, which is weird to me? The docs don't outright state this, but InfluxDB 2.0 has only Flux examples in the docs I can find. It's a better language, however the learning curve is not trivial.
Tickscript is gone in 2.0, though. I'm not in love with Flux, but I certainly won't miss tickscript.
Yes. The learning curve was steep, and debugging was near-impossible. Flux is rough too, but it's already getting better documentation and tooling support than Tickscript had.
- the context that the author already has a very successful career as a well-known developer
- the humility he evidences in most posts on his blog
- the fact that he explicitly highlights the work of others in this post alongside his own
I really don't think Dan is doing this as any form of personal marketing. He has no need of personal marketing, his blog already has several million views per month and frequently shows up on HN as it is, and it isn't really his style.
The post about salary reads much better so might just be an experience thing.
"I did it by myself in one day, well actually it was one week but had I known the stack I would likely have done it in one day. Oh, and by the way, after that week, there was yet another month of work involving at least two other persons from my team and then even more work from other teams. But let's not dwell on boring details".
It's nearly as infuriating as the "Appendix: stuff I screwed up" which doesn't contain actual screw up. It's a shame because the rest of the writing is interesting and doesn't need to be propped up.
Were you able to move past this to read the rest of the article? Because it's a very good article.
If that point needed to be made (and I don't think it really did in this article, that's not the focus), it could have been done more carefully.
Anyone interested in this topic might want to check out an issue thread on the Thanos GitHub project. I would love to see M3, Thanos, Cortex and other Prometheus long term storage solutions all be able to benefit from a project in this space that could dynamically pull back data from any form of Prometheus long term storage using the Prometheus Remote Read protocol: https://github.com/thanos-io/thanos/issues/2682
Spark and Presto both support predicate push down to a data layer, which can be a Prometheus long term metrics store, and are able to perform queries on arbitrary sets of data.
Spark is also super useful for ETLing data into a warehouse (such as HDFS or other backends, i.e. see the BigQuery connector for Spark[1] that could write a query from say a Prometheus long term store metrics and export it into BigQuery for further querying).
[1]: https://cloud.google.com/dataproc/docs/tutorials/bigquery-co...
I might be a great telecomm tech, a genius even, but once I'm out of a job, I can't build my own telecomm system - that would cost billions. I have to go back to some other telecomm system to start making money again.
But, at a startup, a kicked-out senior engineer can actually pretty much exactly recreate the company; they can do the equivalent of a laid-off telecomm employee starting a new, almost-as-good (except for branding) telecomm company.
No billions in infrastructure required: within a month or two, the cloned company could be near-indistinguishable from the original.
So companies have to pay employees more like partners, instead of employees, because either they pay them as equals or they'll be forced to compete against them, as equal rivals.
Title didn't live up to article imho. But I get it. Thanks for sharing your methods.
and anyway:
The missing piece remains to be able to query one system with the query language of the others. For example, query Prometheus using Graphite's query language.
We killed our influx cluster numerous times with high cardinality metrics. We migrated to Datadog which charges based on cardinality so we actively avoid useful tags that have too much cardinality. I’m investigating Timescale since our data isn’t that big and btrees are unaffected by cardinality.
It extends very well to something that we constantly hammer home on my team: using boring tools is often best because it's easier to manage around known problems than forge into the unknown, especially for use-cases that don't have to do with your core business. Extreme & contrived example: it's much better to build your web backend in PHP over Rust because you're standing on the shoulders of decades of prior work, although people will definitely make fun of you at your next webdev meetup.
(Functionality that is core to your business is where you should differentiate and push the boundaries on interesting technology e.g. Search for Google, streaming infrastructure for Netflix. All bets are off here and this is where to reinvent the wheel if you must – this is where you earn your paycheck!)
The single-binary approach is still a problem, though. In my mind any serious telemetry collection stack should separate the query engine and ingestion path from each other - Prometheus has both the query interface and the ingestion/writing subsystem in the same process.[ß]
As for the parent poster: you certainly want to push telemetry out on every event, but the mechanism has to be VERY lightweight. With prometheus the solution is to have a telemetry collection/aggregation agent on the host, feed it with the event data and have prometheus scrape the agent. Statsd with the KV extension is a great protocol for shoveling the telemetry out from the process and into the agent.
ß: you can get around this with Thanos + Trickster to take care of the read path only, but it's quite a bit more complex than plain Prometheus.
Taking your example you could push without sending a packet on every event by instead accumulating a counter in memory, and pushing out the current total every N seconds to your preferred push-based monitoring system. You could even do this on top of a Prometheus client library, some of the official ones even as a demo allow pushing to Graphite with just two lines of code: https://github.com/prometheus/client_python#graphite
In my personal opinion, pull is overall better than push but only very slightly. Each have their own problems you'll hit as you scale, but those problems can be engineered around in both cases.
Disclaimer: Prometheus developer
Or any good resource which discusses possible optimizations in the infra stack at a more theoretical, abstract, generalizable level?
I feel like I have an inception. Should "boring, descriptive, names" be the default in all IT?
The bit about not being able to use columns for each metric because there were too many ....
the classic solution is to have a column called “metric name” and another for “metric value”.
Can’t spot why they didn’t just do that.
Does a column store like paraquet make a good time series DB? Trendy named time series databases I’ve had the displeasure of using would all fail miserably by high cardinality series too, so I’m not convinced there is actually a better thing than files on a lake for this stuff.
So, use some format to name the metric in each row. If paraquet, use dictionary encoding on that column and sort or cluster the rows ... will give go min/max pruning etc.
But presto is currently 5x slower to chew through paraquet vs orc so perhaps simply use orc. Or, for this data, Avro or json lines.
And then when you’ve used presto to discover interesting metrics you can always use presto (or scalding or whatever your poison is) to extract the metrics you have identified you want to examine more closely and to put them into separate datasets etc.
I’m just outlining standard approach’s to these kinds of problems.