Fly’s Prometheus Metrics
fly.io
fly.io
Google has long since abandoned the Borgmon data model for histograms with monarch. The closest non google implementation is probably circonus. Sadly neither is available as open source software.
I can’t really blame fly for not individually building an open source modern metric db. But it’s sort of sad that the infra team I’m most impressed with has to use metric systems from 15 years ago when the rest of their stack is so cutting edge.
They have a self hosted option but it’s not free or open source.
Google's SRE stack is lightyears ahead of anything else out there simply because they're at the scale where they can afford to hire dedicated developers to just write internal ops software where most other SREs are understaffed and overworked just on operational projects.
https://static.sched.com/hosted_files/promcononline2021/33/2...
> We use Lets Encrypt to issue certificates, and donate half of our SSL fees to them at the end of each calendar year.
Any chance of an RSS feed on that blog? [e: ah it's just missing the indicator meta tag thingy, https://fly.io/feed.xml]
I have some pretty important services I run on there and my availability has beaten anything I’ve run on GCP.
I'm not a heavy cloud services user, so the prices aren't that important to me, but fly.io seems to be on par with other providers.
I think of Heroku as fly's most direct competitor. Heroku charges $50/mo for a dedicated VM with 1 GB RAM, whereas fly charges $31/mo for dedicated VM with 2 GB RAM.
I find their billing dashboard leaves a bit to be desired, I'm still very nervous I'm going to get a big unexpected bill.
There were some initial teething issues as Fly only recently added UDP support, but they were very responsive to the various bugs I reported and fixed them. My name servers have been running problem free for several weeks now.
The Fly UX via the flyctl command-line app is excellent, very Heroku-like.
For apps that need anycast the only real alternative to Fly that I found is AWS Global Accelerator, but it limits you to two anycast IP addresses, it's much more expensive than Fly and you're fighting the AWS API the whole time.
:(
Not actually Prometheus-compatible, sloppy code, spotty docs. I have no idea why this dumb product continues to attract users.
https://prometheus.io/blog/2021/05/04/prometheus-conformance...
> Telegraf, which is to metrics sort of what Logstash is to logs: a swiss-army knife tool that adapts arbitrary inputs to arbitrary output formats. We run Telegraf agents on our nodes to scrape local Prometheus sources, and Vicky scrapes Telegraf. Telegraf simplifies the networking for our metrics; it means Vicky (and our iptables rules) only need to know about one Prometheus endpoint per node.
Normally you just use a regular Prometheus server to do this. Why add another, different technology to the stack?
> We spent some time scaling it with Thanos, and Thanos was a lot, as far as ops hassle goes.
It really isn't -- assuming you're not trying to bend Prometheus into something it isn't. Prometheus works using a federated, pull-based architecture. It expects to be near the things it's monitoring, and expects you to build out a hierarchy of infrastructure, in layers, to handle larger scopes.
This is structurally different to what I'll call the "clustering" model of scale, where you have all your data sources pushing their data, aggregating maybe on the machine or datacenter level, but then shuttling everything to a single central place, which you scale vertically from the perspective of your users. This appears to be what you want to do, based on the prevalence of push-based tech in your stack.
Prometheus doesn't work this way. Some people really want it to work this way, and have even created entire product lines that make it look as if it works this way (Cortex, M3db) but it's fundamentally just not how it's designed to be used. If you try to make it work this way yourself, you'll certainly get frustrated.
Our physical hosts have hundreds of services exporting metrics. And many of those exported metrics are from untrusted sources. So we can both rewrite labels and decrease the scrape endpoint discoverability problem by aggregating them in one place.
> Not actually Prometheus-compatible, sloppy code, spotty docs. I have no idea why this dumb product continues to attract users.
Because it works incredibly well, it's easy to operate, and handles multi tenancy for us.
OK, but Prometheus can do all of this just fine?
> Because it works incredibly well, it's easy to operate, and handles multi tenancy for us.
Again, Prometheus itself ticks all of these boxes, too, if you're not trying to force it to be something it's not.
There's an interesting discussion to be had about how our infrastructure works; for example, in the abstract, I'd prefer a "pure" pull-based design too. But things appear and disappear on our network a lot, and remote write simplifies a lot of configuration for us, so I don't think it's going anywhere.
I think you're reading a critique of Prometheus that isn't really present in what we're writing. Prometheus is great! Everyone should use it! Our needs are weird, since we're handling metrics as a feature of a PAAS that we're building.
I'm observing that you've used pull-based, horizontally-scaled tools to build a push-based, vertically-scaled telemetry infrastructure. It can be made to work, sure, but the solution is an impedance mismatch to the problem.
That pretty much works now?
I see the ideological purity case you two are making for "true Prometheus", but it is not at all clear to me how doing a purer version of Prometheus would make any of our users happier.
What am I missing?
Well, with the requisite glue code that would inform each user's Prometheus instance how to scrape the service instances -- yes, more or less.
> is not at all clear to me how doing a purer version of Prometheus would make any of our users happier.
If the only things you care about when you build systems are "works" and "direct impact on customers" then there's not really a point to this conversation. The things I'm speaking about, the architectural soundness of a distributed system, are largely orthogonal to those metrics, at least to the first derivative.
But these are subjective claims! Not everyone thinks the same way!
Exposing a metrics endpoint for customers is nice. How do you manage the cardinality? I haven't used Victoria before, is it just better at high cardinality time series?
FWIW, we are in the process of considering it as a single monitoring plane for multiple clusters with high cardinality metrics
There's a third, "events". Just push an event out whenever something interesting happens, and let the monitoring tool decide whether to count, aggregate, histogram, alert, etc.
Events require less code in the app (no storage, no aggregation, no web server), and allow more flexibility in processing. I have used events to great effect. I am baffled as to why monitoring people still only talk about metrics.
Huh? Who gets excited about writing a parser?
What was wrong with "${key} ${value}" on separate lines?
I’m trying to understand the market they’re operating in. Big ol enterprises would probably want to run on AWS/GCP right? So would startups? What’s the long game? Genuine question.
It's pretty great, we're happy plus subscribers.
(I am a Cortex maintainer)
>if you’ve got a Docker container, it can be running on Fly in single-digit minutes.
I used to laugh at the old Plan 9 fortune, "... Forking an allegro process requires only seconds... -V. Kelly". Guess I'm not laughing anymore?
FWIW, performance of components is the barrier to composition in system design and development. You can't compose modules that take seconds to act, and still have something that is usable real-time.
In whichever case, that doesn't mean we can't compose them into a performant product as the OP suggests, it just means you run them as daemons so you can amortize the startup cost across many invocations. This isn't specific to containers--we do the same thing for web servers, databases, virtual machines, physical machines etc. Anything with a startup cost that you don't want to pay each time.
I’m all for the “k8s is not erlang” position, but in fly’s case it seems like the right tool for the job. Fly is much faster than a new EC2 instance.
Oh boy!
> None of us have ever worked for Google, let alone as SREs. So we’re going out on a limb
Oh.... boy.
> We spent some time scaling it with Thanos, and Thanos was a lot, as far as ops hassle goes.
You know, they have these companies now, that will collect your metrics for you, so that you don't have to deal with ops hassle.
> In each Firecracker instance, we run our custom init,
... in Rust. Yes, the thing that is normally a shell script, is now a compiled program in a new language, that mostly just runs mkdir(), mount() and ethtool(). (https://github.com/superfly/init-snapshot/blob/public/src/bi...). In a few years, when that component is passed off to a dedicated Ops team, and they find it hard to hire a sysadmin who also knows Rust, there will be some poor intern who learned Rust over the summer whose job is to rewrite that thing back into a shell script.
It looks like your init is running the gamut of typical linux init steps: handling signals, spawning processes, handling TTYs, running typical system start-up commands. Then it gathers host and network metrics, reports results, and operates as some kind of websocket server? All of that other stuff should live in a separate dedicated program that the init script runs. Unless I missed something and that's not possible?
We wrote a little more about this here:
I say it will eventually be rewritten as a shell script (minus the stats, which will be replaced by a monitoring agent) because people messing with embedded Linux often try writing their own compiled init to bundle with an initrd. It eventually becomes a hassle, so they either make their own feature-filled init replacement, or they go back to a shell script. Most go for the shell script.
I mean, it's a DSL for composing programs. It's too powerful for what you need to do, but all its other advantages (portability, flexibility, simplicity, tracing, environment passing, universal language, rapid development, cheaper support, yadda yadda) make it a win over time. The only reason I can see not to use a shell script is if speed is your highest priority.