Monitoring Elixir Apps on Fly.io with Prometheus and PromEx
fly.io
fly.io
Would you mind elaborating. I'm definitely looking for a PasS for Elixir/Phoenix + Postgres.
That said, our stack does make things Elixir folks need very easy. In particular:
* Private network and built in service discovery mean clustering just works
* You can connect to your private network with WireGuard and then use iex or similar to inspect running Elixir processes: https://fly.io/blog/observing-elixir-in-production/
* Postgres setup is pretty magical. "fly postgres attach" gets your app all connected and auth'ed to your Postgres cluster.
* And, most importantly, we actually are the best place to run an Elixir app if you're using Phoenix LiveView. You can push your app servers out close to people and make LiveView incredibly responsive: https://liveview-counter.fly.dev/
One slightly useful trick that the docs don't highlight is that you can run the /metrics endpoint on a different port to the rest of the application. If you have that firewalled off, a local grafana agent or prom shipper of your choice can happily run against your application without making metrics public.
That standalone HTTP server in PromEx can also be used to expose metrics in non-Phoenix applications. In the coming weeks you'll also be able to run GrafanaAgent in the PromEx supervision tree so you can push Prometheus metrics via remote_write. Stay tuned ;)!
I understand you can put your application server in any location but generally there is only one storage so are these application servers doing cross region database calls?
Having only worked with single cluster setup web apps, I am always curious about this part.
Is the answer always - use a managed replicated database and send read queries to one near your location and all write queries goes to the primary instance?
We tried cross region database calls for HTTP requests. They were bad. They seem to work ok for long lived connections like websockets, though.
The entire team is active on forums, and the response time are fast and issues get fixed and deployed in hours.
Running a Phoenix LiveView on a fly instance close to your users is the closest you can get to an SPA experience without going full JS.
FWIW these are OCI images, which can be built with other tools besides Docker. These are not the same thing as containers. I think Docker is overrated and am glad to be using a Containerfile instead of a Dockerfile (it's the same format).
Firecracker is great but there is an alternative: https://katacontainers.io/
There seems to be confusion as to what Firecracker is doing. Docker in Docker isn't really Docker in Docker when it's in a MicroVM - except for being run by an OCI image. https://community.fly.io/t/fly-containers-as-a-ci-runners/10... Docker uses LXC containers, and an LXC container inside an LXC container is likely to have performance problems. MicroVMs, on the other hand, are similar to VPSes.
fly.io suggests that "CPU cores, for instance, should only ever be doing work for one microVM". https://fly.io/docs/reference/architecture/ These MicroVM's are not cheap, and maybe it would be nice to be able to share a VCPU between multiple containers in a lot of situations. Fly.io could be burning a lot of investor money.
I think WASM and Deno are good ways to break down a MicroVM with a whole VCPU into smaller sandboxed entities. Also Docker run by an OCI image makes more sense in a MicroVM that has a whole CPU than it does in an LXC container with less than a whole CPU.
I do think fly.io is pretty cool. I hope they find ways to educate people more about the difference between a MicroVM (not invented by AWS) and a LXC container and to get people making more use of their heavier weight "containers".
Edit: I noticed that in their free tier and low cost plan, they have shared CPUs. https://fly.io/docs/about/pricing/
They are charging quite a bit more for one dedicated CPU than DigitalOcean, Vultr, and Lightsail. The cheapest one is $31.00/mo.
It's quite similar to Google Could Run. It looks like Cloud Run doesn't have an option to run less than a vCPU in memory though. That means you can pay fly.io $6/mo for a MicroVM with a gig of memory and not have a full cold start, but have a shared CPU which will of course cause latency but probably typically less than a cold start, but you can't pay Google $6/mo and get that. https://cloud.google.com/run/pricing Also websockets can stay open with a shared CPU, but can't in between the VM stopping and starting. I don't know if the $6/mo 1GB MicroVM with a shared vCPU is sustainable, but if so, it's a pretty impressive way to host an app that uses websockets for super cheap.
Edit 2: I said "They are charging quite a bit more for one dedicated CPU than DigitalOcean, Vultr, and Lightsail." - I was actually wrong here. It is more for self-hosting if you would run a database like PostgreSQL/MySQL in a VPS (they don't discourage it, and it's pretty simple to do) but wouldn't run a database like PostgreSQL/MySQL in a Fly.io VM (they might discourage it, and it's a bit tricky to do).
What does it offers over firecracker?
We picked Firecracker because the smaller scope makes it easier to understand and, thus, trust for our particular workloads.
In fact, a Fly Firecracker VM is not all that different from a DO droplet. We expose them as containers, but you could run Dokku within a Fly app and have a similar set of stuff under the covers.
In general, we treat "vCPU" as a single hardware threads, which is pretty common. Our hosts use Epyc CPUs with two threads each, which is also pretty common.
So a single CPU dedicated VM on Fly is equivalent to owning a hardware thread, which makes $31/mo comparable to other places. These are roughly the same as DigitalOcean's "general purpose" droplets.
Basic Droplets, and most Vultr VMs, use shared CPUs.
Providers use wildly different CPUs too. You'll sometimes find people surprised how fast Fly VMs are because we standardized on Epyc CPUs pretty early. Much of what you can buy runs on either older Intel or consumer grade processors. Which is actually fine! The people who buy our dedicated CPU VMs just so happen to need a lot of power because they're frequently transcoding or generating images.