Fly Machines: An API for Fast-Booting VMs
fly.io
fly.io
Right now I'm doing Postgres stuff (RDS) and dealing with taking 10+ minutes to boot a fresh instance. I'm tempted to try out fly.io and their Postgres clusters but I'm afraid I'd be spoiled and hate my life after (my job has me stuck in AWS for the interminable future).
I would be interested to know where all that time is being spent in on the AWS side. To be a fly on the wall seeing their full, unfiltered logging and metrics.
EC2 has historically not focused much on instance boot time. We did for GCE and drove it down pretty heavily. The post here from fly has a good set of sequence diagrams for "what are the various phases of creating an instance from scratch" that are generally applicable.
I'll note though that different users have different targets. Some people care about "time from request to first instruction ticks over" while others only care about "time from request to ssh'able from the public internet". There's an interesting middle ground of "time from request to being able to talk to other services like GCS or S3".
It's not clear to me what the networking / discovery story is for a Fly Machine that is stopped and then starts. That is, how long does fly-proxy take to update (globally? within a metro?) to add and remove the new Fly Machine? I vaguely recall that only external endpoints support IPv4, so I assume Fly is reserving and registering the internal IPv6 endpoints in the more expensive "create" step and then "start" is just about propagating liveness.
As a heavy GCE user, something weird about GCE to me is that instance boot times can be extremely variable — predictable within a given instance group, but unpredictable with even very small changes to e.g. instance sizes within the same instance family. (And this isn’t the instance blocking on getting scheduled onto a hypervisor; I can recognize that point, because it’s when any quota limits hit to potentially kill the instance provision. That phase of the delay is very stable.)
The variable delays also seem to apply to “reset” of the instance (which won’t involve an even-temporary deschedule, as reset keeps NVMe state) — but not to kexec-reboot, if one opts for that.
Is GCE built on a hybrid of two different hypervisor systems with wildly differing boot-time performance characteristics, where subtle tweaks of your instance config determine which hypervisor you get? Maybe one that relies on a hardware offload for something (PXE kernel signature verification?) that the other does purely in software?
One thing that’s clear to me is that instance types introduced after a certain point (e.g. all n2d instances) are always of the fast-booting type. So I’m guessing this is just old hypervisors being stuck on some legacy config because of long-term customers partially pinning those host machines with workloads that somehow prevent live migration (workloads with NVMe or GPU), and so make those hypervisors unable to be drained for a hard-upgrade.
This is the same target: a machine (that usually only has single app on it) shouldn't take more time to boot than a general-purpose consumer PC/laptop.
The reason it takes so darn long to start in so many cases is just how horrendously overcomplicated the whole cloud setup is internally and externally (sometimes for good reasons, sometimes because we don't know better, sometimes just because it really is just overly complicated and overengineered)
That's an incredibly easy target. VMs can and should boot much faster than that - just look at firecracker hypervisor.
Even with KVM, if you replace systemd with something small and simple [0] (which you totally should, for single-app VMs), boot times of couple of seconds are within reach.
I also remember clicking around Ling (Erlang on Xen, sadly no longer active [1]) where the whole VM could boot up, service the request, and shut down in less time than it takes a cloud to start spinning up an instance :)
only used by Lambda at the moment though, not in ec2
Open source is so cool when it works.
From the button of the comment: https://news.ycombinator.com/item?id=26747701
He is often on HN, look at the karma: https://news.ycombinator.com/user?id=tptacek
(YC-Founders take notes)
only further confirms my assumption that tptacek has an army of clones
although, as a fellow HN addict i wish i’ve had a private army of comment posters so i don’t have to
nice job on karma though, do you still remember at which point they start paying dividends? :D
now, that's my style
:-)
(Post by HNoutsourced.com, why not sign up today?)
I think this is what will make Fly really exciting. Right now (if I understand right) you need to be paying for a VM 24/7 in every region you want your app available in, because it only scales down to 1. So it runs apps in regions close to users that you're willing to pay for 24/7. If they make scale-to-zero work in every region, then maybe you can just make every app global and if you have some occasional users in Australia then it can just spin up over there while you're getting requests. I think it's what will make many-regions feasible for every app.
I honestly don't understand what's going on here. I thought we turned to Docker/containers because VMs were too heavy? Now we've got VMs that run Docker? (Not trying to be dense - what is the advantage?)
Why run Docker in VMs instead of using VM images? Because Docker's build tools are more popular than Packer-style tooling.
They provide any extra layer of indirection which helps with usual exploit attempts, but also introduce new scope. We've had exploits specifically targeting the namespaces API already.
Well, isn't that what happens when you put a shield into place? Someone tries to break it. Why have people concluded that it can never be made properly secure?
It's a moot point, because this is a solved problem. Use containers for single-tenant workloads; use micro-VMs, whichever flavor you like best, for multi-tenant.
If you're reading this and have big thoughts on how we might do GPUs at Fly.io without keeping us up at night about security, you should reach out; we're hiring.
The win would be in the attack surface area. For hypervisors there's a good layer of abstraction to pivot over, whereas with containers it's a much thinner wall.
The container kernel surface is just insane by comparison.
Why run daemons in containers instead of using proper process isolation? Because containers absolve the system administrator from understanding their systems.
VMs are lightweight now, though containers are lighter still; ref: https://fly.io/blog/sandboxing-and-workload-isolation/
Btw, Docker wasn't about security as much as it was about "package once, run anywhere".
> Now we've got VMs that run Docker?
Fly.io doesn't run Docker as-is, but rather unpacks it and runs it in a guest through containerd; ref: https://fly.io/blog/docker-without-docker/
> ...what is the advantage?
This has been discussed numerous times, and here's a link to one such discussion: https://news.ycombinator.com/item?id=26747701
Read also: https://gruchalski.com/posts/2021-03-03-thoughts-on-creating...
The tooling for proper VM creation on the other hand is in the stone-age comparatively--there are just a few tools like packer or a frankenstein of ansible scripts and neither are as nice or easy as Dockerfile creation.
Docker is much, much simpler. Write a Dockerfile that's mostly just bash or shell code. Build and run in seconds to immediately see the results. Once it's working push the container image to a public registry and you can distribute it to anything.
I am so excited about the future. We are seeing a bunch of announcements from multiple companies that make it possible for a single developer or small team to fairly cheaply run a global service without spending a whole lot of time on ops.
I am excited to see what people will come up with.
It’s one of the major pluses of the big clouds yet their pricing isn’t always awesome. Smaller player can help push that down.
See also the DO announcement today. Probably won’t use that but glad about it anyway
>"We're not done. You need something to run, right? Firecracker needs a root filesystem. For this, we download Docker images from a repository backed by S3. This can be done in a few seconds if you're near S3 and the image is smol."
I feel like I am missing something. If an S3 bucket is a requirement and I was interested in the isolation provided by Firecracker why wouldn't I just use AWS Fargate or Lambda which are both powered by Firecracker? If low latency was the concern, I can't imagine there being any lower latency than having my workload and storage being colocated in the same AWS Availability Zone.
Fargate and Lambda are not as consistently fast to boot VMs. Fargate, in particular, can take minutes to get a container launched.
This is not because they're bad services, it's because they make different tradeoffs than we do. When you ask for a Fargate container (or a new Lambda "instance"), AWS actually moves other containers/lambdas out of the way to get you running. Most of the wait time is their infrastructure doing orchestration magic to match your thing to their available compute.
Fly Machines don't do any of this. If you try and start a machine and there's no capacity for it, you get a very fast error response instead. This works well for our early customers. Most of them want to start a process quickly enough for a good UX. Fast errors give them a chance to do that.
I'd love to be able to supply some kind of memory snapshot in addition to the docker image to cut down on cold starts. Probably blocked on snapshot support in Firecracker according to another thread? Eagerly awaiting this since it could make Fly Machine the best of both worlds!
Not a fan of how Lambda makes me scale memory and compute in tandem, when my use case benefits so much more from compute than memory. I basically have to pay for 2+ gigs I'm never going to use to get the compute performance I want. Makes 0 sense.
My understanding is that Lambdas aren't ever really truly warm unless you have a completely steady traffic level. The first request in a while will hit the latency spike. But so will every increase in concurrency. So if you were serving 2 req/sec, and a 3rd concurrent visitor comes along then they will also get a cold start.
If you have a low-latency use case then fly.io's regular VMs are much better than either this or lambda. You get permanently running VM (the smallest of which is $2/month for 256mb/RAM), which can serve more than one request all by itself and will auto-scale with traffic.
For this use case though, I forgot to mention that it needs _much_ faster autoscaling than what Fly's regular VMs offer, with unbounded concurrency, and not ideal to run concurrently in a single VM due to each request being compute heavy and needing full isolation from each other since they run arbitrary customer code.
It's true that with Lambda, some amount of cold starts are probably inevitable with extreme spikes in traffic. But I'm hoping to mitigate most of that by sending artificial concurrent traffic on a schedule to keep a decent buffer of warmed up Lambdas above the current real traffic level. Still to be seen if that plan works out in practice.
Though I think ultimately it's going to be impossible to have all 3 of:
- Ability to deal with large traffic spikes
- Low-latency
- No provisioned resources
This is why I want to see a similar mode of operation for Fly Machines, possibly through memory snapshots, so I can manually provide a "warm" suspend state for it to unsuspend into. In fact this would be even better than the lambda model since there would be _no_ cold starts.
as far as i understand this will let me run VMs with specified Docker images?
i'm thinking of using something Fly.io to offer a dedicated hosting for my upcoming product, so when the customers sign up they get a new machine with an individual endpoint
the workload that needs to be running on those machines is quite intensive (like crawling web pages) and not very scalable when sharing resources
also can you give more details about your Nomad stack?
i was actually thinking of using Kubernetes or Docker swarm as API to deploy these workloads
Machines are designed to work well for your customer hosting! You can install a machine for them, and then turn it on when they push a button, or have it turn on automatically when they visit a URL.
I'm happy to talk about it more. Feel free to send me/us an email!
btw., i've sent a mail already (mish at ushakov), but only got an automated message
i'm already using free fly.io for wikinewsfeed.org and very happy so far
would be awesome if you could send me a tip how i could use fly.io to deploy instances for my customers?
in my use-case i want to run many instances of Chrome at the same time
running like 100x instances of Chrome on a single machine is too resource-intensive and having a bigger host machine won't do the job, so your only option is to have a dedicated vm for each Chrome instance
https://github.com/drifting-in-space/spawner
check out the demo: https://www.youtube.com/watch?v=aGsxxcQRKa4
That looks promising, but I don't want to handle any of this myself :/
I want a service where I can start a new VM fast by posting to an API, have it run some long running JS server code, and the VM should close itself when it's done.
My use case is often CPU bound, so a small VM with a single CPU is just fine.
they're in private beta
...still waiting for a reply 2 years on ;)
E.g. if I need to run a bunch of processing, would it be A) spin up the micro-VM and pull from a queue service B) embed SQLite C) use some kind of in-memory store
TBH I've been waiting for years for someone to do 'firecracker as a service'. I must have searched that exact term about once per month.
For example, we have a queue that handles video encoding. I would like to have 0-N encoders running at the same time, based on demand.
Spin up time is important as well, since I typically provide test renders triggered from the UI.
You set a min and max number of VMs to run. Set the watermark for users per machine. And they automatically scale up and down depending on the number of connections. It will even figure out where in the world all the connections are coming from, and spin up the new VM in a region near to them.
So if your video encoding was requested via HTTP(S) connections then this would be trivial.
Edit: the meme goes something like "no one asks follow up questions if you say you're an accountant"
https://www.washingtonpost.com/technology/2022/04/08/algospe...
> our apps platform (orchestrated by Nomad) already does it
Preserving internal vm state or just restart everything?
LayerCI / WebApp.io did mention they migrate VMs... Fly.io should steal that tech https://news.ycombinator.com/item?id=25980897
1. Lambda has runtime constraints and is billed per request
2. Fly Machines are basically just VMs with some magical startup sugar. You can do pretty much anything you want with them, including run for months.
Fly Machines are a partial answer to "how can I build my own Lambda service"?
Deploy App Servers
Close to Your Users
is there a timeout limit to functions? This piques my interest but I can't tell if fly is a serverless function provider or some way to deploy my docker closest to my user (which is what I am looking for right now)What guarantee is there in terms of average latency for my users? Is there a looking glass of sort where I can ping/see all the locations where my docker images will be running?
a dedicated 4-core 8gb ram is $124.00/month which is 4~6x more expensive than running on KVM vps so I want to know what I am signing up for
edit: I see the list of locations and it makes me think, aren't I already doing what fly.io is doing? I spin up a VPS instance at one of the locations that is closest to my user. It takes about 30~120 seconds. It's far far cheaper
Most people run "apps" on Fly. You control which regions they run in, we load balance to the nearest. We have guides for launching some frameworks here: https://fly.io/docs/getting-started/
The difference between us and a VPS provider is: your app runs in as many of those regions as you want. So does your database, if you're using our Postgres. And we route writes to the appropriate place: https://fly.io/blog/globally-distributed-postgres/
The other difference is probably CPU. The 4 cpu, 8GB RAM instances are 2 dedicated AMD EPYC cores + hyper threads. They're relatively expensive. You may not need them! VPS providers typically run cheaper CPUs and over provision them.
We're shipping shared CPU options with more memory soon. They should be closer to what you see from VPS providers, though still more expensive.
We managed to hack a solution which uses pipes for IPC but it would be nice not to have to do this.
You could already run multiple processes in a guest vm [0] or run multiple guests vms on the same host [1]. If the guest kernel Fly boots your app into doesn't have required modules, I guess you could consider requesting (nerd-sniping 'em) specifically for those, like so: https://twitter.com/dave_universetf/status/14262218974072422...
Fly VM's are just VM's that can start quickly.
They don't seem to currently support a request/response based VM lifecycle.
If you wanted to use a fly VM in a lambda-like way, seems like you would need to have some kind of proxy to coordinate the work, ie start the VM via the API, have your VM process start a web server, once it's booted, send it a HTTP request, once the request is finished, shut down the VM via the API.
Also seems like fly can't suspend a running process, your process needs to start up every time you start a fly VM.
Lambda will suspend a VM between requests, keeping the process in memory for a few minutes.
Sending a subsequent request to a warm lambda is much faster then booting a lambda from scratch, particularly for JIT based language runtimes.
Lmao props to the team for getting this copy out unsanitized by (potentially) unchill bosses.
We're painfully aware that we get a limited number of bites at the HN apple, and we try to spend those on things like Litestream.io, which is an open source project that benefits people who won't ever use Fly.io. Several of our last few blog posts were about stuff we've done "wrong"; so, we've also got no qualms about charting on HN with a post about how much trouble we've had with Raft, or user-mode WireGuard.
Dan Gackle has said a bunch of times that he wishes more companies got the lovey-dovey reaction we seem to get from HN. I've got the cheat codes, if you want them: write posts for the HN audience, and throw your marketing goals out the window. I'm not going to bullshit you and say that we don't benefit from those kinds of posts too, but I hope it's at least clearer why they're received more warmly than a lot of tech company product announcements: we don't write them to be product announcements. (Unlike this post!)
If there was a "No HN" meta tag we could set on our posts, this post would have had it.
Ordinarily I'd be squeamish about dragging us into metacommentary like this on one of our stories, because I'd rather argue about whether you can scale a modern full stack app entirely on SQLite than about our marketing. But, like I said, we've got no skin in how this post ranks here.
To me this actually also suggests that the algorithm has a quality-score for both you as users and fly.io as a domain.
So both your style & the algorithm have like a flywheel with HN. (Makes sense to me.)
* Small, scrappy team taking on the big incumbents that everybody hates (AWS)
* Heavy focus on efficiency / latency
* Deep expertise in their areas of focus
* Lots of Rust
* Chill, non-corporate-y writing style
I just wish I knew about this earlier because from what I read, I think we at Devbook [1] built pretty similar service for our product. We are using Docker to "describe" the VM's environment, our booting times are in the similar numbers, we are using Nomad for orchestration, and we are also using Firecracker :). We basically had to build are own serverless platform for VMs. I need to compare our current pricing to Fly's.