Fly.io has GPUs now
fly.io
fly.io
I do wish they had some features I need, but their support and responses are top notch. And I've lost much less hair and time than I would going full-blown AWS or another cloud provider.
Digital ocean has been terrible for me, some regions just go down every month and I lose thousands of requests, increasing my churn rate.
Fly.io had tons of weird issues but it got better in the last months. It's still very incomplete in terms of functionality and figuring out how to deploy the first time is a massive pain.
My plan is to add Hetzner and load balance with bunnycdn across DO and H
So yes, a lot of people on HN complain about fly's reliability. fly posts to HN a lot and gives them the opportunity. Is it actually meaningful compared to the alternatives? It's very hard to tell.
First: this is 100% a "live by the sword, die by the sword" situation for us. We're as aware as anybody about our weird HN darling status (this is a post from two months ago, about an announcement from many months ago, that spent like 12 hours plastered to the front page; we have no idea why it hit today, and it actually stepped on another thing we wanted to post today so don't think we secretly orchestrated any of this!). We've allowed ourselves to be ultra-visible here, and threads like this are natural consequence.
Moreover: a lot of this criticism is well warranted! I can cough up a litany of mitigating factors (the guy who stored his database in ephemeral instance storage instead of a volume, for instance), but I mean, come on. The single most highly upvoted and trafficked thing we've ever written was a post a year ago owning up to reliability issues on the platform. People have definitely had issues!
A fun cop-out answer here is to note all the times people compare us to AWS or Cloudflare, as if we were a hyperscaler public cloud. More fun still is to search HN for stories about us-east-1. We certainly do that to self-sooth internally! And: also? If your only consideration for picking a place to host an application is platform reliability? You're hosting on AWS anyways. But it's still a cop-out.
So I guess I'd sum all this up as: we've picked a hard problem to work on. Things are mathematically guaranteed to go wrong even if we're perfect, and we are not that. People should take criticisms of us on these threads seriously. We do. This is a tough crowd (the threads, if not the vote scores on our blog post) and there's value in that. Over the last year, and through this upcoming year, staffing for infra reliability has been the single biggest driver of hiring at Fly.io, I think that's the right call, and I think the fact that we occasionally get mauled on threads is part of what enabled us to make that call.
(Ordinarily I'd shut up about this stuff and let the thread die out itself, but some dearly loved user of ours took a stand and said they'd never had any problems on us, which: you can imagine the "ohhhhh nooooooo" montage that took place in my brain when I read that someone had essentially dared the thread to come up with times when we'd sucked for some user, so I guess all bets are off. Go easy on Xe, though: they really are just an ultra-helpful uncynical person, and kind of walked into a buzzsaw here).
When it works, it's brilliant. The problem is that it hasn't worked too well in the last few months.
/s
> it's one person who had issues
Issues specific to an application or one particular account have to be addressed as special cases (like any NewCloud platform, Fly.io has its own idiosyncrasies). The first step anyway is figuring out just what you're dealing with (special v common failure).
> looks like a theatre
I have had the Fly.io CEO do customer service. Some may call it theatre, but this isn't uncommon for smaller upstarts, and indicative of their commitment, if anything.
Are you talking about fly postgres? Because I use it and feel they've been pretty clear that it's unmanaged.
huh? it does what it says on the tin. nothing crazy about it.
They spell out for you in detail what they offer: https://fly.io/docs/postgres/getting-started/what-you-should...
And suggest external providers if you need managed postgres: https://fly.io/docs/postgres/getting-started/what-you-should...
If you are offering a service like Fly I think the database should be managed personally, the whole point of Fly.io is to provide abstractions to make production simpler.
Do you think the type of user who is using fly.io is interested in or capable of managing their own Postgres database? I'd rather just trust RDS or another provider.
Honestly.. kinda, yeah
At least I'm projecting my weird "I want to love you for some reason, Fly" plus my skillset onto anyone else that wants to love Fly too haha
They feel very developer/nerd/HN/tinkerer targeted
and they have a managed offering [2] in private beta now...
> Supabase now offers their excellent managed Postgres service on Fly.io infrastructure. Provisioning Supabase via flyctl ensures secure, low-latency database access from applications hosted on Fly.io.
[1] https://fly.io/docs/postgres/getting-started/what-you-should...
When I tested it I was hoping for at worst early termination of old connections with no dropped new connections and at best I expected them to gracefully wait for old connections to finish. But nope, just a full downtime switch over every time. But then when you think about the network topology described in their blog posts, you realize theres no way it could've been done correctly to begin with.
It's very rare for me to comment negatively on a service but that fact that this was the case paired with the way support acted like we were crazy when we sent video evidence of it definitely irked me for infrastructure company standards. Wouldn't recommend it outside of toy applications now.
> It feels like it's compelling to those scared of or ignorant of Kubernetes
I've written pretty large deployment systems for kubernetes. This isn't it. Theres a real space for heroku-like deploys done properly and no one is really doing it well (or at least without ridiculously thin or expensive compute resources)
https://community.fly.io/t/fly-io-support-community-vs-email...
Have you tried Google Cloud Run(based on KNative) I've never used it in production, but on paper seems to fit the bill.
It's in a weird place between heroku and lambda. If your container has a bad startup time like one of our python services, autoscaling can't be used as latency becomes a pain. Its also common deploy services on there that need things like health checks (unlike functions which you assume are alive), this assumes at least 1 instance of sustained use as well, assuming you do minute health checks. Their domain mapping service is also really really bad and can take hours to issue a cert for a domain so you have to be very careful about putting a lb in front of it for hostname migrations.
I don't care right now but the fact that we're paying 5x in compute is starting to bother me a bit. A 8core 16gb 'node' is ~$500/month ($100 on DO) assuming you don't scale to zero (which you probably wont). Plus I'm pretty sure the 8 cores reported isn't a meaty 8 cores.
But its been pretty stable and nice to use otherwise!
I do get that it is a bare server, but if you deploy even just bare containers to it, you would be saving a good bit of money and get better performance from it.
There is also a $63/month option that is significantly worse.
Our solution was to migrate the service to Kubernetes using an HPA scaling on the number of un-acked messages in the subscription, and then use a pull subscription to ensure reliable delivery (if the service is down they just sit in the queue rather retrying indefinitely).
I'm convinced Cloud Run/Functions are only useful for trivial HTTP workloads at this point and I rarely consider them.
But sweet sweet github triggered deploys. Have you found an easy solution to this?
Triggered deploys to Kubernetes you mean? There's a million ways to solve this problem for better or worse. We use Gitlab CI so we invoke helm in our pipelines (I'm sure there's a way to do this with github actions), but there's also flux cd, argo, etc. etc.
We use Kubernetes (GKE) elsewhere so we already had this machinery in place luckily. I can see the appeal of CloudRun/Functions as a way to avoid taking that plunge
Starting containers on Cloud Run is weirdly slow, and oh boy, how expensive that thing is. I'm getting the impression that pure VMs + Nomad would be a way better option.
What is this about? I assumed a highly throttled cpu or terrible disk performance. A python process that would start in 4 seconds locally could easily take 30 seconds there.
As a long time Nomad fan (disclaimer: now I work at HashiCorp), I would certainly agree. You lose some on the maintenance side because there's stuff for you to deal with that Google could abstract for you, but the added flexibility is probably worth it.
You need blackbox HTTP monitoring right now, don't ever wait for your customer to tell you that your service is down.
I use Prometheus (&Grafana), but you can also get a hosted service like Pingdom or whatever.
I was very excited about Fly originally, and built an entire orchestrator on top of Fly machines—until they had a multi-day outage where it took days to even get a response.
Kubernetes can be complex, but at least that complexity is (a) controllable and (b) fairly well-trodden.
Or to clarify your comment, Kubernetes on which cloud? Amazon? google? Linode?
I definitely understand the comparison between Kubernetes and fly. You have couple apps that are totally unrelated, managed by separate teams, and you want to figure out how you can avoid the two teams duplicating effort. One option is to use something like fly.io, where you get a command line you run to build your project and push the binary to a server. Another option is to self-host infrastructure like Kubernetes, and eventually get that down to one command to build and push (or have your CI system do it).
The end result that organizations are aiming for are similar; developers code the code and then the code runs in production. Frankly, a lot of toil and human effort is spent on this task, and everyone is aiming to get it to take less effort. fly.io is an approach. Kubernetes is an approach. Terraform on AWS is an approach.
That’d be a slightly more valid comparison albeit flyctl is much less ambitious by choice and design. That said, using flyctl to orchestrate your deployments is not the only way to Fly. Example:
The Fly team has worked on solving similar problems to Kubernetes. Ex://fly.io/blog/carving-the-scheduler-out-of-our-orchestrator/
Of course, Fly also provides the underlying infrastructure stack too. If you want to be pedantic, you can compare it to GKE/AKS/EKS.
Kubernetes on any major cloud platform is more mature, controllable, and reliable than Fly.
It looks worse than AWS or Azure to me.
Never used the service, but based on what I hear, I'll never try...
Why should I choose Fly? How come they are so prominent on hackernews? Are they backed by VC and get their default 400 upvotes by backers? I get the impression that Fly posts here are kind of sponsored.
If anyone has any questions, fire away!
But who is the target user of this service? Is this mostly just for existing fly.io customers who want to keep within the fly.io sandbox?
Frankly, I also see the other part of it as a way to ride the AI hype train to victory. Having powerful GPUs available to everyone makes it easy to experiment, which would open Fly.io as an option for more developers. I think "bring your own weights" is going to be a compelling story as things advance.
What have you learned from the exploration?
The application server uses Deno and Fresh (https://fresh.deno.dev) and requires a shared-1x CPU at 512 MB of ram. That's $3.19 per month as-is. It also uses 2GB of disk volume, which would cost $0.30 per month.
As far as post generation goes: when I first set it up it used GPT-3.5 Turbo to generate prose. That cost me rounding error per month (maybe like $0.05?). At some point I upgraded it to GPT-4 Turbo for free-because-I-got-OpenAI-credits-on-the-drama-day reasons. The prose level increase wasn't significant.
With the GPU it has now, a cold load of the model and prose generation run takes about 1.5 minutes. If I didn't have reasons to keep that machine pinned to a GPU (involving other ridiculous ventures), it would probably cost about 5 minutes per day (increased the time to make the math easier) of GPU time with a 40 GB volume (I now use Nous Hermes Mixtral at Q5_K_M precision, so about 32 GB of weights), so something like $6 per month for the volume and 2.5 hours of GPU time, or about $6.25 per month on an L40s.
In total it's probably something like $15.75 per month. That's a fair bit on paper, but I have certain arrangements that make it significantly less cheap for me. I could re-architect Arsène to not have to be online 24/7, but it's frankly not worth it when the big cost is the GPU time and weights volume. I don't know of a way to make that better without sacrificing model quality more than I have to.
For a shitpost though, I think it'd totally worth it to pay that much. It's kinda hilarious and I feel like it makes for a decent display of how bad things could get if we go full "AI replaces writers" like some people seem to want for some reason I can't even begin to understand.
I still think it's funny that I have to explicitly tell people to not take financial advice from it, because if I didn't then they will.
Do you think the process node advantage and SoC/HBM-first will hold up long enough for the software to catch up? High-end Metal gear looks expensive until you compare it to NVIDIA with 64Gb+ of reasonably high memory bandwidth attached to dedicated FP vector units :)
One imagines that being able to move inference workloads on and off device with a platform like `fly.io` would represent a lot of degrees of freedom for edge-heavy applications.
I do know that getting those working in a cloud provider setup is a "pain in the ass" (according to ex-AWS friends) so I don't personally have hope in seeing that happen in production.
However, the premise makes me laugh so much, so who knows? :)
I'm curious to know how Fly figured their own GPU support with Firecracker. In the past they had some very detailed technical posts on how they achieved certain things, so I'm hoping we'll see one on their GPU support in the future!
[1]: https://github.com/firecracker-microvm/firecracker/issues/11...
It looks pretty sweet. Rust & sharing libraries with Firecracker and ChromeOS's crosvm, with more emphasis on long-running stateful services than in Firecracker.
https://github.com/cloud-hypervisor/cloud-hypervisor/issues/...
I would love an example on how much time a request charges. Obviously it will vary, but is it 2 seconds or “minimum 60 seconds per spin up”?
You can also opt to let our proxy stop machines for you, but the most granular option is to just do it in code.
So yes, kind of. You just wait before you exit.
Edit: Found my answer elsewhere (yes).
God that's a horrible idea. Blog time!
$100/mo is about 10 days of speeches a month, how much data do you have?
edit - if the pricing seems reasonable, you can just limit how many minutes you send. AssemblyAI is another provider at about the same cost.
1. Provisioned machines should not die randomly and not spin back up.
Saving you time and money.
AGPL does not mean you have to share everything you've built atop a service, just everything you've linked to it and any changes you've made to it. If you're accessing an S3-like service using only an HTTPS API, that isn't going to make your code subject to the AGPL.
I’ve sold Enterprise Saas to Google and we had to attest we have no AGPL code servicing them. This is for a CRM-like app.
I am not so sure about that. Otherwise, you could trivially get around the AGPL by using https services to launder your proprietary changes.
There is not enough caselaw to say how a case that used only http services provided by AGPL to run a proprietary service would turn out, and it is not worth betting your business on it.
This is a very interesting proposition that makes me reconsider my opinion of AGPL.
(Regardless of how you see the AGPL)
I've never seen someone put this into words, but it makes a lot of sense. I mean, idealistically computers are deterministic, whereas the law is not (by design), yet there exists many parallels between the two. For instance, the lawbook has strong parallels to the documentation for software. So it makes sense why programmers might assume the law is also mostly deterministic, even if this is false
That it's possible to interpret the AGPL both ways (that the prior hack is legal, and that it is not), and that the project author could very well believe either one, suggests to me that the AGPL's terms aren't rigidly binding, but ultimately a kind of "don't do what the author thinks the license says, whatever that is".
Correct, this is a known caveat, that's also covered a bit more in the GNU article about the AGPL when discussing Software as a Service Substitutes, ref: https://www.gnu.org/licenses/why-affero-gpl.html.en
- https://benhoyt.com/writings/flyio-and-tigris/ (discussed here: https://news.ycombinator.com/item?id=39360870)
We run plenty of our own models and hardware, so I get wanting to have control over the metal. I'm just trying to figure out who this is targeted at.
- having the GPU compute in the same data center or at least from the same cloud provider can be a huge plus
- it's not that rare for various providers we have tried out to run out of available A100 GPUs, even with large providers we had issues like that multiple times (less an issue if you aren't locked to specific regions)
- not all providers provide a usable scale down to zero "on demand" model, idk. how well it works with fly long term but that could be another point
- race-to-zero startups have the tendency to not last, it's kind by design from a 100 of them just a very few survive
- if you are already on fly and write a non-public tech demo which just gets evaluated a few times their GPU offering can act like a default don't think much about it solution (through you using e.g. Huggingface services would be often more likely)
- A lot of companies can't run their own hardware for various reasons, at best they can rent a rack in another Datacenter but for small use use-cases this isn't always worth it. Similar there are use cases which do might A100s but only run them rarely (e.g. on weekly analytics data). Potentially less then 1h/w in which case race-to-zero pricing might not look interesting at all
To sum up I think there are many small reasons why some companies, not just startups, might have interest in fly GPUs, especially if they are already on fly. But there is no single "that's why" argument, especially if you are already deploying to another cloud.
But none of this answers my question.
I'm trying to understand the intersection of things like "people who need GPU compute" and "people who need to scale down to zero".
This can't be a very big market.
Given fly is deployed in equinox data centers just like everyone else, fundamentally there isn’t much difference between #2 and #3.
You can even get H100s for cheaper than these prices at $2.24 an hour.
So these do seem a bit expensive, but this might be because there is high demand for them from customers and they don't have the supply.
The whole reason Coreweave is on a fat growth trajectory right now is they used their VC money to buy a ton of GPUs at the right time
Wow, that’s some spectacularly false advertising.
(I'm not sure which of them it was, we are currently evaluating multiple providers and I'm not really involved in that process.)
Custom models? Apart from the Big 3 (in no particular order):
...
Who are the big 3 in this context?
For comparison, the first time experience on Amazon EC2 is much worse. I had tried to get a GPU instance on EC2 but couldn't reserve it (cryptic error message). Then I realized as a first-time EC2 user my default quota simply doesn't allow any GPU instances. After contacting support and waiting 4-5 days I eventually got a response my quota was increased, but I still can't launch a GPU instance... apparently my quota is still zero. At this point I gave up and found vast.ai. I don't know if Amazon realizes how FRUSTRATING their useless default quotas are for first-time EC2 users.
some use cases do so.
but if not there are much cheaper consumer GPU based choices
but then maybe you anyway just use it for 1-2 hours in total in which case the price difference might just not matter
gpu-friendly base images tend to be larger (1-3g+) so that takes time (30s - 2m range) to create a new Machine (vm).
Then there’s “spin up time” of your software - downloading model files adds as long as it takes to download GB of model files.
Models (and pip dependencies!) can generally be “cached” if you (re)use volumes.
Attaching volumes to gpu machines dynamically created via the API takes a bit of management on your end (in that you’d need to keep track of your volumes, what region they’re in, and what to do if you need more volumes than you have)
But at least in theory for deployments you should generate deployment images.
I.e. no pip included in the image(!), all dependencies preloaded, unnecessary parts stripped, etc.
Models likely might also be bundled, but not always.
Still large images, but also depending on what they are for the same image might be reused often so it can be cached by the provider to some degree.
[1] https://news.ycombinator.com/item?id=39353663
Maybe cos it's replicate they might be hesitant to adopt it but it does seem to make things a lot smoother Even with lambalabs' lambdastack I still hit cuda hell https://github.com/replicate/cog
And if it's the best of the big 3 providers, then it can't that bad, right ..... right? /s
*AWS is expensive, always, except if magic*
Where magic means very clever optimizations (often deeply affecting your project architecture/code design) which require the right amount of knowledge/insights into a very confusing UI/UX and enough time evaluate all aspects. I.e. it might simple not be viable for startups and is expensive in it's own way.
Through most cheaper alternatives have their own huge bag of issues.
Most important fly.io is their own cloud provider not just a more easy way to use AWS. I mean while I don't know if they have their own server centers in every region they do have their own servers.
This is the title of one of the sections. Why? Think IT sector needs to stop using such titles.