Fly.io outage – resolved
status.flyio.net
status.flyio.net
Application up time is flakey, but what was worse were fly deploys failing for no clear reason. Sometimes layers would just hang and eventually fail for no particular reason; I'd run the same command an hour or two later without any changes and it would just work as expected.
I'd love to make a monitoring service to deploy a basic app (i.e. run the fly deploy command) every 5 minutes and see how often those deploys fail or hang. I'd guess ~5% inexplicably fail, which is frustrating unless you've got a lot of spare time.
Their support still leaves a lot to be desired even as someone that pays for it but the ease of running and deploying a distributed front end keeps bringing me back.
Always good to monitor your dependencies if you have the time. Then when someone complains about an issue in your service, you can check your monitoring to see if your upstream services are broken. If they are, at least you know where to start debugging.
Looks like it lasted 16 minutes for them.
The entire API was also unusable, not just deployments.
The last post-mortem they wrote is very interesting and full of details. Basically back in 2016 the heart or keystone component of fly.io production infrastructure was called consul, which is a highly secure TLS server that tracks shared state and it requires that both the server certificate and the client certificate be authenticated. Since it was centralized, it had scaling issues, so fly.io wrote a replacement for it in 2020 called corrosion, and quickly forgot about consul, but didn't have the heart to kill it. Then in October 2024 consul's root key signing key expires, which brought down all connectivity, and since it uses bidirectional authentication, they couldn't bring it back online until they deployed new SSL certificates to every machine in their fleet. Somehow they did this in half an hour, but the chain of dominoes had already been set in motion to reveal other weaknesses in their infrastructure that they could eliminate. There was this other internal service whose own independent set of TLS keys had also expired long ago, but they didn't notice until they tried rebooting it as part of the consul rekey, since doing so severed the TCP connections it had established way back when its certificate was valid. Plus the whole time this is happening, their logging tools are DDOSing their network provider. It took some real heroes to save the company and all their customers too when that many things explode at once.
On their careers page [1], the Fly team goes, "We're not big believers in tech debt."
As an outsider, reads like a cacophony of contradictions?
[1] https://fly.io/docs/hiring/working/#we-re-ruthless-about-doi...
If you actually do live up to yours, then you need to adopt better principles.
(personally speaking, I'm humble enough because I can hardly build a toy side-project right!)
We're just people, working on building a thing.
https://www.youtube.com/watch?v=ghNJxYP5Ses
Also: stop calling yourself an "outsider". You follow us as closely as anybody. :)
https://news.ycombinator.com/item?id=41917436
https://news.ycombinator.com/item?id=35044516
https://news.ycombinator.com/item?id=34742946
https://news.ycombinator.com/item?id=34229751
If a cloud platform doesn't really provide reliability, I'd say it's probably not worth it. You could better just rent a (virtual) server and save the cloud tax.
You get edge computing, autoscaling, and load balancing without additional configuration.
Not as flexible as AWS, but also much easier to setup and maintain.
But the reliability issues suck now and then.
Additionally, having machines that turn off when not in use is easy to configure, which I never managed on AWS.
I haven't looked at it recently, but App Runner could do a few of Fly.io esque things (but slightly more expensive): https://aws.amazon.com/apprunner/
Today, Fly.io is more or less in the same market as Lightsail, not AWS. And when you compare it to Lightsail, it blows it away.
This is a bit of a confusing sentence because there are so many pronouns. Do all of the "it"s refer to Fly.io?
For $5 you get:
Latest gen CPUs and RAM
HTTPS
DDoS protection
Cloudflare CDN
Autoscale
Competent support
I'd say the best part is the predictable monthly prices
And while most people probably don't care, they are an established public company, so there is more chance they will exist in 10 years
also, my experience with support was not the same as yours. they were utterly useless for the most part.
for a personal web dev (or similar) project, like, i agree, they’ve got good value.
but having worked in a small biz where DO was what they built everything on — no. bad idea. spend more. use aws (graviton ec2 instances)/azure.
but GO and pocketbase is on record for supporting 10k concurrent requests per second on low powered VPS
5G towers are a ton of compute on the edge to secure and protect the traffic passing through them.
Or if by edge you mean having stuff close to your consumers, every non trivial operation does that.
And no not every nontrivial operation does it to the extreme of an envisioned fly.io deployment.
There is a lot of things we do for our users that we don't need (no one "needs" SPA etc). But if it is easy to make your app faster for your users, why not?
In a world where much web browsing starts with ACK SYN ACK, it is nice if the server is close to you.
For dynamic data I use SWR.
I could use Cloudflare workers but it doesn’t play so nice with Astro.
I also have a “form submission service” where I receive a Post and send an email.
I need maximum uptime to avoid revenue loss.
It’s a go service so I deploy ~6 machines across the US to ensure I don’t drop any requests.
I haven’t had downtime in years.
*Note this is for an instance with only 256MB RAM (https://fly.io/docs/about/pricing/), but it's definitely possible to run non-trivial projects on that. Rust-based web servers like Rocket require only about 10MB RAM. Basic PHP servers should also fit from what I can find.
(Not saying the typical cheap VPS on LowEndTalk has comparable PaaS features. Only responding to parent’s use case of a single cheap instance.)
You can easily get 4 GB of RAM for $5 from the likes of Hetzner or Hostinger, so that's 16x more RAM for 2.5x the price. One relatively unknown provider I have used in the past offers 2 GB of RAM for €3.6/month (if paid monthly, €3 if anually), so 8x more RAM for 1.5-2x the price. I'm sure I could find something even cheaper, but I'm just looking at providers I have personally used.
BTW that dropdown seems to be sorted cheapest > most expensive. If you go to the bottom of the list the price for that same VPS doubles.
There's definitely places that offer it... also 512m
I know because I've personally bought such plans and that was $5-10/yr because I didn't need dedicated ipv4.
One of my VMs had an uptime of more than 1050 days before the infrastructure rebooted it, so in terms of availability they've certainly surprised me.
The only downside I've come across with Oracle Free is that the 'best' regions are typically full. I ended up provisioning my free VMs in another region/country and it works fine.
I suppose another downside (if you want to view it this way) is they will delete idle unused free VMs after a certain time period. You have to add a credit card to your account to "upgrade" your account and run free resource indefinitely. While you're not charged for anything, it makes me nervous forking over a CC number to Oracle.
Fly is mostly (to my knowledge) reselling Netactuate and OVH servers, their main innovation is the developer experience on top, using Docker on a MicroVM based approach. Of course not only that, but I think it’s their main differentiator.
Haven’t used that in a while but Scaleway offered ridiculously cheap dedicated ARM hardware close to these price points, not sure if they still do.
Just get a $5/mo VPS instead if you're really concerned about a few dollars a month.
The perverse irony is that the most common reason cited by cloud providers for not letting people set a hard cap on charges is an insistence that surely the last thing you want in the world is for your service to be taken offline, even if it does means avoiding a $1k–$100k bill at the end of the month.
if you are going to haggle over $2/month then you are better off just connecting your raspberry pi with wireguard/cloudflare tunnel on a residential connection
I had to leave a few months ago after the price raises and how many times my boss saw some issue in the project I had with them.
They also deprecated and removed their sqlite backup service. Back to GCP and not worrying about so many outages now.
But in all seriousness the gall to raise prices before actually fixing the reliability problems is pretty shocking. I understand it's a bit of a chicken-and-egg thing where you maybe are tight on resources but there's no scenario where it's acceptable to have a product with these kinds of problems and then raise prices on existing customers who are putting up with it.
expect to see more of these "post-mortem apologies" from fly.io in the future because it won't be the last
in fact, you can almost get the same thing fly.io does by running firecracker on your own bare metal servers and cheaper too.
I'm afraid the public sentiment towards fly.io has been tainted for good (I can't count how many times they apologized now).
Gotcha. I'll be sure to pass on the good word.
If it helps: all sorts of things can and do go wrong, but the most likely form of disruption you're likely to see here are periods of times when deployments don't work. This outage was a deployments/orchestration outage. We had a total request routing outage several months back, owing to a Rust concurrency landmine we stepped on, but those are very rare.
(Deployment and state-update outages are a big deal, and if you deploy to diverse groups of Fly Machines constantly, as we encourage you to do, that being one of the big features of the platform, they can impact your availability. I'm not downplaying them.)
For accurate updates, follow https://community.fly.io/t/fly-io-site-is-currently-inaccess...
I have had my Railway app online till date without any major downtimes too. I recommend anyone looking for a decent replacement to try them.
We ack'd this and then pretty heavily to making it stellar, so if you're still having issues please let us know (that should not be the case)
Best, Jake from Railway
I understand that end-users want reliability (and Fly gets a bad rep despite pretty significant investment on this front in the past 2 years), but such outages aren't exclusive to one provider & not the other. Building cloud infra is no one's definition of easy.
https://railway.com/changelog/2024-09-20-railway-metal-beta#...
Fly.io seriously needs to get it together. Why it hasn’t happened yet is a mystery to me. They have a good product but stability needs to be an absolute top for a hosting service. Everything else is secondary.
With Fly, we had 3-4 downtimes in 2023 in a span of 4 months.
I guess the secret is to be the incumbent with no suitable replacement. Then you can be complete garbage in terms of reliability and everyone will just hand wave away your poor ops story
That's the core difference.
> Ok.I caught up with our oncall and This seems related to the Fly.io incident that is reported in our status page. Our login does call things in the Fly.io API
> we are already in touch with Fly and will see if we can speed this up
Apparently Turso are going to offer an AWS tier at some point.
In practice, that means if a server goes down, they have to load the last snapshot from that instance from the Backup and push it on a new server, update the network path, and pray to god that not more server fail than spare capacity is available. Otherwise you have to wait for a restore until the datacenter mounted a few more boxes in the rack.
That explains quite a bit the randomness of those outage reports i.e. my app is down vs the other is fine and mine came back in 5 minutes vs the other took forever.
As a business on a budget, I think anything else i.e. a small civo cluster serves you better.
> a fly instance is hardwired to one physical server and thus cannot fail over
I'm having trouble understanding how else this is supposed to be? I understand that live migration is a thing, but even in those cases, a VM is "hardwired" to some physical server, no?
You can run your workload (in this case a VM) on top of a scheduler, so if one node goes down the workload is just spun up on another available node.
You will have downtime, but it will be limited.
On Fly, one can absolutely set this up. Multiple ways: https://fly.io/docs/apps/app-availability / https://archive.md/SJ32K
They mean the storage part. If your VM's storage(state) is on one server and that server dies, you have to restore from backup. If your VM's storage is on remote shared storage mounted to that server and the server dies, your VM can be restarted elsewhere that has access to that shared storage.
In AWS land it's the difference between instance store (local to a server) and EBS (remote, attached locally).
There's a tradeoff in that shared storage will be slightly slower due to having to traverse networking, and it's harder to manage properly; but the reliability gain is massive.
Majority of EC2 instance types did not have live migration until very recently. Some probably still don't (they don't really spell out how and when it's supposed to work). It is also not free - there's a noticeable brown-out when your VM gets migrated on GCP for example.
Generally, you have worse performance while in the preparing to move state, an actual pause, then worse performance as the move finishes up. Depending on the networking setup, some inbound packets may be lost or delayed.
[1] https://cloud.google.com/compute/docs/instances/live-migrati...
Fly might still go down completely if their proxy layer fails but it's much less common.
- MS 365/Teams/Exchange had a blip in the morning
- Fly.io with complete outage
- then a handful of sites and services impacted due to those outages
Usually advocate against “change freezes” but I think a change freeze around major holidays makes sense. Give all teams a recharge/pause/whatever.
Don’t put too much pressure on the B-squads that were unfortunate to draw the short stick.
Sure maybe "unnecessary" changes, but the line gets very gray very fast.
you're right that there's a grey line, but crossing that line involves waking up several people and the on call person makes a judgement call. if it's not important enough to wake up several people over, then things stay frozen.
It needs to get better but it's not there yet.
And then inevitably it turns out that there's a special marketing/product push, with special pricing logic that needs new code, and new UI widgets, causing a huge traffic/load surge, and it needs to go out NOW during the freeze, and this is revenue, so it is critical to the business leaders. Most of eng, and all of infra, didn't know about it, because the product team was cramming until the last minute, and it was kinda secret. So it turns out you can freeze the high-quality little fixes, but you can't really freeze the flaky brand-new features ...
It's just a struggle, and I still advise to forget the freeze, and try to be reasonable and not rush things (before, during, or after the freeze).
https://wa.aws.amazon.com/wellarchitected/2020-07-02T19-33-2... / https://archive.md/uaJlR
Urgent business change needs to go through? Sure, be prepared to defend to a vp/exec why it needs to go in now.
Urgent security fix? Yep same vp will approve it.
It's a no-brainer to stop your typical changes which aren't needed for a couple of weeks. By the way, it doesn't mean your whole pipeline needs to stop. You can still have stuff ready to go to prod or pre prod after the freeze
Sure you can try and reduce those as well during the holiday season, but what if a certificate has to be renewed? What if a critical security patch needs to be applied? What if a set of servers need to be reprovisioned? What if a hard disk is running out of space?
You cannot plan your way out of operational challenges, regardless of what time of year it is.
For example if it's a small feature then it probably makes sense to wait and keep things stable. But, if it's something that itself causes larger imminent danger like security patches / hard disk space constraints, then it's worth taking on the risk of change to mitigate the risk of not doing it.
At the end of the day no system is perfect and it ends up being judgement calls but I think viewing it as a risk tradeoff is helpful to understand.
Reading this, I see two routine operational issues, one security issue and one hardware issue.
You can’t plan you way around security issues or hardware failures, but operational issues you both can and should plan around. Holiday schedules like this are fixed points in time, so there’s absolutely no reason why you can’t plan all routine works to be completed either a week in advance, or a week after, the holiday period.
Certificates don’t need to be near the point of expiry to be renewed. Capacity doesn’t need to be at critical levels to be expanded. Ultimately, this is a risk management question (as a sibling has also commented). Is the organisation willing to take on increased risk in exchange for deferring operational expenses?
If the operational expense is inevitable (the certificate will need renewing), that seems like an easy answer when it comes to risk management over holidays.
If the operational expense is not inevitable (will we really need to expand capacity?), it then becomes a game of probabilities and financials - likelihood of expense being incurred, amount of expense incurred if done ahead of time, impact to business if something goes wrong during a holiday.
Im not super familiar with their constraints but scylladb can do eventual consistency and is generally quite flexible. CouchDB is also an option for multi-leader replication.
Once in a while we'd have a real outage that matched the test we ran as recently as the weekend before.
I was helping a bank switch over to the DR site(s) one day during such a real outage and I left my mic open when someone asked me what the commotion was on the upper floors of our HQ. I said "super happy fun surprise disaster recovery test for company X".
VP of BIG bank was on the line monitoring and laughed "I'm using that one on the executive call in 15, thanks!" Supposedly it got picked up at the bank internally after the VP made the joke and was an unofficial code for such an outage for a long time.
Traders make loads, not the SWEs
It seems like every vendor sales team I work with is an "executive" or "director of sales" even though in reality they're just regular old salespeople.
I don’t envy the difficulty of doing this, but I’m quite confident they’ll iron the bugs out.
If I were in the market for cloud services I’d highly prize a long-term relationship on mutual benefit and fair dealings over a short-term nuisance of being an early adopter.
I strongly suspect your investment in fly is going to pay off.
b7r6@b7r6.net
(Also: not a legend, just loud.)
When you talk technology, I listen, and I doubt I'm alone in that. Keep up the good work with fly.io!
Examples include basically any PaaS, IaaS, or any company that provides a mission-critical service to another company (B2B SaaS).
If you run a basic B2C CRUD app, maybe it’s not a big deal if you service goes down for 5 minutes. Unfortunately there are quite a few categories of companies where downtime simply isn’t tolerated by customers. (I operate a company with a “zero downtime” expectation from customers - it’s no joke, and I would never use any infrastructure abstraction layer other than AWS, GCP or Azure - preferably AWS us-east-1 because, well, if you know the joke…)
But the truth is everybody can afford some level of outage, simply because nobody has the budget to provision an infra that can never fail.
I refuse to believe that this category still exists, when I need to keep my county's alternate number for 911 in my address book, because CenturyLink had a 6 hour outage in 2014 and a two day outage in 2018. If the phone company can't manage to keep 911 running anymore, I'd be very surprised what does have zero downtime over a ten year period.
Personally, nine nines is too hard, so I shoot for eight eights.
Most B2B SaaS solutions have very long sales cycles and a high total cost to implement, so there is a lot of inertia to switching that “a few annoying hours of downtime a year” isn’t going to cover. Also, the metric that will drive churn isn’t actually zero downtime, it’s “nearest competitor’s downtime,” which is usually a very different number.
Meaning empirically, downtime seems to be tolerated by their customers up to some point?
> It's still 99.99+% SLA
But this is simply not accurate. 99.99% uptime is < 52m 9.8s annually of downtime. They apparently blew well through that today. Looks like they essentially had the equivalent of 4 years of 99.99% uptime equivalent this evening.
Four nines is so unforgiving that it's almost the case that if people are required to be in the loop at any point during an incident, you will blow the fourth nine for the whole year in a single incident.
Again, I know it's hard. I would not want to be in the space. That fourth nine is really difficult to earn.
In the meanwhile, <hugops> to the Fly team as they work to resolve this (and hopefully get some rest).
Other circles use "SLO" (where the O stands for objective).
(Anyone know what the details in fly.io SLA are?)
Technically, anyone could offer five- or six-nines and just depend on most customers not to claim the credits :-D
Actually hitting/exceeding four nines is still tough.
Earlier in the year they had a catastrophic outage in LHR, we lost all our data. Yes this is also on me, I'm aware. Still, that's a hard nope from me, we migrated.
Part of that playbook is the old Move Fast & Break Things. That can still be the right call for young projects, but it has two big problems:
1) AWS successfully moved themselves into the position of "safe" hosting choice, so it's much rarer for engineers to have influence on something that's seen by money men as a humdrum, solved problem;
2) engineers are not the internal influencers they used to be, being laid off left and right the last few years, and without time for hobby projects.
(maybe also 3) it's much harder to build a useful free tier on a hosting service, which used to be a necessary marketing expense to reach those engineers).
So idk, I feel like the bar is just higher for hosting stability than it used to be, and novelty is a much harder sell, even here. Or rather: if you're going to brag about reinventing so many wheels, they need to not to come off the cart as often.
Providers will fail. good contingencies won't.
...hears faint sound...I SAID GOOD, QUIET YOU!
About 6 months ago we migrated our most critical stuff from Fly to CF and boy every time Fly has issues I'm so glad we did.
We spent months trying to convince them of problems with their H2 implementation in their LB/proxy (they insisted nginx was at fault, spoiler - it wasn't) but had to leave (we also went to CF, which has its own problems). Eventually one of their employees wrong a long blog post about H2 that made it obvious they finally found and fixed those problems but months too late for my employer at the time.
It would have been infinitely better for us if they could have just fixed their stability problems, that abstraction suited us as did their LB/proxy impl and SNI pricing.
I wish them well, some really smart folk over there but I can imagine these reliability problems are probably really grinding down morale.
Everything is going to be 200 OK!
...with not a single status update from Microsoft in sight.
(There's not a single price on there, why even create the page?)
There's also a link to the pricing calculator https://fly.io/calculator
Anyways we've been dunking on ourselves for not having a proper pricing page longer than anyone else could have. :)