Fly.io Postgres cluster down for 3 days, no word from them about it
webcache.googleusercontent.com
webcache.googleusercontent.com
> Hi Folks,
> Just wanted to provide some more details on what happened here, both with the thread and the host issue.
> The radio silence in this thread wasn’t intentional, and I’m sorry if it seemed that way. While we check the forum regularly, sometimes topics get missed. Unfortunately this thread one slipped by us until today, when someone saw it and flagged it internally. If we’d seen it earlier, we’d have offered more details the.
> More on what happened: We had a single host in the syd region go down, hard, with multiple issues. In short, the host required a restart, then refused to come back online cleanly. Once back online, it refused to connect with our service discovery system. Ultimately it required a significant amount of manual work to recover.
> Apps running multiple instances would have seen the instance on this host go unreachable, but other instances would have remained up and new instances could be added. Single instance apps on this host were unreachable for the duration of the outage. We strongly recommend running multiple instances to mitigate the impact of single-host failures like this.
> The main status page (status.fly.io) is used for global and regional outages. For single host issues like this one we post alerts on the status tab in the dashboard (the emergency maintenance message @south-paw posted). This was an abnormally long single-host failure and we’re reassessing how these longer-lasting single-host outages are communicated.
> It sucks to feel ignored when you’re having issues, even when it’s not intentional. Sorry we didn’t catch this thread sooner.
[1] https://community.fly.io/t/service-interruption-cant-destroy...
If it really got missed, then I don't understand how the thread was made private to only logged-in users?
About a year ago fly awarded a few people in the forums, I think it was 3, the “aeronaut” badge. Basically just pointless bling for a “routinely very helpful” person or somesuch. Still, I can imagine it was cool to get it. No, it wasn’t me.
One person I saw with it absolutely deserved it: this person is, to this day, always hopping in and helping people; linking to docs; raising their own issues with a big dose of “fellow builder” understanding and empathy; that sort of person. My own queries typically led me to a thread that this person has answered. In short - the kind of helpful, proactive, high knowledge volunteer early adopter that every community needs - and a handful are blessed to find.
Then one day I saw this same person had offered — to one random newbie with build problems in one of the many HALP threads — a reply like, “maybe Fly isn’t the best option for you. here are some other places that can host an app”.
The thread was left alone and faded, like many when a lost newbie is involved. But 1 day later, I noticed this tireless early adopter no longer had their “aeronaut” badge.
I still refuse to believe my own eyes about something that petty.
Also, here’s the long forgotten badge, still with 3 people… https://community.fly.io/badges/107/aeronaut
idk man, there's these awfully convenient disappearing forum threads too. The benefit of the doubt is starting to expire.
I see you're a co-founder, so presumably you have some sway on priorities and skin in the game. I think you should take the reputational damage you're accruing here much more seriously than you apparently are. A few more incidents like this and it won't just be you telling people you're a bad option.
* edited to tone down the forum thread disappearance angle. FWIW I do believe that it likely wasn't deliberate. My main point was that these things add up and "of course we wouldn't do that!" starts to ring a little hollow the 10th time you hear it...
FWIW, I do believe them when they say this wasn't intentional. Considering how the Internet operates, they would be incredibly stupid to do something like that on purpose.
That being said, the way the entire affair was handled certainly leaves a lot to be desired.
I was really just trying to point out that this kind of good faith benefit-of-the-doubt has a limit, and fear of reaching that limit should be keeping people at fly up at night a lot more than it apparently is. I don't know how many colossal public fuckups a company can endure before its reputation is permanently ruined, but it's definitely not infinite.
Michael - Don't take the bait.
As someone who has zero affiliation with Fly.IO other than a few PR's to their OSS(I don't even know Michael), I greatly appreciate the contributions they have given back to the community.
There are a lot of great hosting companies. Fly.IO stands out due to their revolutionary architecture and contributions back to the OSS community. I wish more companies operated like this.
It's understandable some are upset about an outage. But Fly is doing really interesting and game-changing things, not copying a traditional vmware, cpanel or k8s route.
Just as a reminder to what this company has offered back to everyone.
SQLite: Ben Johnson's OSS work around SQLite stands out. Fly.IO and his work have really made sqlite a contender. - https://fly.io/blog/all-in-on-sqlite-litestream/ - https://fly.io/blog/introducing-litefs/ - https://github.com/superfly/litefs - https://github.com/benbjohnson/litestream - https://fly.io/blog/sqlite-internals-wal/ - https://fly.io/blog/wal-mode-in-litefs/
Who really considered sqlite as a production option before Fly and Ben? Not me.
Firecracker: Firecracker is amazing, but difficult to debug when something bad happens. There aren't a ton of people in devops who would share what they have. If you've ever used Firecracker, you've really been helped a lot by the various guides they have provided back to the community like these: - https://fly.io/docs/reference/architecture/ - https://fly.io/blog/fly-machines/ - https://fly.io/blog/sandboxing-and-workload-isolation/
Their architecture is beautiful and revolutionary. They're probably the first or second ones to find a lot of the new edge cases as they grow.
It's a lot harder to be the first one over the wall than it is to copy. They've literally given the average developer a blueprint to build scalable businesses that compete with their own.
Well if it's not true then that would be a silly reason to pick to not use them.
https://community.fly.io/c/questions-and-help/app-not-workin...
EDIT: it now appears that the "app-not-working" tag itself has been deleted, and no longer shows up even when logged in.
If I pay for something I want the person I pay money to help me fix problems I get.
I’m back on digitalocean now. I’m not unhappy about it, they’re very solid. I don’t love some things about their services, but overall I’d highly recommend them to other developers.
I gave up on fly because I’d spontaneously be unable to automate deployments due to limited resources. Or I’d have previously happy deployments go missing with no automatic recovery. I didn’t realize this was happening to a number of my services until I started monitoring with 3rd party tools, and it became evident that I really couldn’t rely on them.
It’s a shame because I do like a lot of other things about them. Even for hobby work it didn’t seem worth the trouble. With digitalocean, everything “just works”. There’s no free tier, but the lower end of pricing means I can run several Go apps off of the same droplet for less than the price of a latte. It’s worth the sanity.
collect some evidence, maybe someone wants to do something about it.
This is the one thing I need in every app and don't want to do myself.
The main app is node/express running on Digital Ocean and it connects to directly to the Supabase hosted Postgres for most operations, but then uses the Supabase auth API for auth related stuff.
Saves a lot of time sending password reset emails etc and the entire project costs less than $5/mo in hosting costs.
I don't like self hosting anything that requires its own process. And if I did decide to self host I would choose a more mature project.
This is a very young one man project delegating the heavy lifting to another one man project. And it doesn't appear to support social logins.
Part of what inspired me to give fly.io a shot was that I didn’t love the monorepo deployment story on the app platform. Fly doesn’t have a solution to that, but I suppose I felt less tied to DO at the time because I wasn’t totally content anyways. I’ve discovered since then that I was actually doing it wrong, so I’m way happier. I’m pretty big on monorepos so their whole system fits my workflow remarkably well now.
I’d like to figure out how to prevent deployments when my code doesn’t change in one app, but does in another. At the moment, pushing anything at all will trigger all apps to rebuild and deploy again. Not a huge deal and several orders of magnitude less painful than not being able to deploy at all, haha.
I know that’s probably pretty easy for many, but I was pretty new to k8s and it felt like magic.
Things like they don't give you the postgres root user on their managed postgres. And I ran into issues trying to capture the deployments in code. Their terraform providers are pretty good, but still leave something to be desired. For all its many warts, I'm much happier back on AWS. It did end up more expensive, but it's worth it for the fine grained control in my case.
But I spent the last 5 years as a DevOps/SRE, so... uh... I'm picky.
I’m nowhere near as picky as you are, but maybe I’ll need to be at some point. As it is I mostly just build stuff and send it to the internet. If it builds and it does what I expected, I’m pretty happy! I don’t often need anything too special.
Higher number of outages at Vultr over 5 years, but none longer than a few hours. I can’t remember the last DO outage lasting more than a few minutes.
Experienced a Vultr routing problem that lasted several hours; they communicated about it, but it was still a long time to fix.
DO once did an auto-migration of a server to another cluster with an attendant outage that lasted a few minutes at most. No IP changes, completely transparent.
You boot one up in less than 30 seconds, and get ssh access to it almost immediately. It's very BS-free.
I think fly’s tooling feels better than doctl, but the infrastructure is incomparable at the end of the day. doctl has improved over time too, and with added pressure from newcomers I don’t doubt that it’ll continue to improve.
Additionally, you still keep the full ssh access to the machine if you ever need it.
Contact is in my profile but I'd love to have some more people kick the tires and tell me what they want built the most.
They don't run on AWS. Not sure what sort of rumors are running :(
> data centers from?
The major players e.g. Equinix, Coresite, etc. Varies per location. Even AWS don't build most of their data centers.
I feel like there might be more to it, especially considering the situation with electricity prices in some places in EU recently.
I used (and still use) a Lithuanian platform called Time4VPS which was cheaper than Hetzner previously, yet had to increase their prices somewhat for that reason. Now only some of their plans are competitive with Hetzner, while Hetzner also provides some managed services as well.
Hetzner docs also went into some of the details regarding the pricing: https://docs.hetzner.com/robot/general/pricing/hetzner-prici...
And yet, I can't help but to wonder why they don't give in to the desire to maximize profit margins, like happened to say Scaleway (good platform, but as expensive as DigitalOcean).
Maybe that's not actually the dominant cost, or they've optimized everything else so well they can just eat the electric bill.
like the longtime workhorse was a high performance skylake desktop cpu w/o ecc ram
We used to use nianet to house our hardware in Denmark. Basically these companies does hardware renting and they also do hardware renting with more steps which is where you rent rack space but own the hardware. They provide the place for the hardware and they also have multiple locations so that you have both backup and redundancy, and while it doesn’t scale globally in 20 years I’ve literally never worked on anything that needed to beyond having some buffer caches for clients logging in on their vacations or something like that.
What Hetzner seems to be doing with the DO styled hosting, and this is just a guess, is that they are one or the many EU companies preparing for the big EU exodus from the non-EU cloud. Which is frankly a solid bet these days where both AWS and Azure are increasing prices and are becoming more and more unusable because of EU legislation. Part of this is privacy which Microsoft and Amazon are great with in terms of compliance, but part of it is also national security. I work in an investment bank that builds solar plants, since finance and energy are both critical sectors we risk being told that half of the finance/energy companies in the world can’t use Microsoft because the EU seems it as a single point of failure if our entire energy sector relies on Azure. Which is sort of reasonable right? But what this means for us is that we can’t vendor lock-in, not really, because we need to have up-to-date exit strategies for how we plan on being fully operation a month after leaving Azure. Which is easy when you just containerise everything and run it in VMs or similar, and really annoying if you go full in on things like AKS. Which doesn’t help our Azure costs.
Anyway, right now we are planning on leaving Azure because of cost. Not today, not next week but sometime in the next 5-10 years and a lot of these EU cloud alternatives that actually operate the hardware instead of renting it are likely going to be a very realistic alternative. And that is the private sector, I spend time in the EU public sector which is a massive amount of money and I’m guessing it’ll leave both AWS and Azure by 2050. Some of these EU cloud initiatives is going to explode when that happens, and right now, hetzner is one of the best bets.
To get back to your question, DO rents server space. I have no idea where they’d rent it in Germany but they could potentially be renting it from Hetzner.
You do the same on the other side of the table. Companies like Hetzner knows that EU cloud sollutions are likely to see growth, so it's only natural that they invest in the tech to put themselves in a prime position to jump on the opportunity. Selling a good product while you do so is the way I would do it personally, but you also have EU cloud initiatives backed by VC money going straight for the endgame.
To add on to the comments about Hetzner building their own custom hardware, they also custom built their own software stack. They rejected the hype that was OpenStack and worked diligently on their own hypervisor platform (that they are incredibly secretive about) and that appears to be paying off in spades for them. Most sovereign cloud plays end up being suffocated by the complexity, and incoherence, of the OpenStack ecosystem. It just becomes impossible to ship.
For a fascinatingly different take on how to build a datacenter: https://www.youtube.com/watch?v=5eo8nz_niiM
* Edit: remove speculation about Kubernetes and Hetzner, that was based on hazy memory.
I am asking for this since a while and was told there is no way Hetzner would offer such a service. Certain Posts on Social Media have also never been answered with any kind of indication that they are actually working on it.
Please provide some Details on this.
So huge grain of salt, you are totally right. It could be internal platform work only.
1. Strict rules and strict customer verification. Crypto mining that wastes SSDs is not allowed. Portscans, mass emails, etc. are not allowed. They also don't offer GPUs to the general public because it has been abused in the past. You usually need to send in ID documents just to open an account. My guess is this allows them to avoid most bad actors and, thereby, waste less money on fraud.
2. Extremely long-term investments. They typically build their own hardware and then use it over 10 years. They have their own flea market where you can rent older server models for a steep discount. That means they will have a long time where the hardware is fully paid off and still generating revenue.
3. Great service. With a mid-sized company, I can call their technicians in the middle of the night. The fact that we could call them in case of a crisis has generated A LOT of good will. But I would be truly surprised if they didn't make a profit off those phone calls, as they charge roughly 4x the salary cost.
4. High-margin managed services. In addition to just the cheap servers, they also offer a managed service where they will do OS and security upgrades for you. It's roughly 2x the price of the server and it appears to be almost fully automated. I know some freelance web designers who will insist on using Hetzner Managed for deployment for their clients, because it is just so convenient. You effectively pass off all recurring maintenance for €300 a month and your client is happy to have an emergency phone number (see #3) in case the box goes down.
Does anyone know if that's still the case?
As far as I know, they do not require any ID when canceling the service.
[1] https://registry.terraform.io/providers/hetznercloud/hcloud/... [2] https://community.hetzner.com/tutorials/howto-hcloud-terrafo...
EDIT: For what it's worth, I have had good experiences with app servers hosted on Hetzner Cloud and managed Postgres provided by ElephantSQL (https://www.elephantsql.com/) for Germany-based apps.
I'd love some feedback, specifically:
- Which services do you most want to use/have managed
- What databases do you find yourself using the most
- Concerning caches, do you use memcached or mostly Redis?
[0]: https://nimbusws.com
I want to like Fly, but the reliability is one of those were I feel like every time I investigate moving workloads over I'm disappointed by these stories over and over again.
It's actually the only region to serve the entire AU and NZ population with any reasonable latency. (Ok, Singapore can do in a pinch for at least sub 200ms.)
You'd wanna hope its more than one machine!
You shouldn’t need to multi region a postgres yourself - they should have at least 2 data centre redundancy for the region and it just works.
Hope they get some magic sauce to become better at this.
When I saw them describe their multiregion SQL replication architecture I thought "what crazy person thought this wouldn't eventually open up a spider's nest of distributed systems errors?"
Their license would require a company like fly.io to pay them though, so I'm sure this resulted in fly.io instead trying to whip up an improvised infrastructure on the back of stock Postgres. I bet this cost them a whole lot more than paying CockroachDB would have, but devs have been conditioned that you should never ever pay for software even if it's the result of tons of deep engineering and solves massive brutal problems for you. I also bet there's some not-invented-here ego involved.
P.S. I don't work for CDB but I would absolutely consider them and we may end up using them at some point. They let you do a ton for free. They only charge for stuff you need if you get really really huge or if you are running a SaaS reselling DB services like fly.io would have been doing.
you must be new here ;-)
Smaller companies also do a lot of PR damage control and constantly monitor HN for threads complaining about their services.
You're not wrong but that's how it works.
alt+backspace will wipe that substring in most shells in one go.
I would hope that after a couple of hours downtime, they'd bring up a fresh machine with Ansible or whatever. Hardware or AWS/GCP Vm.
It is not just about a fresh machine which hopefully sits in each datacenter. I can imagine they needed the clone of the system due to the design of the fly.io service and that's where the "fun" begins.
"Oops! That page doesn’t exist or is private."
Edit: Ok, I can see after sign up / log in.
Make it impossible not to do so, and make it frictionless then.
Even if customers are only running one instance, I would expect the whole thing to rebalance in an automated way especially with fly.io being so container centric.
It also sounds like this is some managed Postgres service rather than users running only one instance of their container, so it’s even more reasonable to expect resilience to host failure?
I believe the underlying reason that precludes failing over to a different host machine, is that fly volumes are slices of host-attached nvme drives. If the host goes down, these can't be migrated. I _think_ instances without attached volumes will fail-over to a different host.
Of course, that's not ideal, and maybe their CLI should also warn about this loudly when creating the cluster.
And +1 to the sibling comment; Fly makes it very clear that single instance postgres isn't HA, and talks about what you need to do architecturally to maintain uptime.
edit: Remove triple as not certain about level of redundancy
>Amazon EBS volumes are designed to be highly available, reliable, and durable. At no additional charge to you, Amazon EBS volume data is replicated across multiple servers in an Availability Zone to prevent the loss of data from the failure of any single component. For more details, see the Amazon EBS Service Level Agreement.
https://aws.amazon.com/ebs/features/#Amazon_EBS_availability...
If a read replica fails, I'd expect no downtime (possibly a few errors as connections get cut off abruptly). Although there's always the risk that the remaining instances aren't able to handle the additional load.
If the master fails, you'll get a ~2min downtime
Down time is one thing. Data loss is something else.
Depends on where in your development cycle you are. If you just got started and haven't even figured out what you're actually building (prototyping), you shouldn't really use a hosting provider that randomly lose instances.
If you're on the other hand have done everything to improve your applications performance, had to resolve through-output issues with a distributed architecture and now running 10+ instances, then losing one host shouldn't impact you too much. But you really shouldn't start this way, it's doing web services the hard way and introduces a lot of complexity you shouldn't want to deal with when you're still trying to find product market fit.
I won’t be using it for side projects where I’m okay with paying $5-10/mo but don’t want to have three day outages.
From a technical perspective, could they have "been better" from a technical perspective? I see their name a lot on HN so I know they are doing really cool + advanced things and this is probably some super small edge case that slipped through the cracks.
Could they have added some message / do we as the HN community feel they needed to be like "we're gonna add some extra logging/monitoring going forward so it won't happen again"?
By all means, they probably don't owe anybody in terms of stability + uptime guarantees when it comes to a free tier. Sh*t happens.
The relevance of paid/free is that free (and cheap paid) plans don’t get fly support over email
If I was just searching online or trying to find out what various communities think about Fly.io and see several threads about major outages with poor communications, do you think I will use their services? It would be an immediate pass.
It takes a long time to build a reputation, and you can lose it instantly.
I have an ongoing issue with one of my PG clusters where one of the nodes was failing and all my attempts at fixing it are failing (mainly cloning one of the other machines to bring the cluster numbers back to normal).
I emailed my account’s support email mid Friday morning last week and did not hear back until this past Monday night.
Sucks, because like a lot of others in this thread I like what Fly is trying to do and am rooting for them, but IMO they should use a significant chunk of that funding they just received on hiring a ton of SREs and front line customer support.
EDIT: I should add, the past times I have emailed them the response time was good. It's just this most recent time was so egregious (3 days!) to get even that initial response that I bring it up.
One quote from thread:
> This is the second time I’ve had this kind of issue with Fly, where my service just goes down, Fly reports everything healthy, and there’s literally no information and nothing I can really do other than wait and hope it comes back up sometime
Another user:
> We had four machines (app + Postgres for staging and production) running yesterday, and three of the four (including both databases) are still down and can’t be accessed. I can replicate the issues others have mentioned here.
> This is our company’s external API app and so the issue broke all of our integrations.
> Our team ended up setting up a new project in fly to spin up an instance to keep us going which took a couple of hours (backfilling environment variables and configuration etc, not a bad test of our DR ability).
> There is no way I can find to get the data from the db machines. Thank goodness this isn’t our main production db and we were able to reverse engineer what we needed into there.
> Very keen to hear what’s happening with this and why after so many hours there’s no more info or updates.
Another user:
> As an aside, it’s kind of a kick in the teeth to see the status page for our organization reporting no incidents - the same page that lists our apps as under maintenance and inaccessible!
Another user:
> I’m feeling very lucky that none of our paid production apps or databases are affected currently (only our development environment is), but also really surprised that the issue has been ongoing for 17 hours now with no status page update, no notifications (beyond betterstack letting us know it was down) and one note on the app with not much info as to whats going on.
> It really worries me what would happen if it was one of our paid production instances that was affected - the data we’re working with can’t simply be ‘recovered’ later, it’d just get dropped until service resumed or we migrated to another region to get things running again
> Keen to know whats wrong and whats being done about it
Full thread (as at time of HN post; more has been added since): https://pastebin.com/ebmCSZkC
Someone tweeted Fly CEO: https://twitter.com/SouthPawNZ/status/1682181533673857024
[1] https://community.fly.io/t/service-interruption-cant-destroy...
Their typical response is either silence or so casual ("oh this is what happens we deploy on friday"). The product looks amazing but it's just a nice package around the most unreliable hosting service I've ever used.
You can't just keep breaking people's work every once a week, make them spend their weekend nights trying to bring back their stuff, and give these "we could have done better" answers. This is an excuse for exceptions, not patterns.
Love this
How dare they use AWS' patented approach to having a service outage.
I merely mentioned two characteristics of how they fail, that are spectacularly shit.
- it seems their staff were working on the issue before customers noticed it.
- once paid support was emailed, it took many hours for them to respond.
- it took about 20 hours for an update from them on the downed host.
- they weren't updating their users that were affected about the downed host or ways to recover.
- the status page was bullshit - just said everything was green even though they told customers in their own dashboard they had emergency maintenance going on.
I get that due to the nature of their plans and architecture, downtime like this is guaranteed and normal. But communication this poor is going to lose you customers. Be like other providers, who spam me with emails whenever a host I'm on even feels ticklish. Then at least I can go do something for my own apps immediately.
- Their free tier support depended on noticing message board activity and they didn't.
- Those experiencing outages were seeing the result of deploying in a non-HA configuration. Opinions differ as to whether they were properly aware that they were in that state.
- They had an unusually long outage for one particular server.
- Those points combined resulted in many people experiencing an unexplained prolonged outage.
- Their dashboard shows only regional and service outages, not individual servers being down. People did not realize this and so assumed it was a lie.
- Some silliness with Discourse tags caused people to think they were trying to hide the problems.
In short, bad luck, some bad procedures from a customer management POV, possibly some bad documentation resulted in a lot of smoke but not a lot of fire.
You get to a certain number of servers and the probability on any one day that some server somewhere is going to hiccup and bounce gets pretty high. That's what happened here: a single host in Sydney, one of many, had a problem.
When we have an incident with a single host, we update a notification channel for people with instances on that host. They are a tiny sliver of all our users, but of course that's cold comfort for them; they're experiencing an outage! That's what happened here: we did the single-host notification thing for users with apps on that Sydney host.
Normally, when we have a single-host incident, the host is back online pretty quickly. Minutes, maybe double-digit minutes if something gnarly happened. About once every 18 months or so, something worse than gnarly happens to a server (they're computers, we're not magic, all the bad things that happen to computers happen to us too). That's what happened here: we had an extended single-host outage, one that lasted over 12 hours.
(Specifically, if you're interested: somehow a containerd boltdb on that host got corrupted, so when the machine bounced, containerd refused to come back online. We use containerd as a cache for OCI container images backing flyd; if containerd goes down, no new machines can start on the host. It took a member of our team, also a containerd maintainer, several hours to do battlefield surgery on that boltdb to bring the host back up.)
Now, as you can see from the fact that we were at the top of HN all night, there is a difference between a 5 minute single-host incident and a 12-hour single-host outage. Our runbook for single-host problems is tuned for the former. 12-hour single-host outages are pretty rare, and we probably want to put them on the global status page (I'm choosing my words carefully because we have an infra team and infra management and I'm not on it, and I don't want to speak for them or, worse, make commitments for them, all I can say is I get where people are coming with this one).
edit: I hope this comment doesn't sound accusatory. At the end of the day I want everyone to succeed. I hope there's a silver lining to this in the post-mortem.
If you're running an app on Fly.io without local durable storage, then it's easy to fail over to another server. But durable storage on Fly.io is attached NVMe storage.
By far the most common way people use durable storage on Fly.io is with Postgres databases. If you're doing that on Fly.io, we automatically manage failover at the application layer: you run multiple instances, they configure themselves in a single-writer multi-reader cluster, and if the leader fails, a replica takes over.
We will let you run a single-instance Postgres "cluster", and people definitely do that. The downside to that configuration is, if the host you're on blows up, your availability can take a hit. That's just how the platform works.
If this sounds ludicrous, then I think I probably don't understand who Fly.io wants to be and that's okay. If I don't understand, however, you may want to take a look at your image and messaging to potentially recalibrate what kind of customers you're attracting.
AWS RDS lets you spin up a RDS instance that costs 3x less and regularly has downtime (the 'single-az' one), quite similar to this.
Anyone who's used servers before knows "A single instance" is the same as "sometimes you might have downtime".
Computers aren't magic, everyone from heroku (you must have multiple dynos to be high availability) to ec2 (multiple instances across AZs) agree on "a single machine is not redundant". I don't see how fly's messaging is out of line with that. They don't tell you anywhere "Our apps and machines are literally magic and will never fail".
Just because some customers are less fault tolerant than others, doesn't mean we shouldn't offer those options where people don't have the same requirements or are willing to work around it.
So hopefully as fly.io get's more popular, there will be some compelling managed offerings. I saw comments at one point from the neon CEO about a fly.io offering, but not sure if that went anywhere. I'm sure customers can also use crunchy, or other offerings.
What do you think?
$ fly volumes create mydata
Warning! Individual volumes are pinned to individual hosts.
You should create two or more volumes per application.
You will have downtime if you only create one.
Learn more at https://fly.io/docs/reference/volumes/
? Do you still want to use the volumes feature? (y/N)
(and yes, the warning is already even in red letters too)edit: sibling post shows there is such a message on the CLI. The only other thing I can think of is an "Are you sure you want to do this?" prompt, but in the end you can't reach everybody.
I kid, seems like you guys did what you could.
Hey, even if I can feel sympathetic for the course of unfortunate events, it's hard to not to comment:
if you're using a cache, you should invalidate it on failure!
Btw I am happy I got only small amounts of data in any of bolt databases...
How much is over 12 hours? 12 hours and 10 minutes? 13 hours? 67 days?
This was a server with local NVMe storage. The simplest thing to do would have been to just get rid of it, but we have quite a few free users with data they care about running on single node Postgres (because it's cheaper). It seemed like a better idea to recover this thing.
If they're sad that they lost their data, it's their fault for running on a single host with no backup. By actually performing an (apparently) difficult recovery, they reinforced their customers erroneous expectation that they are somehow responsible for the integrity of the data on any single host.
If you run off a single drive, and the drive dies, any resulting data loss is your fault. But not if something else dies.
This is much closer to EBS breaking. It happens sometimes, but if the data is easily accessible then it shouldn't get tossed.
…proceeds to make a bunch of non-factual statements.
To clarify, we communicated this incident to the personalized status page [1] of all affected customers within 30 minutes of this single host going down, and resolved the incident on the status page once it was resolved ~47h later. Here's the timeline (UTC):
- 2023-07-17 16:19 - host goes down
- 2023-07-17 16:49 - issue posted to personalized status page
- 2023-07-19 15:00 - host is fixed
- 2023-07-19 15:17 - issue marked resolved on status page
The bad news is that I'd be out of a job if I chose your service in this instance. 47 hours is two full days. For an entire cluster to be down for that long is just unacceptable. Rebuilding a cluster from the last-known-good backup should not take that long, unless there are PBs of data involved; dividing such large data stores into separate clusters/instances seems warranted. Solution archs should steer customers to multiple, smaller clusters (sharding) whenever possible. It is far better to have some customers impacted (or just some of your customer's customers) than have all impacted, in my not so humble opinion.
And, if the data size is smaller, you may want to trigger a full rebuild earlier in your DR workflows just as an insurance policy.
The good news is that only a single cluster was impacted. When the "big boys" go down, everything is impacted... but customers don't really care about that.
Not sure if this impacted customer had other instances that were working for them?
There was one physical server down. That's it. They even brought it back.
I've had AWS delete more instances, including all local NVMe store data, than I can count on my hands. Just in the last year.
Those instances didn't experience 47 hours downtime, they experienced infinite downtime, gone forever.
I guess by your standard I'd be fired for using AWS too.
But no, in reality, AWS deletes or migrates your instances all the time due to host hardware failure, and it's fine because if you know what you're doing, you have multiple instances across multiple AZs.
The same is true of fly. Sometimes underlying hardware fails (exactly like on AWS), and when that happens, you have to either have other copies of your app, or accept downtime.
I'll also add that the downtime is only 47 hours for you if you don't have the ability to spin up a new copy on a separate fly host or AZ in the meanwhile.
Combine that with them having tooling for setting up Postgres built on top of single node storage, and you have the downtime problems and unhappy customers as a given.
> If your instance root device is an instance store volume, the instance is terminated, and cannot be used again.
See also the aws "Dedicated Hosts" and "Mac Instances". Those also have similar termination behavior.
The majority of my instances lost are from the instance store thing.
I've never experienced AWS killing nodes forever; at least not DB instances.
> Rebuilding a cluster from the last-known-good backup should not take that long
It's not even clear if that's the right thing to do as a service provider.
Let's say you host a database on some database service, and the entire host is lost. I don't think you want the service provider to restore automatically from the last backup because it makes assumptions about what data loss you're tolerant to. If it just works from the last backup, suddenly you're potentially missing a day of transactions that you thought were there that magically disappears as opposed to knowing they disappeared from a hard break.
That's how other [useful] providers notify their customers that one of their hosts went down unexpectedly. Linode will send me 6 emails when they need to reboot something. Even Oracle sends me notices about network blips. I believe I've gotten one from AWS, but I also know sometimes their gear gets stuck in a bad state and I didn't get a notification, which was super annoying because it took forever to figure out it was AWS's faulty state.
Fly.io messed up, they didn't want to be a Heroku clone, but their marketing and their polished user experience design made it seem like they would be one anyway.
And as a reward now they have to deal with bottom of the barrel Heroku users that manage to do major damage to their brand whenever a single host goes down. Who would have predicted that corporate risk?
>> I get that due to the nature of their plans and architecture, downtime like this is guaranteed and normal. What other cloud providers have downtimes of 20 hours? There must be a lot to call this "guaranteed and normal".
Sadly, I've always felt a good amount of passive aggressiveness in many of the HN threads where fly.io is involved.
* Frequent machines not working, random outages, builds not working
* Support wasn't responsive, didn't read my questions (kept asking same questions over and over again) -- I paid for a higher tier specifically for support.
* General lack of features (can't add sidecars, hard to integrate with external monitoring solutions)
* Lack of documentation -- For happy path its good but any edge cases the documentation is really lacking.
Anyway, for hobby projects its fine and nice. I still host a lot of personal projects there. But I have to move my companies infrastructure off of it because it ended up costing us too much time/frustration, etc. I really had high hopes going into it as I had read it was a spiritual successor of sorts to Heroku which was an amazing service in its day, but I don't think its there yet.
Their free tier is very generous. You can get a lot happening and stay under their billing threshold. But, I like to get stuff done. I have a family. I code in my spare time very rarely, and I need a service that’ll let me just build my goddamn project. This was a small static site built by Node, so nothing spectacular happening.
I do wish them the best though. They have an excellent product in their tooling, and if they could stabilize their infrastructure I’d love to try them again.
(Cloudflare|GitHub|GitLab) Pages should do you nicely!
The answer to _any_ usage related forum question should be:
1. It's in the documentation <here> (maybe I just added it)
2. If you're left with any confusion, let me know and I'll update the documentation to resolve it
Am I right in thinking the platform got bought a little while ago, and it’s being run by a relatively small outfit?
I don't know who owns them but I do get the impression it's a small team. Hasn't been an issue for me so far. Their customer service has been very helpful and responsive on the rare occasions I've needed to contact them
The Fly dashboard reported everything was A-ok, but requests would time out. I had to manually dig into the fly logs to see that their proxy couldn't reach the server, and there was nothing I could do to fix it.
This went on for hours, until I made an issue on their forums. They never replied or gave any indication they read the thread, but it somehow magically got fixed not long after.
I really want them to succeed, but this utter lack of communication and helpless feeling of not being able to do anything has cured me from fly.io for now.
I have no earthly clue why this thread on our community site is unlisted.
We're looking at the admin UI for it right now, and there's like, a little lock next to do the story, but the "unlist story" option is still there for us to click. The best I can say is: I'm reasonably sure there wasn't some top-down edict to hide this thread (the site is public, anybody can sign up for an account and see the thread).
Say what you want about us, but hiding out from stuff like this isn't one of our flaws. When I find out more about what happened with this thread, I'll let you know (or Kurt will reply here and tell me I'm wrong).
I don't know enough about what happened with this Sydney server to be helpful to people who had instances running on it. When I know more about it, I'll be helpful, but I'm just learning about this stuff right now, after getting back in from a night out.
Almost immediately afterwards
It looks like... all the posts in the app-not-working category are "private"? Like it's some setting on the category itself? "Private" here means you need to have signed up for a Discourse account to see them?
Maybe it's hosted in the SYD region
For instance one of those things I've noticed is that most Discourse instances have those nag banners if you're not logged in begging you to log in – and that's one of the least objectionable things they do IMO. I discovered recently that Discourse also blacklists all but the most recent browsers (because Discourse is designed for the next ten years!) and serves up a plain text version on anything older… but not without a nag banner of its own admonishing you for not using a supported browser.
The infinite scrolling… ugh. I'm not a huge fan of XenForo, but as a successor to vBulletin it seems to be far more user friendly.
Obviously, deliberately hiding a negative story on our Discourse is a little like deleting a bad tweet; it's just going to guarantee someone captures and boosts it. We have a lot of flaws! But not knowing how the Internet works probably isn't one of them. No idea what's going on here, still trying to work it out.
Not going to try and guess why or when that tag change happened. Personally, I'm less concerned with this particular thread than with the apparent decision to systematically hide all potentially-negative threads from search engines.
FYI, this doesn't appear to be strictly accurate. The OP commented at 23:52 UTC saying that the thread had been made private, and the reply from "Sam-Fly" was not posted until 02:36 UTC.
I think we can just `noindex` the category instead of making it private?
What assurances could you give the community here that the support would be better next time?
It's difficult for me to say more about what happened here and how you might have handled it, because I don't know what happened with this SYD host, because it's 1AM and the people who worked on it are, I assume, asleep. When I know more, I'll do my best to get you a postmortem.
To be honest, that's enough for me. Sorry I didn't pick up on that.
Trust me, you did the business a favor.
However 'should' is pretty load bearing there and actual results are probably heavily dependent on management culture and the current state of office politics.
A host being down for 3 days isn’t a bug. And you can contact AWS support, even on the free plan, and get a reply. Try it yourself. The great thing about AWS and the other cloud providers? If a host has issues they email all customers with workloads on it so you don’t need to refresh or check a forum.
I understand fly is a community darling. They’re unreliable, with poor support currently. Maybe the dev experience is great and that makes up for it, but pretending like everything else is equally shitty? Not true.
FAR better wrt both response times and technical expertise than you'll get with any large public cloud provider.
I was dealing with some annoying cert + app migration stuff (migrating most of an app from AWS to Fly), and Kurt (CEO) was personally sending me haproxy configs bc I'm not smart enough to know how to configure low-level tcp stuff in haproxy. Not to put him on the spot here -- I doubt he'll have time to do that level of support going forward -- but that's my experience of the company's dedication to support and technical expertise.
If something breaks once it's an accident, if it breaks twice it's bad luck but if it breaks down three times it's broken processes. Based on the comment here things break at fly.io a lot more often than three times.
If you're reading my comments on HN as some kind of official response from the company, you've misconstrued them.
For what it’s worth, this is the reason most companies eventually restrict their employees from making statements about the company; It doesn’t matter if you thought it was clear that is was unofficial, any statement from an employee in a position of power (such as someone with access to the control panel) will be perceived as a communication from the company.
You may have intended it to be a personal remark about your job, but there are a lot of people in this thread looking for any communication they can get about the company.
When you step in to fill that void as a person who appears to have access and power within the company, you are the official communication whether you intend to be or not.
I am a paying customer of fly.io, on the Scale plan.
If you had said "thoughts are my own; I just work there" or something I think it would have been more clear.
I get stuff like this is frustrating. But I bet Fly staff are pretty frustrated too.
For everyone else reading this, we have been running https://changelog.com on Fly.io since April 2022. This is what our architecture currently looks like: https://github.com/thechangelog/changelog.com/blob/master/IN...
After 15 months & more than 100 million requests served by our Phoenix + PostgreSQL app running on Fly.io, I would be hard pressed to find a reason to complain. - Some deploys failed, and re-running the pipeline fixed it. - Early July 2023, 9k requests from Frankfurt returned 503s. Issue lasted 10 seconds. - While experimenting with machines, after many creations & deletions, one volume could not be deleted. Next day, the volume was gone.
That's about it after 15 months of running production workloads on Fly.io.
We mention about our Fly.io experience often in our Kaizen pod episodes, which we publish every ~2 months: https://changelog.com/topic/kaizen. For anyone curious, this is the episode in which we announced the migration: https://changelog.com/shipit/50. There is a detailed PR which goes with it: https://github.com/thechangelog/changelog.com/pull/407. We've been talking about our migration plan from apps v1 (Nomad) to apps v2 (flyd) recently: https://changelog.com/friends/2#transcript-138
I'm sorry to hear that many of you didn't have the best experience. I know that things will continue improving at Fly.io. My hope is that one day, all these hard times will make for great stories. This gives me hope: https://community.fly.io/t/reliability-its-not-great/11253
Keep improving.
Have to admit it's disappointing to hear about the lack of communication from them, especially when it's something the CEO specifically called out that they wanted to fix in his big reliability post to the community back in March.
https://community.fly.io/t/reliability-its-not-great/11253#s...
Are we even trying or just repeating ourselves because we don’t know what else to do?
How can the entire industry keep making the same basic errors?
“Let’s keep it simp… ohh nope we invented a Turing complete language and customer service is terri… wait do we have customer service?”
I get the world turning against SaaS lately.
Computers are so fast now, enthusiasts would be better served DIY; put a beige box in a local colo, use one of the big 3 for big business.
This is just starting to look disreputable and disrespectful to humanity itself putting such resources into one time bomb after another.
Sure, we have most of the day-to-day grunt work for our applications automated. But good operations is just more. It's more about maintaining control over your infrastructure at one hand, and making sure your customers feel informed and safe about their data and systems. This is hard and takes lots of experience to do well, as well as manpower.
And yes, that's entirely a soft skill. You end up with questions such as: Should we elevate this issue to an outage on the status page? To a degree you'd be scaring other customers. "Oh no, yellow status page. Something terrible must happen!". At the same time you're communicating to the affected customers just how serious you're taking their issues. "It's a thing on the status page after an initial misjudgement - sorry for that." We have many discussions ilke that during degradations and outages.
Scared customers seems a bit… puerile? In a Sunday school way? Are we not adults capable of rational discourse?
“Why is line not go up!!” still? Just continues to smell like busy work in deference to a politically mandated hallucination.
Agreed
> enthusiasts would be better served DIY; put a beige box in a local colo
I mean, like, can I provision a zero ops bit of compute from <mystery colo provider> for $20/month?
Edit: looked up colo providers in my city- “get started in 24 hours, pick a rack and amperage, schedule a call now.”. Yeaaah, no. This is why people use cloud providers instead.
You can pay for a business class fiber link too. It's about twice as expensive but they have guaranteed outage response times which is really what you pay for.
We've been adding a ton more hardware lately to stay ahead of capacity issues and as you would expect this means the volume of hardware-shaped failures has increased even though the overall failure probability has decreased. There's more we can do to help users avoid these issues, there's more we can do to speed up recovery, and there's more we can do to let you know when you're impacted.
All this feedback matters. We hear it even when we drop the ball communicating.
https://www.fujitsu.com/global/products/computing/servers/pr...
I used to use them a few years ago in a local data centre, and they were pretty good back then.
They don't seem to be widely known about though.
That does sound pretty useful.
So for yourselves, you rack them then run hardware qualification tests?
I'm still wondering about their hardware acceptance/qualification though, prior to it being deployed. ;)
I think there is an expectation mismatch between what Fly wants to offer and what the market wants from it. Fly wanted to innovate on offering the ability to the devs to be able run their apps from multiple data centers. But without a proper data persistence service, the ability to run apps from multiple data centers is not useful to a vast majority of people.
I think Fly is trying to solve the persistence issue with their SQLite replication, but that means the vast majority of the devs will have to change the way they develop applications to suit Fly platform.
I think Fly needs to choose between what it wants to become. A reliable and affordable Heroku replacement, which is a decent sized market or offer an opinionated way of developing apps which offer best performance to users all around the world.
But opinionated ways of doing things is a double edged sword. (Rails and Spring Boot are highly successful because of their opinionated defaults.) App Engine is an interesting case study in the app hosting domain. It was way ahead of the time and prescribed you a way of developing apps which allowed the apps to scale to very high traffic. But people didn't want to change the way they develop to adapt to it.
They have already pivoted once, no? At their current size (>100M in funding), I seriously doubt they can do it again.
I think they are scrambling hard, putting one fire out just to start another one later. That doesn't give me confidence in their technical roadmap and multiple people have Fly.io in their "check later" list for what now? 2 years?
It's really hard to recover your reputation when people perceive you as unreliable. Especially in the IT space.
It's really not a place to run persistent workloads. If you run postgres there, you need to be prepared to either hot load your data into a new instance, or restore from backups.
They had my account on some sort of shadow ban with no communication whatsoever after asking them to delete my account from their systems. I emailed them and to date never even got a response. I have moved everything over to Railway app and back to Google Cloud Run ever since.
So did you manage to delete your account then attempt to re-register using the same email address you deleted the account with?
Why would a company shadow ban you for asking an innocuous question?
If you are literally overwhelmed with crises, it becomes appealing to make problems go away in this manner. Not saying they are, but this thread is suggesting that.
I remember being blown away by Fly.io's simplicity and how easy it was to use. It was like hosting made simple, and I couldn't help but think, "This is it, this is the one!"
But, as time went on, I noticed little signs of trouble. Downtimes became more frequent, and my deployments, which were once snappy and seamless, turned into agonizingly slow affairs. It was like déjà vu from the time when Heroku's greatness started to wane.
It's disheartening to see Fly.io go down a similar path. As more people flocked to the platform, it seems like its performance began to suffer – just like what happened with Heroku. The more popular it got, the less reliable it seemed to become.
Scrolling through Hacker News, I can't help but feel a sense of disappointment. Others are expressing their frustration too, and it's like we're all reliving that moment when Heroku lost its charm and became a hassle.
I have to admit; it worries me. It's like a cautionary tale of how even the most promising platforms can fall from grace. It's the reality of the fast-paced tech world, but it's tough to accept.
So yeah, here I am, hoping against hope that Fly.io can somehow break free from this cycle and find its footing before it becomes as useless as Heroku was at its lowest point.
Do you think its related to scale? As in, once a company has enough paying customers to become profitable/investable, it has also accrued enough issues to where it starts feeling fresh and exciting like you said, and gradually becomes like the older competitor it once wanted to replace?
This is my experience at least. Once the company goes from a few pizzas to "we've booked a venue", entropy creeps in and adages like Conway's/Brook's law become increasingly evident.
Heroku was a new thing back then, so it took a while for abuse to ramp up—but every subsequent attempt at being generous should not even be considered without either a vicious and expensive anti-fraud department in place or deep pockets to compensate for the initial lack of said department by throwing enough hardware that the minority of honest users don’t notice the overhead.
My impression suggests that Fly does not score high on either of the above. Which is partly why I like them—the above seems like megacorp type bullshit, and they seem to be strictly no-megacorp-bullshit—but I wouldn’t be surprised if engineers at Fly had to spent most of their time dealing with fires or optimizing resource allocation and auto-limiting freeloading cryptominers, scammers, and other abusers rather than focusing on longer term infrastructure reliability or DX.
Durable and available storage are all they really need to draw me away from big cloud providers but this combined with their answer to S3 being "use S3 or run minio" means I'll never take them seriously.
This is a bad look folks, not sure how you can walk back days of silence and hiding threads. Just open an issue and talk to your users.
Is using Cloudflare R2 not an option?
As far as their offering, one should definitely understand that there are limitations and do their research.
[1] https://help.backblaze.com/hc/en-us/articles/360034798433-Ca...
I did have a problem with their dedicated server almost immediately after spinning it up. Noticed that NVMe is broken, and support went like:
- 16:28 -> I contacted them
- 16:36 -> Their first response
- 16:44 -> I sent them SMART data
- 16:48 -> They acknowledged that the NVMe needs replacing and asked me if I consent to that (and loosing of the data that was not already lost -> but running RAID so no problems there)
- 16:52 -> I agreed
- 17:30 -> NVMe was replaced and server booted
I don't have too much experience with hosting providers on that level, but that was freaking impressive response time from them. So a happy camper as well :D
EDIT: Formatting
My assumption based on the creator’s very online hacker news commentary is that they seem to be at least smart in tech. So what’s the lesson here for the rest of us who may want to start a business? Is this a “shots on goal” thing and we’re just seeing these failures more publicly than most so it biases the perception, or is there some je ne sais quoi missing that we could learn from? No offense intended by my post, but I would be very keen to learn whether there’s some X Factor missing from an otherwise ostensibly smart team’s repeated failure that we could learn from.
So why is the link to the thread 404ing and why does this post have to link to google webcache of it? I've grown to like fly.io and use them for my side projects now, and this just isn't sometime they would do. Going through some minor cognitive dissonance right now :/
(b) We definitely didn't make the thread private in response to HN.
(c) It should be public again.
Why would anyone want to become a new customer if all they see is jumble of green, yellow and red?
Green status pages attract business.
Ordinarily, a single-host incident takes a couple minutes to resolve, and, ordinarily, when it's resolved, everything that was running on the host pops right back up. This single-host outage wasn't ordinary. Somehow, a containerd boltdb got corrupted, and it took something like 12 hours for a member of our team (themselves a containerd maintainer) to do some kind of unholy surgery on that database to bring the machine back online.
The runbook we have for handling and communicating single-host outages wasn't tuned for this kind of extended outage. It will be now. Probably we'll just paint the global status page when a single-host outage crosses some kind of time threshold.
At work when it came up in a meeting people went around with horror stories of broken elements while the status page wasn't updated, terrible communication and an overall attitude that nothing is wrong, even when servers go down for days at a time.
I personally at the moment use digitalocean without any issues, but there's always the maintenance overhead of managing a server yourself.
Do you have to use an object store in that case? Or does it have to be separate from whatever application instance?
The peace of mind of managed is nice, all I have to think about is running the app, without having to deal with making sure db and files don't get lost
I didn't use it for much more, but my experience has been great. They deserve way more air time than they currently get.
If that carries over to their customer facing folks and how Render as a team has executed since then I'd absolutely recommend taking a look at them.
And if you need state, then spin up a little RDS with your favorite SQL flavor of choice?
The CI deploy script could even bake in little health-checks so you can do rolling deploys with zero downtime. Depending on how fancy you wanted to get with your shell scripting, you could probably even make 1 of your 3 boxes a canary without too much trouble.
I'm realizing I haven't thought about this in a long time, since nowadays I just get to use the fancy stuff at work. Kind of a fun thought experiment!
Render.com looks like [1] their "$0 + compute costs" plan would work out to:
∙ $25/mo for a single "Web Services" box of 1 CPU and 2GB RAM
∙ $20/mo for a single "PostgreSQL" box of 1 GB RAM, 1 CPU, and 16GB SSD
∙ TOTAL: $45/mo, and you're assuming they'll magically give you zero-downtime
Those are grim numbers, performance-wise, but let's use them as the standard and see what it'd cost in the scrappy AWS architecture I threw together in a few minutes: ∙ $12.10/mo for a single t4g.small box, which is actually 2 vCPU and 2GB RAM [2]
∙ 3x redundancy on that brings you up to $36.30/mo for compute
∙ $16.20/mo for an ALB [3]
∙ $11.52/mo for a single db.t4g.micro PostgreSQL box, plus $1.84/mo for the equivalent 16GB of storage [4]
∙ TOTAL: $65.86/mo for substantially more CPU, redundancy, and control, or...
∙ TOTAL: $41.66/mo for substantially more CPU and control over your infra, if you're willing to drop the redundancy
So it looks like it's pretty comparable in terms of raw dollars.I'll admit there's a little more "devops" overhead with the AWS setup. Though I think it's not as big of a deal as people make it out to be — it's basically an afternoon of Terraforming, and you'd probably spend an equal or greater amount of time digging through Render's docs to understand their bespoke platform anyway.
(Also, once you contemplate bulk pricing for the underlying commodities, it's easy to see how companies like Render make a healthy margin, even on their low-end offerings.)
Anyway, I guess I've nerd-sniped myself, so I'd better stop here. But that was a fun analysis!
[1] https://render.com/pricing#compute
[2] https://aws.amazon.com/ec2/pricing/on-demand/
[3] https://aws.amazon.com/elasticloadbalancing/pricing/
[4] https://aws.amazon.com/rds/postgresql/pricing/?pg=pr&loc=3
Also I have used Terraform to set up quite a few resources and it's only overhead in a small project.
I just wanna git push and see my changes published a minute later. I don't think Render is gonna take more than 10 mins to figure out https://render.com/docs/deploy-rails-sidekiq
I’ve been burned one too many times by ElasticBeanstalk so I bit the bullet and went with Render… and had everything plus PR deploys working in under an hour. Very happy so far.
It's all just docker.
If you are outside of Europe, Digital Ocean or Linode may work better for you.
Can anybody who uses Fly.io explain their rationale? Why do the additional integration with Fly.io, trust and install their special software on your machines and tie your project into their ecosystem?
What type of application are you running? How many users are using it?
Heroku (before its inevitable enshittification under Salesforce) was great for this use case. Sure you will outgrow it at some point, and it did get expensive, but when you just want to throw up an MVP with minimum fuss and maintenance you could do much worse.
You already know how to set up your project locally. Why not just do the same setup on any cloud VM and boom it is online?
That's more than what I would have or need locally.
In my experience, for a simple PHP web application, the smallest VMs already can handle a thousand concurrent users, which amounts to something like a million monthly users.
But if your users are distributed around the world and most requests are read requests then it can make sense to shave 100 or 200 ms off your response times.
You can always squander those gains later by running JavaScript for 5000 ms before showing anything :)
Having used various PaaS services that take this "pain" away from you, I sort of think the tradeoff isn't worth it. For $5/m DO will give you a backed up server. Add $15 for postgres that is a good deal.
I used Heroku for a project mostly because my team didn't have skill set to set this up and I wasn't going to do it. As far as I know they are still on Heroku (with a smattering of AWS services) for that same reason: just works and cheaper than doing it yourself.
Who doesn't? I couldn't imagine having to push to some cloud agent and wait a random amount of time every time I want to test something. With it local I can just save, maybe rebuild or have it auto-rebuild if necessary, and test, then repeat. On a fast machine this can be a few seconds or instantaneous.
Maybe the niche I'm missing here is very "green" developers who don't know how to do any sysadmin work or deploy things.
If this is you, learn it. It pays off huge, not just during development but in being able to have a lot more choice about where you deploy and a lot more control over your own stuff.
I’ve run Kubernetes clusters in multiple Parallels VMs locally with work loads in them to play around.
You also learn a ton about how things work which helps you debug and fix stuff when things go wrong. Even if you use managed stuff it’s always a huge plus to understand at least the basics of how it runs.
They still continue to get love from the developer community who "wants them to succeed". I'm puzzled as to why? Because of some blog posts?
AWS is more expensive than God, but I'll be damned if you can't have a throat to choke in less than 10 minutes whenever something like this happens.
But again, you get what you pay for.
"I'm in Hawaii on my honeymoon and my backup missed your call, so it escalated."
I probably wouldn't have answered the phone. Granted, that's why I don't do that job. But I have always had a real appreciation for the good TAMs ever since.
It was not a wise plan. It did, however, run. Technically.
I recall an incident at my old company where we were under DDOS, it was getting through cloudflare and saturating LBs in some complicated manner (don't recall the exact details) which made it hard for us to fix ourselves. They were on the phone with us for hours, well past midnight their time, helping us sort it out. The downtime sucked, but I was certainly impressed with their truly excellent support.
That is a hell of generous description for a person who sits in your Slack instance and responds with "I have escalated to the team internally and am waiting to hear back on confirmation if this is an issue."
Moving a Level 1 support engineer closer to the customer doesn't give them more information, it just reduces the latency to getting a non-answer.
Opened a ticket and support had it back up again within about 10 minutes, turned out to be a failed CPU fan which caused an overheat condition and made it so the system wouldn't complete the boot. They swapped the fan and it came up. It's the only failure I've had in years of dealing with them and was just impressed how quickly a physical failure event like that got handled.
Datacenters in my country usually had some rooms with tower servers 20 years ago here, well my first colo was for the tower server I brought in the large backpack:-). But density requirements, cold/hot aisles etc. prevailed and towers are generally considered inefficient for the datacenter purposes.
And then you have Hetzner datacenter that probably all people running DCs I know would ridicule, but they would not be able to respond to fan replacement at the same time. I wonder how many rack server chassis are recycled each year because the manufacturer just won't let you reuse them with new motherboard, power supply due to new shape, design, ports placement etc.
yikes
It just means one single person (at the vendor) who you can complain to, or raise an issue with.
anyway, had the same though when typing my sibling comment, felt so disgusted reading that
See also: https://news.ycombinator.com/item?id=35044516 and https://news.ycombinator.com/item?id=34229751
who wants that?
For the unique privilege of being able to build machines out of thin air, I will accept the occasional weekend page
Same goes for Digital Ocean. No buzz words. Just hosting with droplets. They simply say "here pick a linux distro, configure whatever and don't ask us much about app support". I use their Linux distros for my own apps and if want anything extra I just install it and suffer my own actions' consequences. Not theirs.
So this "closer to your users" voodoo is a little beyond me.
Sure you can do that with any cloud (or multiple) that has datacenters in a suitable spread of regions, but I suppose the point (or claimed point, selling point, if you like) is that that's more difficult or more expensive to coordinate. Fly says 'give us one container spec and tell us in which regions to run it', not 'we give you machines/VMs in which regions you want, figure it out'. It's an abstraction on top of 'battle tested proven bedrock' providers in a sense, except that I believe they run their own metal (to keep costs down, presumably).
I mean chat or e-commerce yes, the edge and all.
But for a ticketing system, invoicing solution or such, a few hundred millisecons are not that much of a big deal but compliance, regulations matter more.
But if you are wondering how AWS manages to be so good at it at such scale? Hosting infrastructure is incredibly complicated and AWS employs something like 100k people. Seemingly small AWS services employ more engineers than Fly.io.
That being said my take is that what's happening at Fly.io is a lack of leadership. There are not the right people in the right positions clearly. I've worked infra at companies from 5 people to, well Rackspace, and I'm having a hard time imagining so much time passing with.. Essentially a piece of infra MIA and impacting users.
That boring paas must have gotten a lot of growth from their competitors fucking up
Heroku is going to shit. GitHub integration was down for weeks last year
Fly.io can't get it together. Database being offline for days is just ridiculous.
Anyway, here's an archive link for future visitors: https://archive.is/7lSJA
I've had service issues on Fly that I've escalated to support in the past, and given my experience it feels highly unlikely that they tried sweep this under the rug or somesuch.
At the time we had deployed a small business workload (few 100$/mo in billings) and paid for their $29 support plan, so grain of salt there. We faced service issues and, while the service reliability did eventually push us to migrate, support was top-notch the whole way through. Support was happy to escalate as needed to try to help get a solution, with MrKurt eventually joining in and helping identify root causes. During the entire episode everyone was realistic about where issues could be (i.e., were open to the possibility of it being a Fly issue). As people from Fly have noted, they've historically been quite open about when they weren't the best choice.
Again, while service reliability has been an issue (and Fly has admitted this in the past and is working on it), I think the assumption of badfaith in this thread is pretty unprofessional. It's also a lesson in how hesitant people are to pay for support. $29 for access to a human is not a bad deal; we certainly got good value out of it.
Customers aren't supposed to show professionalism. Service providers are. I didn't see disrespectful comments here.
People here are just poiting this has happened many times and look like a pattern. If you don't fix a communication issue after multiple occurrences, you might not be ill intentioned, but at least careless.
I'd argue the expectation goes both ways. I won't link to specific comments, but I think it's pretty clear that some of them cross the line to disrespectful.
But it's very clear that over the last few months, the (already quite capable) fly team is just in over their heads and have bitten off way way more than they can chew.
I've had nothing but headaches after they auto-migrated our app to v2. My build machine had to be forcibly destroyed to even work. Then it allowed me to easily just delete an app cuz I wasn't able to deploy to it.
Then the deploys kept failing due to some VM provisioning error (it thinks I want to add another app when I just want to deploy to an existing app within the 3 machine limit) and honestly, I just don't care to troubleshoot this anymore. That was the point of using a platform like this... Any time I would've saved by using this platform has been wasted with these random errors that I don't have the time to troubleshoot.
I destroyed the app thinking "ok maybe I'll also recreate that one" because clearly the migration to v2 failed. And now that all my secrets were destroyed, when I try to attach the new app to postgres (with the existing username and existing database), it won't let me.
I genuinely wish you guys the best of luck with what has to be a tough time for your company, and will reconsider if you build something demonstrably more stable. But right now I just can't afford to drown with you with clients breathing down my neck.
I don't read their blog regularly but I always thought they had great content. But not after reading this.
The irony: "What people actually wanted to talk about, though? Databases."
...but apparently not when they are the problem behind said databases?
I actually like some of the concepts, like pods and ingress, but one thing I noticed that I didn't like, as far as I remember, was that there's not really a good way in kubernetes to make your YAML more dynamic. Apparently you're supposed to use these other things like Helm Charts, which isn't even part of kubernetes?
* First was some sort of certificate issue that cost me literally days of debugging that turned out to be their fault. * Then weirdness around their v2 deployments where I just can't grok some of the documentation.
Just use AWS. Your time is more valuable then what you're saving on the fly.io free plan.
Fly seems unreliable, but they offer a deploy region close to me. Does anyone have know of any alternatives?
Companies moving to the cloud are only increasing their operating costs
Self hosted enterprise "starts" (!) at $100/vCPU per month. So, yeah. Not exactly the hobbyist's choice.
As someone not incredibly experienced with devops, I always wonder what is best with databases? Should they be provisioned in Pulumi or do I just manually create them in RDS?
Secrets Manager seems like a bit of a pain point as does IAM which I think I just about understand until I get lost! Giving everything access to ingress and egress also seems a bit overly complex/powerful.
Probably the time to get something working is dramatically shorter than it once was with ChatGPT to help.
At least that would have indicated some competent leadership.
Fly gave me plenty of downtime. "Network connectivity issues".
I'm with them only because of laziness.
The setup was more complicated than a standard setup because of the abstractions and corner cases and lack of examples in my language.
Logging is bad,you can't easily SSH in the machine and inspect things.
Next time I'll setup a more powerful server I'll move all my apps there.
Unfortunately, such an approach is unlikely to produce the same stability you might be used to from other places.
But yeah I cannot complain too much, I pay nothing so I got the appropriate support.
For this very reason.
Out of your control
Out of your control
AWS is way too expensive and complex nowadays. I can’t stand the amount of terminology they invented just to deploy a website and database.
I'm very happy about everything. No complexity, easy to deploy and setup.
If hetzner had this, I would have picked them instead.