A couple of times per mo, redis (and occasionally postgres) times out.
My TLS cert did not renew, resulting in 8 hours of downtime.
You get what you pay for.
A couple of times per mo, redis (and occasionally postgres) times out.
My TLS cert did not renew, resulting in 8 hours of downtime.
You get what you pay for.
- It took several hours to get a 1 GB disk provisioned (their API would time out and I'd get left with an allocated but unusable disk that couldn't be deleted without deleting the entire app it was attached to.) The Fly forums have numerous reports of the same issue, going back months, across multiple data centers. All resolved with the moral equivalent of "oops, try again later?"
- The fly.toml reference is incomplete and internally inconsistent, documenting keys which are not valid, while omitting valid keys that appear elsewhere in the docs.
- Volume snapshots are retained for either 5 or 7 days, depending on which of the multiple conflicting docs pages you believe.
- I wanted to let them know about that last one, so I emailed the address listed at https://fly.io/docs/about/support/, only to get an autoresponse saying that they don't read that mailbox and that I should pound sand unless I'm already paying them to read my emails.
Reading between the lines of their recent blog posts, it seems like a lot of the infrastructure pains are pinned on Nomad, with the hope that moving to Fly Machines / Apps V2 will obviate a lot of those issues.
I hope they're right, because what they're trying to deliver is exactly what I want.
Volume provisioning issues have been a game of whackamole, largely as the result of scaling pains. We've grown like 4x in the last 3 months (thanks, Heroku) and it's put a strain on pretty much everything.
So many problems between me and even giving fly a serious try. Unfortunate, but maybe in a year it will be better.
The SSH and Builder issues are interesting. Those sound like something related to our wireguard stack.
With no way to evaluate the service, I opted to delete my account.
It's since been corrected, and I still have a handful of services running with them, but getting started with the platform definitely was a rougher experience than I was expecting it to be.
Thinking this is just because of a domain name is silly.
Also Fly.io -- You may want to clearly accept and manage trial issues (or at least a subset) via a support path. It seems silly to effectively bounce users with the impression that you have no support because they are not YET paying customers while in the trial phase.
Probably that was it, but my point is, there is no process you (as a customer) can follow to remove the fraud flag.
Also the flag was still there long after the preauth.
That's just unacceptible imo.
The downtime you see is our proxy taking time to cutover- however, we have https://docs.railway.app/deploy/healthchecks that will only cutover once we have a 200 from your API. This way we can keep your old deploy live and serving requests if and only if it's live.
Theres more we can do to make it magical, but this should help in the meantime.
This could be done automatically by pinging the root and searching for a header set by railway's default page. If it doesn't exist then the service is live.
We are hiring Network Engineers for exactly this reason :')
What I think is happening is that the UI shows the status of the underlying container and there is some fault with the reverse proxy they are using to expose the container to the internet.
Fly.io doesn't have this issue but I am leery of all the complaints here.
This is only true for the free tier, and (hopefully) won't be true for much longer.
3 separate issues prevented me deploying. One was my fault, but hard to diagnose because error messages contain no useful information.
The other two were the platform falling over, with the support forum having no resolution but plenty of other people in the same boat.
Lots to love about fly. The latency is noticeably better than anything I’ve used except cloudflare. Prices are good. Hopefully they’ll get the growing pains sorted and we’ll have another contender.
If you end up trying again, give the machines based apps a shot. It's much simpler infrastructure: https://community.fly.io/t/fly-apps-on-machine-prerelease/10...
Redis/Postgres timeouts sound like they might be a problem we can help with, though. Both those services work through a load balancer with a TCP idle timeout. If you get timeouts after a period of inactivity, it's usually a simple config tweak to have a driver handle that seamlessly.
The timeout issues surprised me, more drivers than we expected can't handle a DB connection timing out when it's idle.
My server response time seemed to vary from <100ms to >300ms for the same operation.
I don't know whether this is a platform issue with Fly or something else, but it was kind of annoying to have such a high variation with seemingly no way to fix it.
it was very frustrating to see un-answered problems reported, fly.io down for what seemed to be many users, the same time seeing a fly.io co-founder arguing about politics all day long here on HN...
I think they are too engineering culture. Target is devs, sure, but they could use some product guys.