Migrating from AWS to Fly.io
terrateam.io
terrateam.io
I’d be happy if they can wrap Neon’s serverless postgres service into their networks somehow.
We are working on becoming first party service on several major problems now. We will be happy to provide this for Fly. Kurt and I have been chatting.
Fly wants cross region replication which is coming soon. Once it's there we can integrate.
With that Neon is still on AWS and can be close to Fly, but not run on Fly servers. It's relatively straightforward for us to run on Fly, but the S3 part will still be Amazon.
For those out-of-loop: https://news.ycombinator.com/item?id=7382151
The idea of having to do my own upgrades by updating the docker container scares me.
I have literally never run a stateful service off docker and don’t see the point, thanks to AWS.
RDS has been extremely stable and boring. Maybe it’s out together by similar duct tape under the hood but it’s been incredibly solid for nearly a decade across multiple database stacks.
If Fly builds it, it will create another Goliath who tries to eat everything.
Servers sitting less than 1km to each other tend to have lesser latency between them, compared to servers that could be anywhere on the public internet.
But if you have a zero-trust architecture and everything communicates over wireguard, then technically public or private won't matter right?
Not saying public/private IPs don't matter – since almost all the servers would still only have private IPs.
In a sense this is more like VPC peering.
I remember a failure of one transatlantic line that was used for one direction of packets we were sending between EU and US. I spent a night on a phone trying to convince AWS and our other datacenter operator to work around the problem by changing their routing and to push the backbone operator to fix the problem - we didn't have any relationship with the backbone of course.
You don't want such disruption to happen in a core part of your app. These disruptions can also be intermittent and "random" and you have no way to fix.
I think I know what you're getting at, and if I'm right, this isn't really a subjective or squishy subject.
The only reason I think we might not be is the selling point is less vendor lock-in, not something managed significantly better. For the most part, the major cloud providers have mature, feature rich product offerings for common use cases like DB, distributed queues, running services, load balancers, etc.
A lack of integrated security and authorization? Should be solvable though.
The expense of external traffic? The latency to a DB in a different datacenter?
With Fly.io (at least today), this isn't as straightforward given your favorite database is likely not deployed there.
Higher database round-trip times are fixed by using database stored procedures instead of making multitudes of tiny compound SQL queries.
The best option I could find was using Digital Ocean Managed Databases. The cheapest costs $15, but you can host multiple databases with good insights and backups. You can choose where you host it, and place it close to the region where youre Fly.io apps are, with low latency.
Only caveat is there isn't an easy way to automatically update the IPs whitelist of the database, to support the Fly builder and deployed app(s).
As a software consultant focusing on designing, building, and deploying software on AWS, more often than not, the infrastructure cost issues I've seen have less to do with the underlying infrastructure provider — AWS in this case — and more to do with the application itself that's being deployed. Most recently, we were able to shave off 50% of a client's bill that was using Lambda for compute. The issue? Several, including sleep timers (on a service where you pay by the millisecond!) as well as pathological code (i.e. consume message from queue, enqueue same message, rinse and repeat).
Yes. You can save money by transitioning to another provider. But I'd start with reviewing the underlying code and architecture first
Fly makes it really easy to do the 80% of things that most small/medium operations need to just get done now.
Give fly.io a couple years and it will end up forced to take on all the complexity and edge cases AWS and Azure have to deal with today.
A shiny new codebase is always nicer than an old one, because the old one has had to accommodate all those pesky customer needs.
Their SLA's a bit too iffy for my use case atm but I wish them well in this noble endeavor.
Wrote up some stuff back then: https://f5n.org/blog/2022/trying-out-some-hosting-options/ but unfortunately none of my pain points/caveats seem to have been fixed
* deploying (re)builds your container but doesn't even tag it locally, so if you just built one, it will build the exact same one again and on the other hand you are left guessing which hash is the container that the tool just built.
* if you start with your own Dockerfile the mandatory options in fly.toml are not described overly well, so a bit of fiddling
* recently I redeployed and without any real change (just a new code build) it wouldn't deploy, I had to remove some of the settings in fly.toml that I had guessed at to make it work in the summer to make it work now - no changelog anywhere
* if your app is coming close to some assumed limit on the services.concurrency (which is not documented well), again you need to guess and redeploy until it works
Maybe I've run into an edge case with a JVM app where 256MB of RAM is tight, but my problem is more that all the feedback I get from the tooling is a little meh and doesn't make me confident to ever try production workloads here. Not complaining about the free service though and the dashboard/metrics stuff seems nice.You do read about interesting hacks where someone will set up rds in a region that may be single digit milliseconds away from a fly region. then, presumably, you could put PG bouncer on a sort of bastion host that connects to the fly wire guard VPN. But obviously, there’s no guarantee that the latency will always be that good.
That guarantee already doesn't exist with RDS. The RDS master is in one AZ, your server may be in another (and if it isn't, well, RDS will fail over to another AZ eventually). The network latency between AWS AZs is usually good, but it can be arbitrarily high, up to full outages between AZs.
This is just a measure of degrees. Your application and/or business already has to handle some degree of network issues (aws network outages), so it's a tradeoff of whether the extra hops increase the chance enough to matter.
It isn't "Managed Postgres," but the differences are minimal. RDS is ultimately a solution for people who look in the mirror and confidently say "you don't know how to run a database."
I think this is a terrible oversimplification and something tells me that you haven't had to deal with a complex database setup from an operations perspective. RDS reduces a huge overhead in terms of operations (ha, backups, upgrades and clustering being the first ones that come to my mind). Between RDS and running a database on a virtual machine(s) and manage it with, let's say, Ansible for providing the four aforementioned features I would chose RDS any day of the week.
(Disclaimer: I'm a TAM for AWS.)
Count me as someone who confidently doesn't know how to run a database and doesn't really care to. At least not at a level where someone would hire me to do it in production.
Hah, I wish you didn’t need to know how to run a database to use RDS. So much is dependent on settings/parameter groups.
I've often found it to be the opposite actually.
My experience is with RDS MySQL, and on that, RDS heavily restricts what you can do. Want to do partial replication? Nope. Want to install a database plugin? No access provided to do so.
Used to have a MySQL instance on EC2, but the rest of the team joined the cargo-cult of 'everyone else uses RDS, so it must be good'. I used to be able to grep replication logs to find problem queries, but RDS doesn't give you access to those. I used to use various I/O and CPU monitoring tools to help pinpoint bottlenecks in queries/performance, but you only get the few metrics RDS gives you (e.g. RDS only gives you aggregate CPU usage, not per-core usage).
Even stuff like killing queries gets annoying - standard MySQL GUIs typically issue a `KILL` statement, but you aren't given permission to execute that. RDS provides a workaround via a stored procedure, but that means you have to break into a console and remember the name of the SP.
Which leads to my next point - I think "managed" is a big misnomer. RDS is nothing like a truly managed DB with a DBA. AWS isn't assigning someone to optimise your tables, or won't help you look at your queries to see what can be done better. If something goes wrong, it's on you to fix it. IMO, RDS is more like a pre-configured database. It saves you from having to initially configure the database, and saves you from having to set up automated backups, HA etc.
My opinion is that, if all you need is a cookie-cutter solution, RDS is okay. If you need a complex setup, stay away.
In the same way that a RDBMS is ultimately a solution for people who look in the mirror and confidently say “you don’t know how to directly write to disc while guaranteeing the validity of relational data in spite of concurrent writes, power failures, etc.”
Use the tools that make sense, but don't be afraid to pick the right tool.
Which is hopefully virtually everyone whose full-time job isn't DBA.
But, realistically, we'd like a managed Postgres provider on Fly.io hardware. It's a much better developer experience, we need DBs in every region, and our private networking is pretty dang powerful. I think we're close, but we may need to get a little bigger before we seem relevant to them. We're weirdly closer to managed MySQL than Postgres.
https://github.com/fly-apps/postgres-ha
It has some direct `flyctl` integration (which is also open source), but it's not doing anything you can't do yourself if you want.
Pre migration, the Terrateam AWS monthly bill was about $200/mo. With Fly.io, we're paying around $90/mo.
From your post you say you needed a simple service plus a Postgres instance. I think you can have that for around $50 on AWS using a small EC2 and Aurora Serverless v2 (min $40/mo)
The initial motivation to migrate was cost, but at the end of the day we're just happy with the Fly.io developer experience. It's a better fit for us.
2.Free tier oracle sometimes feels like you have a lower priority on the given resources, and network speed is less than on hetzner, but works fine now for a year. Wouldn't trust commercial things running on it though without a plan b
Oh and FreeBSD runs on it (i think it's under partner-images)
I also really like the dashboard much better than AWS or GCP.
Only, of course, it really, really can't, and their current offer is very much contingent on that not happening.
There is a reason why they need to have this offer to attract a few flies. Personally I would stay as far away as I can from Oracle's products.
High licensing fees: Oracle is known for its high licensing fees, which can be very expensive for businesses. This can lead to a financial burden on companies, especially smaller businesses.
Poor customer support: Oracle has a reputation for poor customer support, with many users complaining about long wait times, unhelpful responses, and a lack of follow-through.
Complex software: Oracle's software can be difficult to use and understand, which can lead to delays and frustration for businesses.
Compatibility issues: Oracle's software is not always compatible with other systems and can cause problems with integration.
Poor security: Oracle has had a number of security breaches in the past, which can lead to concerns about data privacy and security.
Overall, Oracle's high costs, poor customer support, complex software, compatibility issues, and security concerns make it a risky choice for businesses.
Moreover, Oracle has a history of suing companies that it believes are using its software without proper licensing or permission. This includes sending out lawyer letters to companies that it believes are infringing on its intellectual property rights. Oracle has been criticized for its aggressive tactics, which some believe are designed to intimidate and bully companies into compliance.
I wrote this answer with the help of ChatGPT. I think the AI did learn Oracle’s bad side pretty well.
I will just add a link to an archive of the famous blog post that was thankfully deleted but that show that the reputation is not completely inaccurate: https://web.archive.org/web/20150811052336/https://blogs.ora...
For anything serious you would want to run a dedicated vCPU which is several times more expensive (though still quite affordable)
$200 -> $90 a month is peanuts. But we're 100% bootstrapped and frugal.
You mentioned there was a case where stolon didn’t failover. Have you created an issue on this for GitHub?
(Have been a user and contributor of stolon; hence curious to know)
> if you configure your application to expose a Prometheus endpoint, those metrics will automatically show up on your Grafana dashboard
But that's such an amazing and simple idea to integrate observability. Love the approach.
Nothing special vs vanilla prom/grafana, but seamless “no-ops” integration vs DIY.
Fly.io went with "you know grafana and prometheus? yeah, we'll do just that". And I think that's perfect.
I set up a machine - a few weeks ago this involved a few direct api calls using curl - maybe `flyctl` can do it all now.
Then I deployed on that machine a version of my app that gracefully shut down if no requests were received for 60s.
Works great. Shuts down after a minute, boots back up in under a second when a request comes in.
See the docs for details
I daresay they'll be wonderful once the team get the machine stuff working under flyctl fully.
Their generous free tier does make it easy to play and assess though.
I was really excited since fly seems dead simple, but I haven't gotten it running yet...
Though, it's a "partner" provider and it's in your official github account, plus there's no notices saying it's not officially supported... it would probably help people if that were a bit clearer.
It also probably needs an update to use api.machines.dev. It is currently trying to connect over wireguard to hit our private API, because we didn't have a public API endpoint when it was built. This is probably brittle. :)
Can I just sub `https://api.machines.dev` in for the api hostname everywhere?
EDIT: See this comment for an example on how to make the provider use the public endpoint https://github.com/fly-apps/terraform-provider-fly/issues/42...
EDIT EDIT: I pasted the wrong link, sorry about that.
Looking at the code, it uses GraphQL? I can't find any documentation on your GraphQL API either though, aside from a web editor that doesn't appear to be linked anywhere official.
Pretty tiny cloud costs either way, probably safe to assume they’re a small, early stage startup. But even still, it was likely a bad call to spend significant engineering effort on saving $110/month, which is a rounding error even for small startups. Hard to imagine that time couldn’t have been better spent on things that grow the business, like new features and bug fixes.
However, the blog post author responded, noting that this migration took only 1 weekend, and they like other aspects of Fly than just the cost. So fair enough, 1 weekend is a small investment, less than I expected, I would have guessed this took 2+ weeks.
All of my apps are low-volume hobbyist web apps that run on a single machine, so I don't do any Terraform, k8s, or Postgres, so my use case is a little simpler than OP.
My biggest complaint has been with outages,[0, 1, 2] but that's gotten better. The original architecture meant that larger servers could evict other users' running servers, which they admitted was a bad idea.[2] I'm not sure if they've fixed that since, but I haven't seen it happen in a few months.
>The container logging solution provided by Fly.io is basic. It's easy to view logs with the Fly.io CLI and via the Fly.io dashboard. However, there's only a small window of logs that are kept, forcing you to create a remote logging solution either internal to your application or via a separate Fly.io application that ships logs to an external service. This is a piece of operational overhead that I'd like to see removed.
I very much agree. This is a big feature gap I feel when using Fly as well.
>Fly.io is not a support-first company. If this bothers you, then Fly.io is probably not for you.
>When emailing support, and email is the only option, it can take from hours to days to get a response. Sometimes they don't follow up. It does not feel like their level of support is on par with other cloud providers.
This hasn't been my experience. I go through the forum rather than email, but I typically see responses within hours, even on the weekends.
By contrast, when I reported bugs to GCP through their designated support channels, there was multi-day latency, and they'd just keep dismissing my report.[3]
[0] https://community.fly.io/t/cant-deploy-my-app-in-iad/7746
[1] https://community.fly.io/t/application-vms-down-without-any-...
[2] https://community.fly.io/t/app-stuck-in-pending-state-after-...
Outages is not something specific to Fly. For example, if you have an app on Azure and it has only one instance then... well, you asked for it to go down from time to time.
Maybe not a big deal for hobby projects, but for business critical stuff one should always consider deploying several instances of the same app when possible.
I'd rather have a host that can do around three nines and just accept that as my upper bound.
Once that simple requirement is met, the scaling becomes easy.
My apps are self-contained in a single Docker container. Just Go, SQLite, and Litestream:
Biggest negative so far: no easy way to reset the database [1]. I'd typically like to blow it away and re-create it frequently as I tinker with it e.g. adding/removing/renaming a column or changing a data type. Those small changes now require migrations, which are a fair bit more work (but not a huge problem tbh).
Biggest positive: ability to select a zone meant response time for Australian users was very fast compared to heroku; making a big difference to UX.
I think that's the real niceness here, a CLI that is pretty straightforward (not perfect by any means! but I can still wrap my head around it)
Wondering whether they would be a viable alternative to Postgres in the typical "Database + Docker Container" architecture.
As for LiteFS, it's still early beta stage so I don't know of many people using it in a production capacity. I would expect it to be another six months or year before confidence builds enough to see more regular usage.
Is Fly.io really than much cheaper for hosting CDN/Postgres/Docker Container?
I see closed connections on low traffic redis connections - for now a dumb catch-and-reconnect is doing ok
https://community.fly.io/t/redis-socketclosedunexpectedlyerr...
Taking this thread's advice, I started self-hosting on flyio and have only seen 1 timeout in ~2 months.
People who proxy through cloudflare hit this a lot, unfortunately.
I suspect what may have caused the issue is I have multiple fly hosts on subdomains, but each subdomain had a wildcard certificate against the root (*.example.com).
Here are a few support issues with this problem:
https://community.fly.io/t/ssl-certificate-did-not-renew-aut...
https://community.fly.io/t/certificate-expired/9143
https://community.fly.io/t/ssl-cert-expired-and-did-not-rene...
I tested Fly, DO app platform, Render. Ended up with DO (but it's meh, can't mount disk volume, can't deploy to a specific sha code git)
I'm going to [stress]test their wss:// again this year, I'm thinking to finally move a commercial one to fly (already have hobby apps there)
They are the only provider that has nailed the killer feature: turn off shit while it isn't in use.
My friend and I have different ISPs and yet we’re both routed to the east coast instead of the much closer Los Angeles edge server.
So our requests end up going to the east coast(edge), then back to the west coast (app), back to the east coast(edge), and then back to west coast (origin).
The latency adds up and makes requests to a nearby app almost as slow as to one hosted across an ocean.
I’m worried that other users will hit this and make hosting on Fly terrible for apps that require lower latency :/
Routing issues like this are typically the result of some weird network provider politics. The fix is typically ti change how we announce IPs to force a route to do something we want. It's a dark art.
Traces and ISP involved are in there. Thanks for looking into it!
Edit: I just double checked and it's still the same route to the east coast as posted in the thread
1) Record the IP address and edge of connecting clients 2) Use an IP API to determine the rough location and ISP of clients based on their IP 3) Spatially query the edge closest to each client and graph/log serious mismatches (>x miles off) 4) Cast the spells required to route offending ISP better 5) See if the data improves :)
For example, the IAD edge would record that I connected from San Diego via Cox Communications. My closest edge geographically is LAX so that's ~2500 miles of excess travel!
I also documented steps how to run Rails application if someone needs it: https://businessclasskit.com/docs/how-to-deploy-rails-sideki...
> When emailing support, and email is the only option, it can take from hours to days to get a response. Sometimes they don't follow up. It does not feel like their level of support is on par with other cloud providers.
This enough to keep me on AWS. You can say a lot about AWS but their support is great
You can't even ssh until the container started and they have some requirements on hostnames (0.0.0.0)
At least a few of their examples are broken.
The documentation could be better but it's pretty good
Is it better than Aws? Yes of course, you need to spend way more to get a dx as bad as Aws.
Setting up a VPS is a smoother experience than fly, though.
Once it's running, it's lovely and reminds me of heroku.
In fairness, I'm struggling to think of a way they could get you to ssh to a container that hasn't started yet
> and they have some requirements on hostnames (0.0.0.0)
Can you clarify?
A reminder: we don't run "containers". We take containers, unpack them, and transform them into VMs. There's no host OS for you to access.
With that said, I succesfully deployed a simple Rust/Actix/Sqlite app with a Dockerfile to Fly.io. I then thought I'd try out litestream. For reasons I'm still not sure about, having `ENTRYPOINT ["litestream replicate … -exec myapp"]` resulted in immediate kernel panics.[1]
As I was debugging, my instinct was to want to `fly ssh console` into a running container to see if running the same command from the shell produced any clues for further debugging. Though I understand this is the wrong mental model, the thought was something like "well, I know everything else works, I just need Fly to ignore this failing command so I can have a minute to poke around." To do this, I ended up just removing litestream from the ENTRYPOINT so the deploy would succeed, then I could SSH in and play around at the shell to see what was going on.
Again, I have no idea whether this is the same sort of problem the other person was having, but for my case, what would be helpful is probably not changing how `fly ssh console` behaves, but perhaps some documentation of suggested debugging techniques in case your app is failing to start.
[1] Putting the exact same command into a "start.sh" and making that the ENTRYPOINT worked fine so that's what I ended up doing.
We can clean this up, so that you get a clearer, simpler error ("your entrypoint exited, here's the exit code, there's nothing else for this VM to do so it's exiting, have a nice day"), and it's been on the docket for months. We'll get it done!
We could conceivably add a flag for our `init` to hang around waiting for you to SSH in after your entrypoint exits. But that's clunky and complicated. Usually, you want your kernel to exit when your entrypoint fails, so that your service restarts! What you should do instead is push a container that has enough process supervision to hang around itself. Here's a doc:
But when I'm trying to get a service running for the first time and I'm not sure that I have the right command in the entrypoint, the right arguments to that command, or the right supporting files in place, or the right libraries installed, or the right file permissions, …, well, then I don't want things to just blindly restart, I want a handle and some information so I can figure out why it isn't working.
ETA: I recognize that your link to docs about running a supervisor addresses this problem. For me this raises some interesting questions. Like, I understand why Ben would implement `litestream exec` but maybe it would be better to steer users to a proper supervisor? Separately, what if it's the supervisor that's failing? Now I'm back to seeing kernel panics and not having error messages or a shell.
This is a common thing in CI platforms, and the way they usually expose this is "run tests with SSH enabled", and they keep it open for 30 minutes/2 hours/whatever until a session closes.
So if I have some app failing, being able to run `fly restart --ssh-debug`, having that first just sit around waiting for the app to boot, and then dropping into ssh would be a very helpful piece of UX. The main thing is cleanup, but y'all charge for compute! You can be pretty loosy-goosy on that one honestly.
Also a way to reuse the definition but run a different command. For example for a rails app, web worker (puma) and job worker (sidekiq) differs only in the command run. In fly, I would have to duplicate those. If there is an easier way, do let us know.
The first, we're not completely sure about yet. Probably someday, though.
Just like Heroku.
Thank you btw for all your funding/support of the Erlang/Elixir ecosystem.