From Go on EC2 to Fly.io
benhoyt.com
benhoyt.com
A couple of times per mo, redis (and occasionally postgres) times out.
My TLS cert did not renew, resulting in 8 hours of downtime.
You get what you pay for.
- It took several hours to get a 1 GB disk provisioned (their API would time out and I'd get left with an allocated but unusable disk that couldn't be deleted without deleting the entire app it was attached to.) The Fly forums have numerous reports of the same issue, going back months, across multiple data centers. All resolved with the moral equivalent of "oops, try again later?"
- The fly.toml reference is incomplete and internally inconsistent, documenting keys which are not valid, while omitting valid keys that appear elsewhere in the docs.
- Volume snapshots are retained for either 5 or 7 days, depending on which of the multiple conflicting docs pages you believe.
- I wanted to let them know about that last one, so I emailed the address listed at https://fly.io/docs/about/support/, only to get an autoresponse saying that they don't read that mailbox and that I should pound sand unless I'm already paying them to read my emails.
Reading between the lines of their recent blog posts, it seems like a lot of the infrastructure pains are pinned on Nomad, with the hope that moving to Fly Machines / Apps V2 will obviate a lot of those issues.
I hope they're right, because what they're trying to deliver is exactly what I want.
Volume provisioning issues have been a game of whackamole, largely as the result of scaling pains. We've grown like 4x in the last 3 months (thanks, Heroku) and it's put a strain on pretty much everything.
So many problems between me and even giving fly a serious try. Unfortunate, but maybe in a year it will be better.
The SSH and Builder issues are interesting. Those sound like something related to our wireguard stack.
With no way to evaluate the service, I opted to delete my account.
It's since been corrected, and I still have a handful of services running with them, but getting started with the platform definitely was a rougher experience than I was expecting it to be.
Thinking this is just because of a domain name is silly.
Also Fly.io -- You may want to clearly accept and manage trial issues (or at least a subset) via a support path. It seems silly to effectively bounce users with the impression that you have no support because they are not YET paying customers while in the trial phase.
Probably that was it, but my point is, there is no process you (as a customer) can follow to remove the fraud flag.
Also the flag was still there long after the preauth.
That's just unacceptible imo.
The downtime you see is our proxy taking time to cutover- however, we have https://docs.railway.app/deploy/healthchecks that will only cutover once we have a 200 from your API. This way we can keep your old deploy live and serving requests if and only if it's live.
Theres more we can do to make it magical, but this should help in the meantime.
This could be done automatically by pinging the root and searching for a header set by railway's default page. If it doesn't exist then the service is live.
We are hiring Network Engineers for exactly this reason :')
What I think is happening is that the UI shows the status of the underlying container and there is some fault with the reverse proxy they are using to expose the container to the internet.
Fly.io doesn't have this issue but I am leery of all the complaints here.
This is only true for the free tier, and (hopefully) won't be true for much longer.
3 separate issues prevented me deploying. One was my fault, but hard to diagnose because error messages contain no useful information.
The other two were the platform falling over, with the support forum having no resolution but plenty of other people in the same boat.
Lots to love about fly. The latency is noticeably better than anything I’ve used except cloudflare. Prices are good. Hopefully they’ll get the growing pains sorted and we’ll have another contender.
If you end up trying again, give the machines based apps a shot. It's much simpler infrastructure: https://community.fly.io/t/fly-apps-on-machine-prerelease/10...
Redis/Postgres timeouts sound like they might be a problem we can help with, though. Both those services work through a load balancer with a TCP idle timeout. If you get timeouts after a period of inactivity, it's usually a simple config tweak to have a driver handle that seamlessly.
The timeout issues surprised me, more drivers than we expected can't handle a DB connection timing out when it's idle.
My server response time seemed to vary from <100ms to >300ms for the same operation.
I don't know whether this is a platform issue with Fly or something else, but it was kind of annoying to have such a high variation with seemingly no way to fix it.
it was very frustrating to see un-answered problems reported, fly.io down for what seemed to be many users, the same time seeing a fly.io co-founder arguing about politics all day long here on HN...
I think they are too engineering culture. Target is devs, sure, but they could use some product guys.
i have no doubt they'll come along if fly continues owning mindshare at this level. when its obvious on HN but not yet obvious on the marketing sites you just have to wait a couple years (with some exceptions to the rule, e.g. people often bring up the failure of light table which got significant HN pull and seems to have died out https://hn.algolia.com/?dateRange=all&page=0&prefix=true&que...)
Case in point, the last big thing we had run here was, theoretically, part of a series of recruiting challenges --- for a role we've paused hiring for. If we had any strategy, we'd have held off on it for a couple months and ran it when we were actually hiring. But Ben really liked the piece, and so did the rest of us, so up it went.
> i have no doubt they'll come along if fly continues owning mindshare at this level. when its obvious on HN but not yet obvious on the marketing sites you just have to wait a couple years
Seems like an accurate observation.
Now please elaborate as to why? And how does one shorten this?
I am unskilled at devrel and wish to benefit from your experience further. Cheers. :)
as for shortening... i'm writing https://dx.tips/ and have a couple of related posts coming up on that
I did have occasional outages (lost db connection). They seems to usually resolve themselves. I ended up shutting down the (unused) app since it was under attack.
Which is exactly why companies hire marketing, PR, DevRel, business development, salespeople, account executives. Hiring a bunch of smart engineers and hoping everything will automatically turn out great is a losing strategy, but that is the one fly.io seems to be focused on.
One thing from heroku in 2021 should have been learned here: Communicate before your clients learn about the issue. Communicate open.
I hope they improve in the next year
It's probably all in line with the ToS and everything, but I find it a bit shady that I have to investigate this, rather than having a free offering that doesn't suddenly charge me.
- Up to 3 shared-cpu-1x 256mb VMs
- 3GB persistent volume storage (total)
- 160GB outbound data transfer
If you spin up, say, 4 VMs you will pay for it.
But that is not all
My account is now "inactive", with no option to activate it and with no option to even pay and upgrade to paid tier.
The account is in some broken state where half of the dialogues produce dialogues like "Authorization failed or requested resource not found", like I had account with permissions denied to everything.
Even the "upgrade to paid tier" option gives
> You do not have access to perform this operation. Please contact your cloud account administrator for additional access.
Utter fucking disgrace
No thanks Oracle even for 24GB free VMs!
So I don't really care if one service costs x2 if the difference is between 5 and 10, but I care a lot if I could accidentally end up with a $2000 bill one day (and warnings don't count - not much good if I get a warning if I'm away from my laptop in a log cabin or something).
That's why I stay away from e.g. AWS for side projects, too many histories of a stupid mistake that costed quite a bit of money when all that is needed is some breaker if usage is exceeded beyond a point.
You may want to check if you also have a domain name purcahed through route53 that is set for automatic renewal every 3 years or whatever
If you're getting charged you should run `fly apps list` and see if you have more than two apps running.
One thing that trips people up is Postgres. Postgres is just another fly app. So if you moved two apps with DBs, you're probably running four total apps. We prompt you for database info when you do this, though, it's not done in the background.
One problem we've come across is expectations from Heroku users. We are very different (like the Postgres thing, there's no shared tenant postgres). People expect us to be the same. We haven't figured out how to expose those differences to people in a way that sticks.
Sounds like there may be room for improvement in your documentation.
Apps
-----------------------------------------
myappname shared-cpu-1x
myappname-db 1GB shared-cpu-1xSecond paragraph
As a company with a cushion, hosting a profitable app, it might be worth doing the usage-based pricing to try and save some money. But the risk is just not worth it for smaller projects
Worrying about having to configure systemd? "restart: always" in compose. Caddy updates? Caddy provides an up to date docker image. You can't get around caddy's config, but tbh, you do it once, and then you're done. It's way better than nginx. Deploying a new version of your go application is just "docker compose up -d".
Just wondering about having a "native" NGINX reverse-proxy + multiple apps ran via docker-compose on my VPN, seems the most hassle free. Currently using it for integration tests and it works great.
Did you maybe consider k8s, maybe as a learning experience?
Fly.io, 8 vCPUs, 32GB RAM (dedicated-cpu-8x): $398/mo
AWS, 8 vCPUs, 32GB RAM (t3.2xlarge, us-east-2): $240/mo
So, 65% more expensive. Quite steep.
https://docs.aws.amazon.com/AWSEC2/latest/UserGuide/burstabl...
When you consider ARM vs AMD vs Intel vs current-gen vs last-gen, AWS sells several 8-core 32-GB machines that have wildly varying performance characteristics.
It used to be that the manner of virtualization impacted performance, too. I think that might be less the case now.
To know if you're getting good value, you really need to try a few different vendors and benchmark then, either with real workloads or synthetic ones.
This is actually a benefit of AWS and other large public clouds. You have a choice due to their size.
If AWS's investment in Gravitron design works well for your workload, it's a ~30% cost savings. They're able to invest in designing and producing the chips due to their size. A company like fly.to is limited in selection and heavy R&D.
The OP was asking what was "most comparable" to Fly's offering. I assumed they meant "what would similar performance cost elsewhere".
By switching from Intel to ARM, they're buying something fundamentally different. Maybe it's a better fit overall! But I didn't understand that to be what they were asking.
I wasn't sure what the dedicated instances for fly.io were running on (can't find it on this page: https://fly.io/docs/about/pricing/#plans). To truly do an apples / apples compare there's networking capabilities, storage, and other considerations also.
> The EC2 burstable instances consist of T4g, T3a and T3 instance types, and the previous generation T2 instance types.
Your point still stands.
Here's a m6g.2xlarge with 32GiB of memory and 8 vCPUs (AWS Graviton2 Processor / arm64):
https://instances.vantage.sh/aws/ec2/m6g.2xlarge?min_vcpus=8...
Monthly cost: $224.84
Fly compares better on the low end of scale.
AWS t3a.nano 512MB (2vcpu but 1h12min burst only, you can normally only consume 5% of that!) is $3.431/mo
Fly shared-cpu-1x 512MB 1vcpu $3.19/mo
They're charging you more for the convenience of `flyctl region add lhr` or whatever it is vs figuring out the routing & VPCs for example in AWS yourself.
It's not all roses though, I find the docs a bit lacking (good on networking - bad on security & storage) and the price of that convenience (aside from literal price) is that it is limiting - as soon as you think 'well I just want a very simple firewall/security group wall' or 'what's the equivalent of allowing only this x to y with an IAM policy', nope, can't do it.
Surely the amortized costs of the Fly.io platform don't scale linearly with the power of the VMs you rent from them. A 64GiB server won't be twice as expensive to maintain as a 32GiB one. So why not create a model that e.g. consists of a cost per VM/autoscaling group/etc plus whatever the VMs cost?
I'm genuinely interested, I could very well be missing something.
It's not just RAM, though, our physical servers have a consistent ratio of RAM/CPU/disk available. You're basically just buying slivers of hardware we've already paid for.
That said, our pricing isn't far off of AWS. They have much better economies of scale than we do, and can afford to do Graviton, but you get pretty close horsepower per dollar for comparable instances.
I used it in above comment as a shorthand for the fact that such an API/functionality exists that you can just say to Fly.io (whether via CLI or ~better~ other means) 'hey stick this in a second region too'.
flyctl is a reasonably simple GraphQL client. They haven't really documented the GraphQL API (see e.g. https://fly.io/docs/app-guides/custom-domains-with-fly/#grap... for one slice of it), but that's all flyctl does.
You can get equivalent hardware for much cheaper elsewhere and save way more than 9$. Add some load balancing across cheap hosts and you can achieve good reliability.
I run things on Digital Ocean, hetzner, ovh, linode, fly: they all cause downtime from time to time. I have clients on azure / Aws / Google and it's a much rarer occurrence.
Fly is definitely among the unreliable ones, so you're paying for a simplified deployment / ops experience here - the heroku model.
Even at that, configuration is not the easiest and there are gotchas you need to know. Documentation is also pretty sparse and the example images were stale. Getting a build to work was a huge pain.
The free offering is nice, but I'm not recommending it for production.
Apps get static IPs, at no point would you need to make a DNS change to route traffic to a new instance.
VMs can definitely fail silently. When apps fail, we're not very good at notifying you of what's going on. I'm hoping that improves this year.
https://fly.io/docs/reference/configuration/#the-statics-sec...
Not sure why giftyweddings.com needs to serve static files with an ETag in the HTTP response header: a hash of the file content is already included in the URL, which means it will invalidated when the content changes, just based on the path, regardless of the ETag.
Do you have a strategy for that situation? Your old Ansible branch hopefully is still there and compatible with your recent production code changes. And your sqlite DB is replicated to an external storage and your data is reachable even if you have lost control over “your” PaaS account.
You have found a hosting provider where this is not a risk?
a) Cloud providers; b) Hosting providers; c) Collocation; d) On-prem;
The only things that help protect you against this are (1) having named contacts on your account who will vouch internally at a supplier; (2) negotiating these clauses in contracts to prevent instant suspension in all but the most serious cases; and (3) utilising multiple suppliers.
I agree that in all cases, you depend on someone, but you can lower the risk to the point where other dangers that threaten your business become more probable.
Fly is different enough that moving to or from it takes much more time and effort.
It is? Fly seems to be only docker containers. You can't get much more standard than that these days!
Have you actually used fly? It's pretty vanilla. Having migrated to and from fly I don't think your assessment is accurate.
It uses Docker files or Heroku style procfiles - not particularly exotic. Your apps live in a private network, which is managed for you.
Your level of abstraction may vary. I stick to VMs after realising how much overhead I had with the cloud as a one-person shop.
Teams can interface with k8s which is also standard nowadays, taking you can spawn a cluster fast.
This seems like a trivial thing to add. If I wanted to use a domain specific configuration file then that's always an option, later.
Fly devs have confirmed this is in the works, think macaroons.
The only defense here is that Fly isn't the only party being incredibly sloppy with their API keys. Cloudflare is just as bad.
Doesn't that force you to use your hosting provider as your CI provider?
I've always avoided the automatic "deploy on git push" solutions because I want to keep my CI provider uncoupled from my hosting provider. That way, I can switch hosts without having to rewrite all my CI code.
About 1/3 of my deployments are just static sites, and the rest all fit into a single Docker container, so integrating Compose would be a big increase in complexity.
That being said, you can use Watchtower without Compose. It's enough to run
docker run -d -v /var/run/docker.sock:/var/run/docker.sock containrrr/watchtowerNote that my load test wasn't to see how much I could squeeze out of this, but to determine that it's "more than good enough" for this use case. The typical qps I get from real customers on the site is about 0.5 qps, so I think I'm good. :-)
Edit: apparently you can select the region, and you can possibly get a pre made DPA document template.
Customers with HIPAA compliance needs may need an Business Associate Agreement (BAA). That BBA thing is only included in paid plans.
But I also don't think you can build a from-scratch cloud without VC money. It's very expensive.
Linode did it 20 years ago (somehow). Dunno if there's a more recent example though. :)
I don't think you have to be "lazy" to take advantage of the value proposition of saving time or complexity. I think Dropbox and similar are successful because they have put a higher emphasis on their target audiences time than other competitors in the past.
I have a question regarding the change from a external periodic process (author mentions two: backup database, and send post-wedding email to couple) to goroutine:
- if fly.io ever needs to scale your app up to > 1 instance, won't couples receive duplicate emails? (similarly, any other background process like database backup would be run in multiplicity)
Because of all their generous free tier, I'm willing to pay more for my personal projects and literally host everything at fly lol.
Fly.io?
AWS Lambda?
As an example, node 12 was released in 2019 [1] and will begin deprecation phases in another month [2]. So it looks like a 4 year support cycle, although old runtime versions will still run. If you're using any of the AWS security scanning tools, these will flag any vulnerable dependencies and then you'd update them. Those seem fairly common in the node ecosystem, so other languages may see lower toil.
[1] https://nodejs.org/en/blog/release/v12.0.0/
[2] https://docs.aws.amazon.com/lambda/latest/dg/lambda-runtimes...
Edit: and only one app per server. Also only the Hetzner Cloud offering with managed backups and managed firewall.
As I'm reading it, your setup is the exact opposite of that.
Managing a server is nowadays with offerings like Hetzner the easiest way to run apps which just need to work and don't need to scale to exponentially over time. Why add anything on top of it? It's done really within 5 minutes if you know what you do.
If you need something which scales and doesn't require file storage / databases, then I'd go for the provider-agnostic Serverless framework. If you need storage and/or databases then it really depends, as Serverless gets costly at these things and more complex and there are more provider-lock-ins. Containers may make mor sense in such cases.
At our company we use these two systems to reduce maintaining basically to zero.
func getEnvOrDefault(name, defaultValue string) string {
value, ok := os.LookupEnv(name)
if !ok {
value = defaultValue
}
return value
}
for some of my projectsI normally do early returns, so it's somewhat surprising I wrote "value = defaultValue" instead of "return defaultValue", but whatever.
Is volume mounting the hardest problem in cloud computing ?
All buzzwords till now is scammy in the sense that, it doesn't make any sense when you can't cache your data in your local file system.
It kind of is.
If you don't provide a mutable filesystem, you get a whole bunch of benefits:
Need to scale to handle more traffic? Start another instance and add it to the load balancer.
App crashes for any reason at all (including a hardware failure)? Fire up a fresh instance on another physical server and load balance there instead.
Want to provide zero-downtime deployments? Start a new instance running with the new code, switch the load balancer over to serving from that instead, then shut down the old instance.
Each of these becomes a LOT more complicated if you offer read-write volume support.
Is there no facility for a CDN in fly.io ?
The point of embedding is similar to making a jar from an ops perspective. You “bundle” everything into one thing and you deploy that thing as single binary/archive. Less moving parts, less complexity.
Secondly, providing etags/modified headers for serving static assets is not a hack. It’s just not built-in. But it’s perfectly fine to serve those directly from your app server with the appropriate headers, so proxies can serve them directly.
In fact I would ask the reverse: why use a CDN when you can use canonical features for deployment (Go embedd, Java jar etc) and serving (HTTP cache headers)?
Personally I wouldn’t necessarily use a 3rd party library for this but to each his own.
* different work profiles (streaming large amounts of compressed data vs small JSON requests/replies) means that different software is more optimal for each -- an application thread usually is a lot more powerful and does a lot more (and thus is slower in the general case) than an nginx thread
* you're using your own compute on your own dime to do what your service provider is probably offering to do for cheaper
* bloats the image size -- there is a productivity difference between working with a 150MB image and a 1GB one (debugging, easily diffing changes)
Admittedly, these aren't problems for a one man startup or a very low traffic website, but these things do eventually matter as your system grows in size.
With the appropriate headers and configuration you would have the same kind of performance characteristics and resource utilization. Am I missing something here?
As for the last point: you have 1GB of stuff in either case. And all of that stuff needs to be looked at. The difference is really the deployment: do you prefer to move everything in a single piece or would you rather make individual, fine grained deployments? Either has advantages and disadvantages.
I’m not sure how Go let’s you inspect embedded files in a binary though. There might be gotchas. An archive is probably more transparent and more general tools can be used.
Also serving only the current version's assets will cause failures during a restart. If you have more than one container, there's a chance the page request will hit the new one, but the resource request will hit the old version - without the needed asset.
And once you hash your assets, you normally only deploy a couple of extra files, not the whole 1gb. I don't get why people would want to keep them embedded (apart from convenience for a third party who will run it)
> Fly.io looked more geeky and command line-oriented, which suited me, and their prices are also ridiculously low: free for up to three small virtual machines (I only need two), and $2/month for small VMs after that.
Their free allowance allows up to 3 small VMs and 3 GB of permanent volume space. I'm under that (2 VMs and 2 x 1GB volumes) so I get both my apps for free, and I could add one more. Even if I had to pay the regular rates they charge after the allowance, it'd be only $4.18 per month.
In any case, I was paying AWS $9/mo, now I'm paying Fly.io $0/mo, so that's what I saved.
Saves you traffic costs
for low, intermittent traffic sites, go on lambda might be a better comparison:
https://github.com/nathants/libaws/tree/master/examples/simp...
There is even a story of the Federal government forcing these companies to make sure that the letters had postage even though they were NOT being delivered by the USPS. Customers would actually pay a premium over the postage to the private company because the overall experience with private company was better than the USPS.
I am wondering if something similar is going to happen with AWS: more and more private companies will charge a premium for services built on top of AWS (and therefore reliable) but much easier to use.
(I imagine this is probably already happening and I'm sure HN folks will point out examples).
PS In the end, these private mail companies were deemed illegal and went out of business.