How much I’ve spent so far running my own Mastodon server on AWS
micahwalter.com
micahwalter.com
- Cloudflare DNS
- 2x CPU 2GB RAM application server with snapshots
- 2x CPU 2GB RAM postgres server with automated backups
- 2x CPU 8GB RAM elastic search server to provide a full-text search for the known Fediverse
The search is not mandatory, but super nice to have. It is about 20€ a month, not counting the domain costs due to having it already for other purposes.
Akkoma is written in elixir and is much better in handling async messaging. No need for Sidekiq or Redis. My CPU usage hovers around 6-12% and memory around 55%. Postgres uses about the same.
I've never tried Mastodon, but having been a Ruby developer in my past life, I needed to find something else to install when I joined the Fediverse. Couldn't be happier with my setup. It just works, I never really need to do any maintenance.
All the normal clients work with it.
Probably? (Maybe it explodes if you try.)
Even GotoSocial now that they've put special handling in to work around their bizarre key-handling design decision.
fediverse.observer points to https://gotosocial.fediverse.observer/gotosocial.sequentialr... as being open.
[1] Unlike honk, which is also in Go, but was written with little regard for readability.
Note that “If your monthly egress data transfer is greater than your active storage volume, then your storage use case is not a good fit for Wasabi’s free egress policy”. There are reports of bill shock from mastodon users as a result of the costs due to the policy.
I recommend turning on tiered cache on cf [0]. Our b2 bill went from $7 to $4.
[0] https://developers.cloudflare.com/cache/about/tiered-cache/
You'd likely be in the free tier.
they have cheaper storage and up t0 150GB is free so maybe try this one as well? i am on the fence for a dirt cheap S3 myself but B2 and cloudflare R2 is in the $10-15/TB range and something cheaper would be nice for a small usecase
That's where I actually like having the sidekiq component separate and scalable. Because it means I can make use of dynamic scaling in a cloud to work through the queues, then scale down when the queues are empty.
The Mastodon web service can run on minimal resources, even with 1000 MAU's, it's only Sidekiq that needs to scale, and that also only when there is large amounts of messages to federate.
I don't really know how Mastodon does this with the Sidekiq, but what I've seen people with a very small amount of activities (we're talking about hundreds of thousands, or a few millions), need to do manual scaling of Sidekiq processes or threads. As a user, you should not be tuning these metrics to handle this kind of activity. A single beam process could handle all the WhatsApp messaging back in the days, so it can easily handle our relatively tiny fediverse servers with much less users.
Yes it can run one big process of everything, and does when I find it. But some routes or subservices are slow or hog cpu or use way more memory than the other routes. We'll have a handful of infrequently used routes that use >1GB while typically the monolith stays inder 200MB of ram. We'll have some expensive route that hog the even loop for a second. Mixing these routes in with the well behaved normal routes means we have to way over-provision, and means sporadic performance degredation once se lesser behaved routes start firing.
Separating your stuff up is a huge win. Being able to see ingest, egress, and web traffic separately & to manage them separately is a huge win. Im not sure what the win is from a single process. I have a hard time thinking of advantages. There's some small wins to memory, but a well designed forking architecture works around most of these advantages.
A service such as a small community Fediverse server does not need microservices and complex architecture. What I've been seeing, people are kind of struggling with Ruby and Sidekiq even when having servers that are not super big. In the past, WhatsApp was having a small group of engineers with a handful of servers, serving hundreds of millions of users with Erlang. To get people running their own Fediverse servers, it should be as easy and as cheap as possible. Akkoma is a good start, and in a few months I hope we see the first proper systems made with Rust and Go, running with the minimal amount of resources, being single process and federating everything without the user needing to do any maintenance.
Mastadon aint fast, but, this current architecture has been invaluable in letting folks degrade gracefully. Sidekiq for ingress gets backed up, but often the web service is totally responsive and egress sidekiq (if seperately configured) works fine. Picking a more effective runtime might give you a lot more headroom, but I'd still want a plan & ability to cope & degrade gracefully when the system does come under pressure, and Im not sure Erlang/Beam or any other single process system is going to have such innate ability to have the different things dealt with differently.
Bulkhead pattern is a truly great one- keep the whole ship from sinking when one part floods.
https://learn.microsoft.com/en-us/azure/architecture/pattern...
Mastodon has chosen a Service oriented architecture for this, and I'm happy with that decision because it means I can keep the web and DB queries running on minimal resources, while scaling the queue processing up and down depending on current activity.
One could say that incoming http requests from end users is a fairly predictable level of activity, while queue processing is a very unpredictable level of activity. Because it's tied into people's posts, their followers, the number of instances their followers are on, and so forth.
And with Mastodon v4 the resources required for end user http requests are even fewer, because they've shifted more processing onto the client with Javascript.
So I think the freedom to be able to scale the queue processing separately is fundamental to me. I can't imagine what it would be like to run a different software where I had to scale a monolithic service.
It doesn't really matter how fast the software is because queue processing is about blocking tasks, and they're blocking because they involve contact with other services and servers on the internet.
So yes I'd love to see a sidekiq replacement in Golang, Rust, or even Elixir, but I would not like to see it bundled in with other services. Mainly because queue processing deals with blocking tasks, and incoming HTTP requests should never be blocking.
What mastodon is missing is a good catalogue of reference deployments for 1, 10, 100, 1000, users on the major platforms (AWS, GCP, Azure, DO, OVH, Hetzner). Of course, “users” is a rather poor proxy for “activity” but it’s a start.
There are “scaling” articles that say “the first thing you’ll need to address is…” and breathless blogs of “what we did when we doubled registered users overnight”. But not so much if you’re starting an instance for a community that you know will be a certain size quickly. How many cores should I expect sidekiq to need to service each thousand typical users?
Sure, it’s a piece of cake for sysadmins or an AWS solutions architect, but it makes me wonder how many fragile instances there are out there that have grown fast but could also fail fast.
The way I read the post showcased how easy it was for the AWS SA to make insane billing mistakes with unneeded things like CloudWatch. They also hide the cost of EC2 because there's no charge for that instance type until 2023. This isn't a positive advertisement for AWS in my mind. This is in line with experience I've had with former AWS consumers I've converted over to DigitalOcean in 2022. I saved those conversions, on average, 70% of their prior AWS bill - with no loss of functionality. AWS has spun out of control with the nickel and dime aspect to billing. I can't say I'd recommend it to small or very large clients unless you have money to burn. At the end of the day "traditional" architectures on AWS are not cost effective without excessive oversight of choices.
If you sit down and work out their cost of data processing, it costs them cents to provide a service for which they charge tens to hundreds of dollars.
I figure the reasoning is that large enterprise requires logging for compliance, and it’s a hidden charge that’s easy to overlook in a bill.
If you already own a domain you can use a subdomain and cut even that cost although it might make up for a long URL.
I think the main cost here is time. Probably far more expensive than all of those costs. Even on AWS.
- Cloudflare for caching, DNS, domains. - Hetzner, Linode or DO for a cheap instance
See what does/doesn't work. Storagebox on Hetzner is cheaper storage hosting and you can have redundant backups for cheap on backblaze as well or offshore cheap VPS doing encrypted backups with rclone/rsync.
I don’t get why people assume others are just mindless fruit flies in these threads. Maybe OP just doesn’t want to fiddle with technology and pays Amazon to simply get it done. That’s what I’d do, there’s no need to turn every project into excercise in orchestrating open-source software.
In fact I have to do a bit of AWS for work and I hate it with a passion. It's spread over a ton of services each with their own console (that look & feel like they were designed by completely different people), lots of overlap between services, complex billing. Configuration often needs to be done with complex json files. Ugh. Perhaps it makes sense for a major enterprise where you have orchestration for everything. But I would never pick this total mess for a personal project.
Azure has a bad name but at least everything is integrated in the same console pretty nicely.
But really if you just need an instance of something for personal use, DigitalOcean, Scaleway etc are so much simpler. And cheaper, and most stuff is just included in the monthly price so you don't get billing surprises.
AWS - "Free to get data in, expensive to get out"
What’s sneaky? OP states at the beginning of the post they’re a solutions architect at AWS… and provides reasons you may not want to use AWS if you want your own single user mastodon instance.
First of all, a 5USD/mo instance cannot run mastodon. Mastodon needs a lot of resources from the get go, even with one user.
Second, you’ll want to enable backups, which is 20 percent of instance cost as extra.
Third, you will want more disk space than is available on a base instance. Mastodon is huge and disk space is taken up pretty quickly. With AWS you can set up autoscaling drives, DO doesn’t. So, more costs…
Finally, if you want to use DO Spaces, that is a flat 5USD/mo. AWS S3 is far cheaper for what would be used in this instance, on the order of free tier or couple of cents. The OP did the set up incorrectly which is why their costs are far higher. I actually did set up AWS S3 on my DO instance and the costs are extremely low.
So why am I using DO? I have credits on it, that’s all. I’ll move to AWS and probably set up Takahe at some point: https://docs.jointakahe.org/
In Linode you can get additional block storage for a relatively low cost: https://www.linode.com/pricing/#block-storage So a 5 or 10USD instance should be enough.
I will say that S3 is very expensive not because of the upfront cost. It's because of traffic. Also cloudflare is more affordable than AWS.
I also use DO and first of all Mastodon itself, the web services, streaming API, can run on minimal nodes. I have almost 1000 MAU's now and only need about 600M RAM for each web service container instance.
And backups, for that I use postgres-operator (yes it's kubernetes), which does S3 backups and those cost almost nothing. (I know cost can go up if I need to restore because of egress traffic)
Media is also in S3, with a Cloudfront proxy. As soon as I enabled the CF proxy my S3 costs went down to 0.1 dollars basically.
And finally, thank the gods for kubernetes because it allows me to dynamically scale up sidekiq when the queues are backed up.
So my base cost for 1000 MAU's on DO is about 150USD/month. But it goes up if Sidekiq scales up temporarily, so it's a fluent final cost based on how many times Sidekiq had to scale up in a month.
Most importantly for me is that nothing in this setup is vendor locked. S3 compatible services are plenty, cloudfront web proxy services, plenty of those out there, AWS SES for outgoing e-mails can also be replaced. And of course kubernetes with proper IaC can be deployed into any managed k8s solution.
DigitalOcean no longer has a USD 5 instance, it's USD 6 now.
Also, there’s a new $4 entry-tier droplet that has the same CPU, the same amount of ram and the same amount of storage as my $5 droplet used to: https://www.digitalocean.com/pricing/droplets
The included bandwidth is only half as much (500GiB), but I suspect most people who would have gotten $5 droplets in the past will now be getting $4 ones.
The AWS instance in the post, t4g.small, has 2 GB of RAM. So it's not really comparable to the $4 DO droplet.
Also, FWIW, I’ve load tested different VPSes between AWS, Linode and DO and found that for reasons I don’t understand, AWS ones perform considerably worse than the other two platforms VPSes with identically listed stats.
That said I assume the author is not trying for cheapest-possible.
Lastly, the more people run their own instances, the worse everyone's experience will get. The fediverse almost entirely relies on direct communication between instances with no DHT or other mesh routing overlay, so messages are O(n^2) relative to number of instances. Caching works OK for assets but not for ActivityPub RPC, so many instances don't even implement it at all. One of the main reasons the large instances struggle so much is that these communication patterns drastically increase traffic volume and queue length, which is further exacerbated by the fact that Sidekiq (the part of Mastodon that handles those queues) is a poorly implemented tuning nightmare.
The best bet IMO is to join one of the medium-sized instances aligned around a community or vibe that you like (I'm on hachyderm BTW and it's awesome for me). Smaller ones have all of the problems I've mentioned above, while the very largest are sub-optimal in terms of both technical UX and social atmosphere. Goldilocks wins again.
I quite enjoy Mastodon, but the whole thing feels.. incorrectly aligned with technical needs. I hope the attention it's getting leads to core implementations more aligned with the community.
I tried, and gave up, setting up Mastodon several times in the past as I got choice paralysis from all the different ways to set it up. That was a few years ago, but from what I hear, it is still a somewhat daunting task.
If you really need scale, AWS can be cheaper than rolling your own infra and abstractions, granted you make it a mission to keep laser focus on AWS costs.
Primary ideal use case for AWS is internal Amazon usage. Charge internally but write off as expenses, for infra you had sitting idle anyway. Pretty sure AWS and rest of Amazon are separate legal entities to take advantage of tax laws.
Renting is usually more expensive than owning. The larger you get, the more expensive renting is compared to owning.
For anything else, there are always going to be cheaper and better solutions.
You probably shouldn't use it for day to day ops but it's a really good spillway for extremely temporary & particularly for unexpected spikes of traffic.
Of course you could argue that AWS is particularly bad but I think this applies to most non-bare-metal instance cloud services.
OTOH, if you have an understaffed infra team that gets by by accumulating technical debt on infrastructure, AWS takes away that option at several levels of abstraction.
I think we pay 16.81 Euros/month.
I am fairly sure a mastodon server should run on a server with half the stats and price.
But once there are a few hundred users, Sidekiq jobs will use a lot of compute and egress data will add up. Might be better in this case to use a provider with generous free bandwidth.
There are S3-compatible object storage providers that have no per-request fees, and significantly lower bandwidth and storage fees. Examples – OVH, Scaleway.
I had an use case for object storage where my service would be uploading tens to hundreds of small objects per second, 24-7. Running some napkin math, it looked like AWS S3 costs for me would have been in the hundreds. I went with OVH, and have a €1 monthly bill for the object storage :-)
$20/month, not including compute, is their starting point on AWS. As they get more followers and follow more people that will continue to grow.
My biggest suggestion is to disable the relay. I took experimented with using a public relay, but the stream of content it adds is very low value. Better to create a second account on the instance that is more aggressive about following people your only tangentially interested in. This will make your federated feed interesting, but keep your main feed more focused.
To the people suggesting DO/Linode. I'm a big supporter of Linode, however article has no egress or compute costs. Likely because it's a single user instance running on the free tier. Their costs are almost all storage related.
[0] https://github.com/g3rv4/FakeRelay [1] https://github.com/g3rv4/GetMoarFediverse
What was the exact attack vector in the case of nyevgeny?
Yevgeny did exactly this. He compromised a web server running Apache or something, used that to brute force his way into a networked work machine, which took weeks, but went undetected. That machine turned out to be the physical host the VM running the web server was on. Then he used the work machine to break into LinkedIn using credentials and VPN profiles found on the iMac.
This vector would have been impossible had that open server not been running so close to the work machine.
If more people sign up to his instance, or he just uses it, the costs will keep going up due to the caching Mastodon does for assets (and especially due to the fact he also doesn't have purging enabled)
That will add up. The cloudwatch costs are just ridiculous.
Even with the high energy costs in Europe I'll just be migrating to a more efficient machines and I will lower my power budget from 100 watt to 50. For now, with the current 100 watt I can run 4 nodes each with a hard drive and ssd, 16 gb of mem. These are hp 6300s, old but low power and if needed they can stil pack a punch with the i5 CPUs.
So this year I bought at old Dell T7910 workstation and fitted it out with a 14 core E5-2680 V4 (second one coming), 1TB nvme, 64gb RAM and has cost me around $400 USD. Runs at around 95 watt and I already have a good fibre connection any way.
For a similar spec machine I would be paying this every month and now I can just keep throwing things I’m interested in at it and so far has been overkill for a relatively fixed price. More messing around with hardware but something I mostly enjoy.
The incredibly inefficient Rails framework Mastodon is written in annoys me with its wasted performance but honestly the load for one or just a few users isn't even enough for me to start looking for an alternative. I'm sure I'd need to switch to a more optimized ActivityPub server if I'd get 10k users on my server, but for a small private instance you really don't need anything all that fancy.
Electricity per kwh is quite expensive these days, at €0.43/kwh because of war. My hardware is quite inefficient too, being cast off gaming hardware that is way overkill for the use case. Still, it averages 200W under load so that's (0.43 * 24/5) or €2/day.
So €65/mo, with some of the world's most expensive electricity. With 2021 energy rates the electricity cost would have been one third of 2022, at €25/mo. Use more reasonable hardware like a pi 8gb or a nuc or something rather than a used gaming PC and you could get even lower, though obviously you'd want to depreciate your hardware cost.
Meanwhile comparable compute from AWS is $1/hr
1. I want to write text-only posts and have them distributed to whoever is following me on Mastodon.
2. I want to receive all the posts of people I follow, and then just display/store the text and nothing else.
Does something like this exist? Seems like it should be super slim and easy on resources.
Minimalist and opinionated but is a single binary with an SQLite db. Source is, uh, esoteric but is definitely hackable (there's a whole bunch of forks.) Seems to work ok with a few minor weirdnesses[1].
Should clarify that it doesn't have an API - you can't use Mastodon clients against it (until/unless someone writes on) - it's basically web only.
[1] I can't follow someone on my GotoSocial instance from honk but I'm pretty sure I've debugged that as being GTS's fault[2] rather than honk.
[2] They made many odd design decisions which cause(d) issues with federating to other systems.
[edit: clarify that honk as no API]
[1] https://github.com/superseriousbusiness/gotosocial/issues/11...
Ugh!
Is there a fork that just uses whatever stupid key/keysigning choices are needed to play well with the maximal number of other federated thingies?
I'm just going to be sending out a single post per day, tops. And perhaps reading whatever single digit number of people I decide to follow (maybe double if I'm really bored).
If I have to generate and send private keys in the clear to get this done I honestly couldn't care less.
Edit: clarification
For added entertainment, I installed it manually on FreeBSD, because I like playing old-school 'I do things by hand' sysadmin at home.
It's connected to the world through a free Cloudflare tunnel. I have a residential fiber connection. It doesn't use S3, it just dumps files to the filesystem. I did clamp down on federation settings, and it purges old content aggressively so federated content doesn't make the hard drives metaphorically explode. One user following someone on mastodon.social and you can see storage usage climb by gigs a day with defaults.
After the initial setup it's been honestly quite painless. It just works. If people start relying on it, I'll probably move it to masto.host.
I was considering running mastodon too but someone mentioned this here and it does indeed look great
AWS is the most vendor-lock dependent company, IMO. Every single "architect solution" of their is a recipe of how to never get your data back for free.
Thanks for sharing.
its
You come to me and ask "how much does it cost to do X", and i know? I'll tell you. I'm a CC0 / public domain kind of person. I don't get the secrecy. Look, how many full cores/threads and how much ram/cache, and how much storage/io, and how much ingress and egress. Sure, those you can obfuscate a bit, but can i run a 10k user instance on 2 dell rackmount machines from 2014? on a single Ryzen? A single Epyc? a half dozen raspberry pi?
the chances of being able to parry "i run a php webservice" into retirement levels of money is nil, so i don't even understand the premise!
To answer your question more directly: I run misskey on a 4 thread, 8GB mem, 80GB storage VM, where i interact directly with the hypervisor (proxmox). my cost is nothing, but i guess this would be maybe $10/month? I'm not including the monthly cost to host the hyper, the /24, etc; we used to own the metal, but my partner decided there are things more important than owning the metal, right now. The upside is, when i take a snapshot, that isn't managed by disks that i can physically touch. Makes them slightly more reliable to me. Other people may have weirdness about this.
of the 4 "cores" the daily average is about 3% CPU usage, with monthly spikes up to 12% average usage. as far as bandwidth, the daily average is under 9kb/s and the monthly averages are under 22kb/s. a small fediverse instance with a single user is likely completely doable with a co-lo raspberry pi!
You buried the lede:)
That's great to know. Maybe not worth it to cloud host. Are there any issues to worry about with having downtime of my own node, other than my own ability to post and read? I haven't read enough about ActivityPub to understand the impacts of a server being down on the protocol.
Mastodon is php, and as such it requires php-level stuff. Misskey is nodejs (i think) and python, so it depends on if you're used to wordpress-y stuff or not. I prefer misskey because it has a "drive" - you can upload files and hotlink, or make "tweets" with them.
thankfully, activitypub and Ostatus "protocols" are open source, so one could write a server implementation in whatever language du jour makes them happy, but only a few people have.
I see a lot of the same issues with realtime applications like matrix (for chat) and audio/video streaming. Some php implementation (or python), and then no one else can make something that adheres to the standard. Whether that's the standard's fault or not isn't for me to say.
if i get a jpeg image file, and the joint picture expert group's spec for jpeg format, i can write something to view that in a day or so. same with whatever RFC IRC was. or http. or ftp.
try writing a toy matrix or fediverse implementation? meh.
Wat. It's Ruby/JS.
Not always. @kris-nova here is the public face of hachyderm.io and has done an excellent job communicating around architecture, incidents, etc. Everyone on that team deserves kudos.
you are in luck... i've used this and it works out of the box