The Architecture Behind a One-Person Tech Startup (2021)
anthonynsimon.com
anthonynsimon.com
The best tech stack when starting a startup is one you don't have to learn.
There's a million things you learn when starting a company. Don't make yourself learn an entirely new tech stack on top of everything else. This advice means you'll be using whatever you've used in the past which might not be the sexiest or newest technology, but your users won't care. Your users want a working product. Choose the stack that will result in a working product the quickest.
Refactor and migrate to a better stack later, if necessary. (It rarely is)
For me, that meant deploying to Elastic Beanstalk (I know, boring, no one talks about it, but it works!) and using Mongo because I was already comfortable with it. It also meant not using React, at first. This was the right answer for me, but might be the wrong answer for you! Build your app on the technology you know.
This just comes down to a napkin calculation of expected value of building a startup given P(success). If your startup is looking like a real business out of the gate, using the simple stack you know is the way to go, and if it's more of a hobby project that you'd like to monetize eventually, learn something new.
What makes the most sense to be is to be really selective about what new technologies to use, and try to really learn ~one thing per project. E.g. my current project is a small search engine, and I've spent a lot of time exploring / figuring out how to use LLM Embedding models and vector indices for search relevance (vs. falling back on using ElasticSearch the same way we use it at work), but I'm using tools that are familiar to me for the UI/db/infrastructure.
I use Symfony (Php) and have not used a full SPA after I retired AngularJS (v1) like 10 years ago. What people now call server side rendering (SSR) is just how Symfony works with its regular Twig templating language (heavily inspired by Django's templating language).
As I gained more experience, I rewrote it. Once from vanilla PHP to Laravel, then later to Symfony.
I've not been exposed to any of the two and only dealt with PHP as part of messing with WordPress in the past.
I'm having some trouble finding analogies.
But maybe Symfony would be something like Linux Debian, has all the building blocks, it's modern but stable and well documented. Laravel is like Linux Ubuntu, it bases many things on Debian, but adds many things to make stuff a bit easier for the user. It's "shinier" and it has better marketing. You can add Debian stuff to Ubuntu, but you can't necessarily add Ubuntu stuff to Debian.
Symfony is more modular, you can add the components to any PHP project. Whereas Laravel uses many Symfony components and adds some syntactic sugar, but once you go into Laravel, it's difficult to stray away too far from the "Laravel way". Laravel uses many Symfony components, but Symfony can't easily use Laravel components.
Self-hosted Wordpress would maybe be comparable to a rooted Android phone. It has a very specific use case (for Wordpress it's fundamentally a Content Management System). You can add all sorts of plugins and additions. But it's also easy to accidentally break something. And once you added too many things, it might be difficult to update without breaking many things.
In the end, they're all Linux based, but living in very different ecosystems (just as Symfony, Laravel and Wordpress are PHP-based).
In programming terms, Symfony might be similar to Django (Python) or Spring Boot (Java), whereas Laravel is "cousins" with Ruby on Rails.
The main difference is that there are IDs that allow the client side code to seamlessly attach event handlers (this is called hydration) to the DOM - and that there is no difference between server and client side code.
In your case, you'd have to do that manually - a huge difference (speaking as someone who used to do that with jQuery).
I certainly didn't want to be manually creating instances, installing Java / Node / Docker, starting / stopping services and doing deploys manually because that's also annoying work :) Elastic Beanstalk + the Github actions deploy plugin is dead simple and just works.
Has kept surprisingly well over time, with some hiccups around platforn version changes.
Am I happy we moved off, yes. Was I impressed how big we took a very simple "starter" setup, also yes. Again one of the biggest reasons was that all of our competence in the company was no longer on beanstalk but on ECS where we had solved the issues we had with it with more complex setups more in our control. If I were starting again would I use Beanstalk? no not me personally but when we started ecs didn't exist and beanstalk was better than bare EC2. now I know and am comfortable with ECS. Did it help the company? very much so, we leveraged the heck out of Beanstalk to scale the company horizontally and vertically.
cant remember the chapter in "Getting Real" - but this was one explicitly mentioned in their famous book!!
I always wonder if python, auto scaling, and cloud change the kinds of businesses (and margins) required. Running 60+ pods of gunicorn for an app incurs lots of overhead, especially if one is using small VMs.
I can't count the number of war stories I've heard from SRE friends who joined some company and realized they were taking over a Django stack for the core of a business that was blowing gaskets left and right.
This sort of thing just leaves me wondering if margins have to be high to support a slower language, slower framework, and complicated deployment model to work around the Python GIL.
Call me jaded but I just wonder if it's worth the time to market to do Python these days versus say Go and be able to deploy a static binary and use all your CPU cores simply. K8s is way less tempting for a "one-process-architecture " (minus database and maybe nginx) until the system is much larger (instead a couple of systemd units would sort you)
Are you saying that to run this 8 processes you suddenly need 8 VMs running in the cloud - yeah. That does seem expensive. It’s very tempting to say one should not start in the cloud but host locally, but that seems an unpopular view.
But maybe as we start to see more and more local workspaces catering to people working from home, we will see more and more “mini data centres” - it might even by Oxides sweet spot.
Edit: “””
because it actually winds up being less DevOps work, on average, to support open-source systems running on bare VMs, than to try to keep up with Google’s deprecation treadmill
“”” https://steve-yegge.medium.com/dear-google-cloud-your-deprec...
Mr Yegge makes my point much better ofc
It’s kinda like using a map reduce cluster when a beefy server with a few Linus commands piped could have handled the “big data” needs. Both have their time and place, but it’s amazing how far a simple setup can go, so don’t overengineer prematurely.
Hell there’s even first class support to compile assets into it if you wanna throw your whole site into the binary, or many other use cases that would otherwise make docker nice.
Damn I miss the startup that was a Go build and deploy to 3 static VMs over my current company’s dozens of microservices + ecs + kafka + aurora postgres.
Startup A had the servers all sitting around at 2% cpu all day, trivially scalable by adding more vms. No downtime over 3 years.
Startup B had 10x the users, but an elaborate docker + ecs + dozens of microservices (that could have just been libraries, avoiding all that complexity) + aurora postgres. Downtime all the time. Most recent was so hard to detect: a db server running fine at 40% load and our application around the same, yet db timeouts left and right so considered an official outage. Aws had throttled the network throughout between the DB and App with no warning! We even had an AWS support person on the call and took them 2 hours to figure it out.
Anyway, it's patently obvious that increasing your unit costs by 7 orders of magnitude will constrain the kinds of business you can run. A few people will deny that, but the cloud is so impactful that those are very few nowadays.
What a lot of people will deny is the relative importance of those two. If you take your C code from an auto-scaling environment and replace it with bare metal Python, you often get a few orders of magnitude gain.
Likewise, you only hire a SRE if a.) you have the budget to hire at all, which the vast majority of startups do not and b.) reliability is shitty enough that you need to hire someone to fix it. All the Django setups that hum along fine do not need SREs; the founders continue to rake in the money and don't need to hire anyone.
This setup is very similar to the startup that I run. We have used k8s from day one, despite plenty of warnings when we did, but it has been rock solid for 3 years. We also decided to run Postgres and Redis on k8s using Patroni + Sentinel and the LVM storage class. This means we get highly performant local storage on each node and we push HA down to the application.
Was there a learning curve? Yes, it took a solid week to figure out what we were doing. And we've had the odd issue here and there, but our uptime has exceeded what we would have had with a cookie cutter AWS setup, and our costs have been a fraction (including human capital costs).
- Every service is deployed via a Helm chart and using containers
- GitHub actions build the container and deploy the helm chart
Some of the details that matter:
- ACK is used to create AWS services (RDS, Redis, etc.) via Helm charts (we also have a container option for helm charts as it's faster and less expensive)
- External Secrets is used to create secrets in the new namespace and also do things like generate RDS passwords
- ExternalDNS creates DNS entries in Route53 from the Ingress objects
- Namespace name is generated automatically from the branch name
- Docker images use the git hash for the tag
Some things that are choices:
- Monorepo although each service aims to be as self-contained as possible.
- Docker context is the git root as this allows for a service to include shared libraries from the repo when creating a container. This is for case where we break the previous rule.
Have you deployed and managed k8s before? What did you find more complicated than learning the ins and outs of the cloud variants?
The former situation does suck. K8s is amazing and once you understand it, it is easy to work with. But if you haven't learned the concepts and core resources, it will definitely appear as a black box with a ton of complexity.
But I think for many of us who have used Kubernetes a lot, it is a no-brainer in a lot of situations as it doesn't matter which cloud provider you're using (for the most part), you get a common and familiar interface.
It is funny (or depressing, depending on your point of view) to catch people doing this by asking them when was the last time they did X or how they know whatever they’ve stated and then seeing the gears turn.
How about this equivalence: I appreciate how extremely sophisticated GCC is, and the very well optimized output it generates. It is still is 100x more complex internally (and thus, error prone, buggy, more complex to modify when needed) than TinyCC, for instance.
I'm also not really sure what the relevance of LoC here is though? The Linux Kernel is a large codebase...but surely you don't object to using that?
100% agree.
At a mega corp they were mandated to use Azure(worst decision) and we choose Azure container apps.
Nothing worked: from the deployment ARM templates to scalability.
It was a disaster.
This was after we had the highest level of support from Azure.
The engineer working with us in this almost admitted this shouldn’t have been released in it’s current form.
I can’t for the life of me understand how Microsoft releases such products!!!
Depending on training and experience, it may be easier to setup a GCP than a Kubernetes based solution, but it's still not exactly trivial. And, once a fitting Kubernetes platform is up and running, it's almost trivial to add new services, as the article describes.
I think, even in 2024, and even for a one-person business like the authors', starting with Kubernetes is not a bad decision. I'm currently building my own Kubernetes based platform, but on a less expensive service (Hetzner Cloud). I use a separate single-node Kubernetes "cluster" with Rancher for management, and ArgoCD for continuous deployment from git. Currently I only have a single-node "playground" cluster up and running for my "playground" workloads, to save costs, but I will upgrade this to a proper HA cluster before a service goes live.
Later, I can still upgrade to GCP, or in my case probably Azure. Until then, I don't need to worry about large cloud bills, which is also a good thing.
I remember building a todo app as my first SaaS project, and choosing something called Stormpath for authentication. It subsequently shut down, forcing me to do a last-minute migration from a hostel in Japan using Nitrous Cloud IDE (which also shut down). Just pain upon pain.[1]
Now, you can just pick a full-stack cloud service and run with it. My latest SaaS[2] is built on Google Cloud, so Authentication, Cloud Functions, Docker containers, logging, etc straight out of the box.
Not to mention, modern JavaScript and CSS are finally good. With so many fewer headaches, it’s a great time to build.
[1] Admittedly, I was new to software dev and made some rather poor tech choices
I’m lucky that my work is event-based, as is it used by in-person live events so my usage comes in waves (pre-sales a month or two out, steady traffic the week leading up to the event, and high traffic the day before or week/day of the event). This means that at worst I only have to ride out the current “wave” and then I have some amount of time before the next event (gives me an opportunity to fix run-away costs.
One of my big runaway costs was when I tried to use something like Datadog/NewRelic/Baseline. You work yourself up to the cost of the service, make your peace with it (the best you can, since it’s also hard to estimate), then get hit with AWS fees (that none of the providers call out) for things like CloudWatch when they are pulling logs/metrics out. I’ve had the CloudWatch bills be 4-6x as expensive as the service itself and it’s a complete surprise (or was the first time). Thankfully AWS refunded it when it happened. I caught it after 2 days and had run up a few hundred dollars in that time, I could have handled it but thankfully they refunded it for me.
The second runaway cost was Google Maps, once you fall off that free tier the costs accumulate quickly. In just a few days I had a couple hundred in fees from that. I scrambled a switch to ProtonMaps and took my costs down to a couple dollars a month.
All these services are predicated on the idea that you never want your site to go down, and will pay anything to keep it running. So if you start logging gigabytes a second, in the old world your VPS would've started failing and your website's buttons would start showing errors to the user. Now, the user doesn't see an error, but you get charged hundreds or thousands of dollars a month to keep it up, even if your website generates you $50/mo.
Many people would say a detached, Typescript frontend is less simple than SSR. At this point in my career writing a bunch of ad-hoc, per-page Javascript would be more toilsome than writing components. Thus, the prevailing idea of "simple" would be less simple for me.
In that same way, if you've spent a large chunk of your career in Kubernetes then managed Kubernetes is probably a lot more simple than maintaining an image production pipeline and a disjointed Terraform deploy. You'd have to make a choice between immutability and non-immutability with VMs as well, and design lifecycles for them if you choose immutability. That choice is made for you with Kubernetes and has well-established patterns.
That's to say, simple is highly subjective. Like another comment said, simple is the technology you know best today (and that you can easily hire for).
I'd like to emphasize the yolo from the GP. I does really not imply image maintenance and immutability. (By the way, "non-immutability" is called "mutability".)
And no, if you know kubernetes to your heart, it's still very often easier to start the grug way and build OPS up only after you have something to run on it. (But, of course, there are exceptions. There are always exceptions.)
The trouble with "single vm and yolo" it is tech debt of the worst kind. It's not a shortcut you might have to expand on someday, it's the kind where whole platform changes are necessary. Someday that vm process won't be enough and that'll require changing everything about the deploy, potentially at a time when big changes aren't desirable. If the business is starting to pick up so that we need something more reliable, that's the wrong time to want to replatform. I'd rather but in a small amount of upfront knowing that I won't have a big tech debt to pay back later.
So I keep it real simple, and keep away from parts that I know will give me a headache. I use a managed service and plan to give it to somebody else when we're big enough to have that somebody else. Then they can worry about the parts I didn't.
It's wise not to do too much upfront, for sure. But it's wise not to back into corners that are hard to undo later, too.
For a surprisingly large set of cases, that day will never come. I've scaled services to millions of users on a single VM. The best debt is the kind you never have to pay off.
FTFY
My only complaint Fly.io doesn't have managed DB (https://fly.io/docs/postgres/getting-started/what-you-should...) and task queue.
At work we use GCP, but try to be as high as possible in the abstraction layer (e.g. use Cloud Run and Cloud SQL).
Moreover, they still allow 1 free db per account:
> Supabase offers one free, resource-limited database per Fly.io user
Are you concerned that it is not directly managed by Fly.io?
* also Kafka: https://fly.io/docs/reference/kafka/
There are no production dbs in k8s, no service meshes, monitoring is only side-cars that send data to an off-cluster backend,
The automatic deployment with flux seems pretty nice!
Especially if there's a chance I need to run more than one client, or environment, on a given server, which is more important for "solo" enterprise than for a bigger company because I do not have money or time to care of multiple servers.
Personally, I think that containers are a good choice for most webdev projects.
If you are just starting out, then Docker Compose is probably very much sufficient for anything that touches upon a single node.
Realistically, once you need to most past that, I feel like most folks would be served just fine by something like Docker Swarm + Portainer, you'd still need to run your own web server in front of your containers if you want something similar to an ingress, but in my eyes that's a plus, given how simple the config is (if you've ever installed Apache/Nginx/Caddy it's very much like that), or maybe go with Traefik if you must. There's scaling, resource reservations and limits, overlay networking between nodes, port mapping, storage, restart policies, configuration and pretty much most of the things you might reasonably need. Something like Hashicorp Nomad could also be mentioned here, but Docker Swarm seems like the simpler option to me.
Past that, it's not like you can't run a bit cut down versions of Kubernetes, that will decrease its surface area somewhat - K0s, K3s and others all try to do this. I personally rather like K3s, it also integrates with Portainer/Rancher nicely if you need a dashboard, you can take advantage of Helm charts and all that other good stuff. I found the choice of K3s using Traefik by default (when I last used it) a bit suspect (configuring a custom non-ACME wildcard cert as the default for the ingress wasn't exactly well documented), but overall it was an okay experience and the resource usage wasn't anything crazy either.
For my personal stuff nowadays, I still run a Docker Swarm cluster and I've no complaints there. It doesn't have autoscaling out of the box, but then again, my workloads are too boring to need that and I enjoy the predictability in billing I get.
You can get to gigantic success and extraordinary scale by growing vertically, without needing to add a very complex layer that kubernetes is.
Think of the many literal millions of lines of code, and thus, potential bugs and liabilities one inherits by "just using kubernetes".
It's not a requirement, neither problematic. Its simply unnecessary for anything remotely resembling a "one-person tech startup".
Besides unnecessary features like horizontal scaling, necessary features like restarting a service already exist, and most likely are already installed in the base distro and being used by the rest of the system. Its not like this couldn't be done before, all kubernetes provides is a conveniet and streamlined package for all that, at the cost of black-hole levels of extra complexity via several added indirections.
It adds vast complexity and is not needed.
Wonder if they kept it as is, and how many people manage it nowadays.
We ran Panelbear for about a year before merging it into Cronitor's codebase.
Cronitor is also a Django monolith but on EC2 and was already humming along nicely for several years through frequent traffic spikes.
We had no reliability issues with K8s, but came down to: as a small team we decided to have less moving parts.
The new setup is not too different from what I describe here: https://anthonynsimon.com/blog/kamal-deploy/
I wonder how they make time to manage it all _and_ iterate on projects _and_ have a life outside of that.
This is the biggest problem I've long struggled with. I don't know of a good solution.