How Instagram scaled to 14 million users with only 3 engineers
engineercodex.substack.com
engineercodex.substack.com
How can they know the internal infrastructure but have to assume the app language?
Edit: So the entire piece is taken almost verbatim as-is from a couple old articles on instagram engineering blog.
It might as well just redirect to: https://instagram-engineering.com/what-powers-instagram-hund...
This is against the guidelines:
> Please submit the original source. If a post reports on something found on another site, submit the latter.
[0]: https://instagram-engineering.com/sharding-ids-at-instagram-...
Size isn't a bad thing anymore since price has dropped exponentially since the inception of Instagram.
I am positive they would use another modern technology today if it was present in the past.
Fantastic read though.
It's virtually impossible for anyone to hotspot in a meaningful way with this system.
The solution in most cases is a simple database that acts as a pointer database user db -> user's db. That is generated on the creation of a user.
From here you create some simple cold storage models ( if user isn't active ) and some warm models which will scale out the db if the user's db grows it shards and replicates for more read access. But the last thing you want is to slow replication or have one DB that can't move to balance resource utilization. There are some new DB tech that does this without even sweating the deets.
Just my own brain reading through old talks and articles from Instagram engineering and Excalidraw for the diagrams.
I did my best to put together all the info I learned from them into a comprehensive and simple manner.
Sometimes when a paragraph I write reads a little too harsh to the ear, I ask ChatGTP to rewrite it - it's still my original thought.
It's really effective, but I tend to tone it down a bit to sound like myself since the output can be too formal, dry, and "academic".
At that time that was the only solution that would have made sense given the achieved behaviour, performance AND development effort.
I did spend a lot of time Cordova(PhoneGap) and all the other HTML5 app thingies for iOS at the time.
Not sure why that particular in my opinion pretty obvious choice bothers you that much. That is very much the reason why they didn't even bother releasing it for Android until almost two years later.
You have to change to Azure, because we are Microsoft partners and we have free credit.
The credit is not too much though and we have to spend the same money on useless trainings so we keep being partners.
4 core and 8GB should be plenty for your dev VMs, that’s the largest we can run on free MSDN accounts.
Have you tried the managed API gateway? Why not?
You should use managed caching on the edge.
We already have an on-prem SQL Server, use that to cut costs. Yeah it runs on Vmware and network storage.
Do we have support contracts for Nginx, Ubuntu? We have RHEL licenses, so you should use that.
Can we run this in our OpenShift cluster instead? To cut license cost it will be co-deployed with the developer envs of other product, but just set resource limits. Yeah, we only have NFS storage.
S3? Just use a PVC, it’s the same.
We decided on Datadog for unrelated product A and B, so you must use that. The license is expensive, so only log errors please.
We use Kafka for the workqueue, but limited to 2 cpu cores to make it cheap, so please make sure not to send too many notifications.
Python is not for production things, we will assign an offshore team to rewrite it in Spring Boot. We target Java 11, because our productivity increasing libraries are not yet updated.
Minor change needed for deploy: every service should build it’s own RPM package in it’s dedicated git repository.
You need to submit the architecture diagrams and service documents next week, thanks for the meeting.
What did I miss?
The security team is getting weird reports from their internal nmap scanner that is constantly scanning this network. No, they don't know which of their vendor automated scanners are doing it, no they can't turn it off, so please add code to specifically ignore those scans.
Our VMware cluster has plenty of compute left, but it is out of storage space because there are hundreds of VMs from other teams and nobody knows what is still being used, so you can't get any more dev instances until we get more storage added to the NAS (in the next quarter's budget).
Any unrelated director somehow ended up seeing this diagram and had some "suggestions". Please implement these changes asap.
this user gets it
From the top of my head...
* A bunch of opinionated devs pushing for microservices because they heard that's what cool kids do now. No technical reason behind at all, just pure bias flamed by a couple of blog posts/YouTube talks they watched one evening.
* A bunch of devops with intentions of implementing "industry best practices and modern tooling ©" that will end up creating a house of cards in Terraform that nobody, themselves included, will dare to touch in six months down the road.
Am I sounding too cynical?
Seven red lines, two with red ink, two with green ink and the rest – with transparent. One in the shape of a kitten.The thing is that with a small team, you need to focus your engineering efforts on things that add values, i.e. mostly functional improvements of your product. Everything else is a distraction. Anytime we hit a problem everything grinds to a halt until we fix the problem.
Non value adding activities like devops are things that we do on a need to have basis. I have used things like teraform in the past but I opted not to on this product. Reason: I'm not planning to take down my servers and recreate them. And when I do that anyway, it takes about half an hour. Not worth spending days/weeks automating. There's a reason companies employ full time devops engineers, it's actually a lot of work. I do most of the devops in my team. When I have time and when it adds value. Which is not often. I do it well and I keep it simple.
We deploy multiple times per day though; that's worth automating. Automating things you only do once just isn't. We might get around to getting some automation for setting up our infrastructure eventually. But it doesn't really solve a problem I have right now.
Like instagram, we can scale if we have to. We're in google cloud and we use a load balancer and a scaling group of simple vms that run the docker container. One nice thing in Google cloud is that you can create a vm and parametrize it with the docker container image and it will run it. No need to install anything. The default os on the vm is docker ready. So, that's one less thing for me to worry about: installing shit on vms and making sure that is updated. I simply build my docker containers, push them to the registry and tell the scaling group to apply a new instance template with the new container. It does the rolling restart. That's just 2 simple gcloud commands in a Github action.
Simple is essential. Easy to understand. Easy to fix when it misbehaves. Easy to explain to others. Simple ensures you don't waste time.
Never knew about this feature!
Look how X has diminished in quality as Elon started slashing team sizes.
Then when you design more features, security and other various systems to serve the customers it will creep in complexity. You can not escape that no matter what you do.
I generally agree with you on your comment. I am a casual user of X these days and the site seems to be humming along without any user facing engineering issues.
With his products such as Tesla, X and Starlink touching millions of people daily Elon is an easy to reach punching bag.
Even if he is somehow instrumental in solving the massive feat of putting humans safely on Mars there will always be people on the side lines having shots at him.
That's not a bot if you are not a bot? I know it's 'normal', but I still it as a bug. AI classifying me as something I not is a bug.
That's the most directly obvious thing that has happened post takeover.
There is a case to be made that Twitter would have been profitable if it didn't torch money on unnecessary complexity, but the crashing ad revenue suggests Musk is not the business genius that you might expect.
what happened?
And one massive security breach.
Most common thing people use it for is posting and reading. If you fake feeds as if they are real-time then the viewer will never know there was a system outage.
I am almost positive there was a massive security breach too.
That's a very serious allegation, as it would be illegal to not report it. Can you add some details?
We hit maximum reply lengths but yes there has been. And a massive security breach.
He had a point there though: He said that there seem to be "3 managers 'managing' one engineer", and I believe this is a common problem in the industry. VC-funded startups are terribly overstaffed and over-inflated.
A lot of tech companies have bloat in the form of AI ethics people, DEI people and so on. They need to go. But Musk probably hurt twitter a lot in short term by firing a lot of engineers and making it a place that made people unhappy.
These do actually have a proper job. They do ethics laundering for the tech companies and are very valuable.
Introduced failing ideas that made him ridiculous world wide, ... to finally hire a CEO.
And now he is using the platform to influence elections and events - free for all. Sure.
Not to mention HR managers.
I have seen situations where there are 10% engineers to 50% "assorted management" in tech companies. (the remaining 40% being a mix of sales and support staff such as office management).
When you cut the company to 1/3 and keep foreigners because of their visa status, nobody is gonna say anything of course, but that says a lot about you!
I do not believe that half the company was just "overstaff". I have been in situations where 1 manager had 1 reporter/reported, but they were single cases - it can't be spread to the entire company and nobody does anything.
That... you care about people regardless if they are foreigners and that you try to help those that would have the most problems if they were let go, especially as these problems are a consequence of your hiring of them?
I'm not familiar with the story, but from the way you presented it it sounds like a proper thing to do.
EDIT: not taking Musk's side, just pointing out the issue with the parent's argument.
Is it though? Small, lean, teams have fewer processes, less distractions, better communication, and more flexibility in what they can do. I've been in such teams and built such scalable systems and there was nothing 80-100 hours about it. It turned worse once the company was acquired and management and "specilised" workers were brought in.
Has it? I've had the exact same experience.
Ideologically or from a service perspective? I don't see any noticable drop in the latter. Some disruption is expected as huge parts of the team is fired/leaves, but I see it going on in business as usual otherwise.
The point is you can.
Choosing not to do is probably more important than being able to do.
Organizations are reflected in the products they create. We shape our teams and thereafter they shape us.
This doesn't mean brain teasers or other arbitrary metrics with standard bell curve distribution so you can pick the statistical outliers and claim you've done this. That's totally wrong because that's not what you're fitting.
Those are filters that produce stochastic results with merely the statistical properties of these rules of thumb.
If you're looking for a programmer, here's a better test: think if some famous programmer walked in and sat down to do your process. Could they pass? If the answer is "dice roll", meaning you'd say, turn down Rob Pike or Larry Wall, then you're doing it wrong.
As far as X, Musk is insane and drunk with power, that is not this.
True but different: even ignoring how many of them were customer support and moderation, once the tech stack gets complex, you can't just snap your fingers and act like it's a simpler stack.
Even if a fresh 3-good-graduates team can reach feature parity with your now-1000-person team, when they've had 6 months from `git init` and you've been at it since the new team were in Kindergarten, the only way for the big corp to do the same is to buy out the new team and then leave them alone.
efficiency shrinks by more team also you get other people. if you are working on something you stay alive until it's done.
≠ a job
100 million?
Meanwhile, all these orgs with essentially a CRUD app, with 1,000s of engineers..? That I never understood.
Scaling is not the only challenge engineers face, but somehow it's the one that is mostly praised.
They also need to quickly respond to downtime, because unlike IG if some of those CRUD apps go down in B2B world you are often losing customers actual money not just ad views
But anyway there was never a period where Instagram had a tiny 3 dev team and handled ads at the same time. 3 devs only worked back when there were no customers, no ads, no profits and no real responsibilities.
Still, between 3 devs and thousands, that's 3 orders of magnitude. Let's take the core team and 30x them for a broader project. Add 50 (?) for billing. Another 50 for web apps. Another 30 for general devops? That takes us to a "mere" 220 devs.
Then I read that Uber had around 2,000 SEs in 2020, and Airbnb at some point was 1,200. I don't know how to grok that.
You mention AWS, my impression was that even many very large orgs run their their hardware in the cloud (surely not all but still).
Dealing with payments globally to hosts/drivers and specifically compliance is complex in itself.
To get traction from users is the real challenge.
Unfortunately yes. Scaling is a problem I would love to have :D.
You've got a company paying off influential people in a space where people are looking for guidance, convincing them that they too have hard problems that cannot settle for simple solutions.
Selling the narrative that developers need to be all in on the most irrelevant aspects of building a product, and ignoring the fact that if you instead focus on building simple, easy to maintain software, the fact your LCP isn't hyper optimized by some newly invented mental model for app development won't matter: Google (or any search engine for that matter) will not ignore the fact people just actually want your content.
They do not care how great your web core vitals are if you waste a bunch of time bending over for some irrelevant bullshit problem instead of talking to users and iterating.
So not only are you wasting your time on all the wrong things, you're also gaining nearly no benefit from it
I'm pretty convinced that all these shiny new hotness tools and frameworks of the past decade are actually just meant to sabotage competition and small companies
They introduced the concept of a "server component": one that is rendered once and cannot update state.
And instead of making that opt-in, they made it the default.
—
That is to say, the default of React is to no longer allow updating state in components.
No value proposition is safe from the forces of financially incentivized thought leaders.
I'm not even going to get into how ridiculous it is that the default config chosen was to break people's builds for a feature that isn't mature enough to have anything more than a single reference implementation from Vercel.
I’m trying to find old software engineering gems and explain them as simply as possible, so I’m glad you found it simple to understand.
Also, it’s definitely possible to make a clone, but the hard part is getting the users :)
You are probably right. The Instagram Engineering blog [0] points to the Fabric documentation on Read the Docs [1] which is empty.
"Read the docs" links to the GitHub repo [2] but it's 404 and even the GitHub organization zwsyff888 is gone.
[0] https://instagram-engineering.com/what-powers-instagram-hund...
I've built quite a few projects using almost exactly the same stack over the last 15 years. Almost all are still running, those that aren't are for business reasons not technical.
The problem is, they're not exciting The Next Thing real-time javascript somethings, so a lot of devs won't want to use it.
> In July of this year, Meta launched its latest mobile app, Threads, a microblogging service and new rival to X, formerly Twitter. In the first five days following its launch, the app achieved 100M downloads – a new record for the company by some margin. Meta’s previous record for new app installs in the first 5 days after launch was 1M.
Built with a slightly larger team, considered an agile team: 3 product managers, 3 designers, about 60 engineers.
- monetization
- finance
- analytics
- ad placement
- ad bidding
- android client
- apple client
- browser client
Just off the top of my head in 2 mins that's a few of the extra concerns I can come up with...
I remember meeting someone who was responsible for writing an ultra performant JS WhatsApp client for firefox OS. When you have 2B users, the long tail is long...
OG Instagram was pre-monetisation. So yeah maybe a simple image hosting service needs fewer engineers but maybe a profitable business that can monetise the service needs more...
Iirc in Singapore, Netflix has a reasonably sized team just dealing with payments.
Presumably for a global business, payments is a pretty gnarly long tail of integrations (never worked with payments so pretty ignorant on the subject) not to mention I wonder whether they also are onboarding content distributors onto some kind of automated payment platform so they can very quickly churn content libraries or attribute some kind of viewing related payment...
I guess tldr doing business is complex and needs lots of engineers vs making a small and focused app.
Don't get me wrong, I think it's great that we live in a world where there's infrastructure lying around that makes it possible for a three-person team to achieve so much so quickly, and that team deserves a ton of credit for doing it. But it's not the "these companies are so bloated" lesson that some people are going to try to take away from it. That's just Tall Poppy Syndrome.
So they got to 30mm users with 6 engineers...scaling linearly with 5mm customers/engineer. Incredible.
[1] https://review.firstround.com/how-instagram-co-founder-mike-...
But when your design relies on many services to provide a wide variety of features you need to break out this design to allow teams to operate independently.
Mini monoliths are more popular today than traditional monoliths of the old.
Not necessarily, because if you scale only one of your services all the other services do not benefit at all.
Having microservices would only be better in that instance if they actually consume the resources they are given.
In my opinion it makes the backend way more resilient than a monolith.
Don’t kill me for this opinion please ;)
There'll always be critical microservices that keep your app running. It doesn't matter if all your other services are running if the one serving up core functionality goes down.
If your engineering rigor is so poor that you can't get reliable failovers with a monolith, god help you keeping microservices running.
Of course when you replace function calls with network calls, make everything asynchronous and eventually consistent, there is a lot of work to do to not end up with a less reliable system.
A monolith doesn't force a single process. IPC is still simpler and cheaper than network calls.
A monolith doesn't force never having an external service for a specialized use case, or FFI.
I specifically called out the extra complexity of network calls in microservices, not sure if you read the full comment.
> A monolith doesn't force a single process
I'm not convinced; if my small/specific code has it's own process, I would say it's a microservice. Sure, we can have replicas for redundancy, that doesn't mean I won't have reliability issues when my process is crashed.
> A monolith doesn't force a single database.
> A monolith doesn't force never having an external service for a specialized use case, or FFI
True, sadly it doesn't usually work this way. People take the path of least resistance.
Also once you add multiple DBs you start to get into eventual consistency; which is one of the harder parts of microservices.
Calling out networking doesn't preclude me from mentioning IPC. IPC isn't limited to network calls, it can be as simple as shared memory and hit millions of OPS: github.com/OpenHFT/Chronicle-Map
> I'm not convinced; if my small/specific code has its own process, I would say it's a microservice.
And you'd be wrong. A core tenant of microservices is being able to individually deploy your microservices. If I spin up a new process for some high risk, highly memory intensive process I've introduced a fraction of the operational complexity of a seperate server and retained the core value proposition of reducing its blast radius if things go south.
Of course again, if you're having so much trouble handle writing software that's reliable that you being to consider isolating instability as a top benefit from your IPC setup instead of a tiny value add... it might be a sign you're not ready for microservices.
_
> True, sadly it doesn't usually work this way. People take the path of least resistance.
> Also once you add multiple DBs you start to get into eventual consistency; which is one of the harder parts of microservices.
You're making my point: If you don't have the engineering chops as a team to make a robust monolith, you definitely don't have the skills and resources to start looking at microservices.
Eventual consistency is not inherent to having multiple databases. If I have an oft changing ephemeral set of data that only affects one feature and it's creating an impedance mismatch with our main datastore, nothing is stopping us from pulling in Redis for all the queries we were previously sending to Postgres, and as far as anything relying on that feature is concerned, nothing at all changed.
With even half decent engineering, Redis going down doesn't break any differently than it would have for a microservice: you define the same error boundaries as before and the failure case ends up the same.
I mean seriously, if your team can't handle having a second data store, imagine the bedlam when you're trying to handle multiple languages across multiple data sources in a non-centralized manner?
_
Microservices are a pattern for companies where a "microservice" gets the kind of development and devops support that would justify spinning off a new mid-sized enterprise.
When you're Netflix your `api/movies/[movieId]/subtitles` endpoint is serving the kind of traffic most companies will never see in their lifetime and needs optimizations that maybe 100 companies in the world will ever need.
For the rest of us EC2 has 224C/488T CPU 24,000 GB RAM machines with 38 GBPs I/O bandwidth. If your business ever scales so far that you outgrow that, throw some of that X Billion dollar valuation money at the problem and build your microservices.
You made the same point I made as though it was in contradiction to what I said. Adding a network call adds complexity, yes.
> A core tenant of microservices is being able to individually deploy your microservices.
And why would you not want this to be independently deployable?
> You're making my point: If you don't have the engineering chops as a team to make a robust monolith, you definitely don't have the skills and resources to start looking at microservices.
Firstly, you never made that point. Also, I never argued against it, in fact I agree completely.
> Microservices are a pattern for companies where a "microservice" gets the kind of development and devops support that would justify spinning off a new mid-sized enterprise.
Disagree, netflix has >1000 microservices.
Ah, wait it is.
> And why would you not want this to be independently deployable?
Because FAANG has more engineers devoted to managing deployment/observability/version skew/DX/scaling/security than you have engineers. Simplifying your needs in those realms helps you greatly.
And to top that off, it 100% can be independently deployable if it's a big enough separate concern: that's just SOA without the 90's XML/SOAP/RPC spin that was ESB: https://aws.amazon.com/compare/the-difference-between-soa-mi...
_
> Disagree, netflix has >1000 microservices.
That says exactly nothing. At Netflix scale their most random "trivial" endpoints are easily doing scale that entire SMEs won't ever deal with.
When FAANG is your case study in any technical discussion in a public forum, you're default wrong. I work at an AV company, I'm not about to start telling people the insane architecture we need to support ingesting petabytes of data is something that anyone else needs.
Any useful technical discussion needs to be grounded in what the 99% need, and microservices are not it.
Again, at no point did I make an argument for microservices.
> that's just SOA
Absolutely not, from your own reference:
> Each service provides a business capability.
Spinning a high memory task off into it's own process is not a business capability. Microservices are more granluar than SOA services, your describing a microservice.
> That says exactly nothing. At Netflix scale their most random "trivial" endpoints are easily doing scale that entire SMEs won't ever deal with.
You said microservices are for when a microservice would have the support equivelant to a medium enterprise, this is not true even at netflix scale. They absolutely have services owned by very small teams, or else they wouldn't have more than 1000.
> When FAANG is your case study in any technical discussion in a public forum, you're default wrong.
Well who do we use as a case study on microservices then?
> Any useful technical discussion needs...
A technical discussion requires nuance, not turning into a black and white one side versus the other.
Yes, you can have multiple DBs in a monolith, but you tend not to. In microservices you are basically forced to.
It's a crude and expensive way to force modularisation. However, that is still what it often achieves, it gives you infra that you can keep other people away from and lets you be in charge.
It's important to keep in mind that this isn't a technical problem. It's an organisational problem. In fact, there is no technical reason why monoliths would be an anti-pattern, which is likely why they are still being taught as though they weren't at many universities where professors still naively think that the MBA's aren't going to cost-cut IT at every opportunity even though their entire organisation is made up of employees who spend 100% of their working time on IT devices of some form. Similarily, Microservices, aren't really the "technical" response to this. It's how IT and digitalisation had to evolve to keep up with business demands and better generate value. The simpler and more decoupled you keep things, the better you'll be able to respond to business needs. Sure, there are a gazillion different ways to do Microservices wrong, and if you do it wrong, then you'll likely be in the same mess that you would be with a monolith, only so much worse, because now you have 9 million tiny monoliths and shared databases.
Luckily we still live in a world where everyone is somehow still OK with IT not working. We went to an appointment that isn't relevant the other day, and they had a tablet where you could register your license plate to avoid getting a parking ticket. It didn't work, so we talked with the receptionist who was like "yeah, it does that all the time, don't worry, if the systems are down then they can't give out tickets"... Fine for us, but think about that... It turned out the system was down in my entire city, which means that all those hundreds of employees who are out handing out tickets had nothing to do while their IT system was being fixed, hell, the entire company wasn't generating income for my city while their IT was down, and this was a regular occurrence? What my point with this is, is that you can do things really wrong, and still be a "successful" company, it's just that the companies who manage to generate value better (which is frankly always microservices of some form) tend to simply do better. But like I said. You can do "microservices" in a million different ways. Running two different django backends to handle different parts of Instagram could be considered having two microservices after all. The importance is how you deal with the needs of your organisaiton in a rapid fashion.
But for anything where there's healthy competition, this completely changes. Errors, bugs, conceptual problems, etc absolutely will have an extremely negative impact.
As an example I once worked for a company selling tickets online, but there were numerous bugs, and the system would often crash under load. Long story short, we lost many users to competitors, that company is no longer independent, and all that code is now legacy.
Compare with the monopoly situation of Ticketmaster, they are far worse than this company ever was, and are quite successful, with a large user base. That hates them ;-)
Same thing happens in microservices too. You just need good planning an organizing.
Twitter: just sharing a bunch of text.
Netflix: platform with static content, where sharing isn't even possible.
TikToķ: Instagram for video, so basically just bigger files.
WhatsApp: they also pulled it off with a small team.
If you talk about Roblox, now there's a challenge! And they pulled it off with way less engineers than Twitter.
I find it interesting that Meta choice largely this same tech stack for their newly created Threads service.
P.S. Here's more information: https://instagram-engineering.tumblr.com/post/10853187575/sh...