Meta’s Hyperscale Infrastructure: Overview and Insights
cacm.acm.org
cacm.acm.org
Kind of more impressive to have kept this ability to ship fast than anything else. A lot of work is needed to not let the bureaucracy increase and stop the lawyers or other functions from creating approval gates everywhere. Or at least to be able to have war rooms that can get it done so quickly.
Is it really a skill to very quickly release a dud app?
I don't know the answer to that. Bypassing bureaucracy seems like heaven, but it feels like it also bypassed the product folks entirely.
I don't know that it's a bad effort, but it's one that rose and died seemingly the same day.
I feel like more time with a good product person would have given more thought to fit, advertising, release, and so on.
Its fediverse integration stuff isn't panning out because nobody in the fediverse is stupid enough to let them federate. The single thread per message instead of hashtags thing doesn't seem to add a lot either.
purpose. value. actual users.
I'm talking specifically about its launch.
It's reasonable to assume they're worse. Bluesky doesn't have Facebook's network or surveillance apparatus. Neither does Twitter, except it's a higher-value target than Bluesky and Threads combined.
I hate meta with a passion, but I don't deny they have some great infrastructure and engineers to enable the bad things they do to the world
All form and no function.
threads as a product was DOA when that didn't work. you need a network of interesting important people for it to be useful. when the migration didn't happen, you ended up with a bunch of instagram meme influencers reposting their content across two apps instead
I think their strategy combined with an open offer exclusivity bonus could have given them the stickiness. up front 5k, 10k, 15k, etc to a twitter user that matches their follower of at least 25k, 50k, 75k, etc count on threads and agrees to exclusively post there for a year. people weren't getting paid on twitter so this would have been alluring
even if it was spun off to be stand alone product without Instagram it would still be worth billions
It's pretty cool to have a version of Twitter to communicate with the world, which isn't run by a literal sieg-heiling nazi
This comment is perplexing.
Reportedly Threads has 130 million monthly active users (and growing) vs. 550 million accessing X (and shrinking). This alone already places Threads as one of the top social media services in the world. It's projected revenue for 2026 is over $10B.
What exactly are you talking about?
https://news.ycombinator.com/item?id=43010781
You wrote a wall of text saying and refuting absolutely nothing.
The only remotely tangible argument you made was the blurb on “Specifically, as it pertains to monetization, we don’t expect Threads to be a meaningful driver of 2025 revenue at this time,”. Considering Meta reports revenue around $40B and Threads, considering the EU snag, was basically launched last year, this is far from being the failure you are trying to spin it.
But if you had any point worth making, and you didn't had an axe to grind, you wouldn't grasped at straws such as the "As for the wider positive effects to the society at large, feel free to point out any."
"our new feature reached 100m users in 5 days" sounds a little less impressive (especially given meta has multiple billions of users to start with).
What this leaves out is that the real customers, advertisers, already have a fully-functional hyperscale system creeping on users and pushing ads in their face. It's akin to crowing about how fast you can put up new billboard ads when the billboards are already built--someone else has already put in the power/network lines, poles, screens, routing, etc, and you just hooked up your existing ad feeds to it. Facebook doesn't give two shits if the billboards were in the middle of nowhere, as long as they can charge their advertisers some $$, they'd deliver seizure-inducing hallucinogenic drugs right into users' veins.
I don't think I've ever seen a viral Threads post that escaped the network?
Meanwhile Facebook screenshots still circulate as memes (derisive, but hey, the motto of the modern era is that the world hating you is better than them not noticing you). And Instagram is fairly entrenched as a marketing channel.
Is Threads the new Google+?
BlueSky feels a lot more adult.
Threads also has a huge oversupply of fragile mediocrities who love the drama and/or haven't worked out it's a terrible medium for self-promotion, and has generated some of the most jaw-droppingly stupid takes I've ever seen online.
(And I used to be on AOL...)
A good few people left Threads when the censorship started to become too obvious, and a lot more will leave when the ads become too heavy-handed.
I'm sure Meta has some internal reputation metric for Insta accounts, and it would have been trivial to give all accounts above a threshold immediate access to threads.
But nope, you have to create a threads account to view threads.
It’s certainly not a flop. It’s almost as big as X globally.
There’s a tremendous locality bias around Meta’s products. Social media use is very localized by geography and age. So if all your friends stopped using Facebook ten years ago, you assume that’s probably true everywhere, while in fact they’ve added several billion users on FB since then.
I've seen several interesting posts from Bluesky referenced elsewhere already, but literally nothing from Threads despite it having had more users for longer.
(She's 20 years old, her father is the richest man in the world, and he regularly does interviews saying his child is literally dead to him. Enough drama for an Orson Welles script and a second-tier social media service.)
It retains important bubbles like Silicon Valley VCs and media people, but globally it’s just not significant with its single-digit market share.
If it's got that level of usage, and not just people clicking on the occasional recommendation from Facebook/Instagram, it's remarkably insular.
At least that's what Meta tells advertisers and shareholders.
Zuck has a long history of fudging the numbers.
I honestly find high pressure work more relaxing than these kinds of meetings.
There's a famous poster from the Facebook days - don't mistake motion for progress. I always thought they should make one for re-orgs.
If it's tied to Instagram, and these users aren't using Instagram while they're on Threads, isn't it basically the same difference?
The secret of keeping deployment velocity high is having nothing important to lose. Facebook can drop or repeat a few posts and pics here and there and nobody will bother really. Anyway, who are they gonna call?
The moment you start dealing with transacted data that requires immediate consistency and guaranteed persistence then yeah, managers get nervous, lawyers get involved and there's no way for guerilla ops to hold up.
Nobody's paying for Meta's stuff except advertisers and something tells me that _their_ part of the pipeline is handled with _much_ greater care than the "consumer"-facing apps.
That said, I'm a young single guy. If I had a family and no lead time, this doesn't sound great.
Shout out to the HBase and ZippyDB teams! This is the first public acknowledgment that ZippyDB was converged upon.
It's also super cool to see the Developer Efficiency pushes called out. 10,000 Services pushed daily, or every commit is so impressive.
When I left FB, I couldn't find anything close. So, I'm building the infra that I was missing as a startup. Batteries Included. https://www.batteriesincl.com/ https://github.com/batteries-included/batteries-included/
Lots of companies are hosting old monoliths on Amazon Fargate, for example.
The flexibility of knowing that any machine can instantly run the code for your API gives a lot of flexibility to rapidly scale up an API.
Nothing is “serverless” to everyone. Especially when you run the data center. But being “serverless” and even sitting above the “language runtime” gives API developers a lot of freedom to focus on business logic.
By that definition, you could argue Amazon's detail page is "serverless".
CGI is also serverless
I wrote a couple articles with these comparisons, and related issues:
Comments on Scripting, CGI, and FastCGI https://www.oilshell.org/blog/2024/06/cgi.html
Comments on Shared Unix Hosting vs. the Cloud https://oils.pub/blog/2025/02/shared-hosting.html
---
I'm not the only one who thinks this: https://www.devever.net/~hl/mildlydynamic
AWS Lambda: CGI But It's Trendy. Recently we've seen the rise in popularity of AWS Lambda, a “functions as a service” provider. From my perspective this is literally a reinvention of CGI, except a) much more complicated for essentially the same functionality, b) with vendor lock-in, c) with a much more complex and bespoke deployment process which requires the use of special tools.
People who don't think this is true probably never used PHP or CGI ... You don't manage the server; you just upload your code!
Architecturally, it is the same. From the user experience, it is the same.
PHP and CGI also scale infinitely, because they are stateless. Meta scales PHP to a user base of half the world's population. They could also scale CGI/FastCGI.
If only which that were true, it is more like building our own digital prison.
Maybe the company in charge is bad in many ways but all of the things in the article are astounding to me.
I am not an engineer like many of you so maybe the article is old news to you guys but I couldn't help but say "wow".
I feel like if you took some old science fiction writers from back in the day and showed them this article, you would find sheer awe in their faces too.
I can't tell if this is satire. The science fiction writers from back in the day wrote about humans exploring other planets, building sentient machines, encountering aliens or maybe developing psychic powers. (looking at you two, Philip K. Dick and Theodore Sturgeon)
Building the world's biggest ad-serving infrastructure might impress them. It impresses me! But "sheer awe" seems a few orders of magnitude off the mark.
To be in awe of software instructure but not espouse the same feelings about say, biology, is sad.
This approach appears to be one of the most optimal designs for software networking (service mesh) and for storage (database operations) for organizations with large server counts. I was surprised to see their IP networking to follow the same model, rather than primarily relying on BGP.
It was omitted in this paper, but I would expect for local caching to be used to reduce load on L7 routers and for improved latency for database queries. Clients can invalidate caches and perform another lookup to the service mesh after a reasonable timeout (100-500ms).
Using Erlang to write serverless functions seems like avoiding all the huge benefits BEAM can offer.
>Additionally, product engineers predominantly write code in stateless, serverless functions in PHP, Python, and Erlang for their benefits in simplicity, productivity, and iteration speed.
>To boost developer productivity, Meta has adopted continuous deployment universally and enabled more developers to write serverless functions rather than traditional service code.
Ten or so years ago I remember going to a larger technical recruiting pitch at Facebook where they discussed their logging complex (I have only a vague recollection of the details). Honestly it was one of the most beefy implementations of system/application logging I remember having seen at the time.
I have been very disappointed since moving from FB to Google at how much worse the data and measurement platform is here.
I'm reading something like this for the first time. Is this common across industry or only typical to Meta? In contrast, we use multiple instances for sub-components of our ML training pipeline.
People in my industry have a tendency to over-optimize (technical folks love nothing more than to tinker), and in doing so create very unique deployments per group that require a significant amount of infra and operational support, and drastically slow down the rate of progress/change (not to mention the cost). When you peel back all the requirements it turns out we really only need three unique deployment options. Makes it all significantly less cumbersome.
Say I want a 1MB image, wouldn't it be faster to serve me the 1MB image over a slow connection with 100ms latency, than going through multiple hops of increasing latency, with multiple round trips?
Say I request the image directly:
me -- 100ms --> datacenter
datacenter -- 100ms --> me
Say I now go through Meta's system, assuming that goes to the same Datacenter in the end, and there's no FTL tech:
me -- 10ms --> CDN
CDN -- 10ms --> PoP
PoP -- 90ms --> datacenter
datacenter -- 90ms --> PoP
PoP -- 10ms --> CDN
CDN -- 10ms --> me
Consider also that the CDN can fetch the whole image async after the first request, but the nature of TCP (that you'll be using to fetch objects) is serial, meaning you can only fetch about 1k each round-trip.
Now, in reality the situation is way more complex, because various hacks have been added over the years, for example your browser might grab the size header and just create a bunch of download threads for your image with various offsets, the TCP scaling window might come into play; but largely most people set 1500b for their MTU on consumer hardware, and there's some overhead from TCP/IP and so on.
CDN <-> PoP <-> Datacenter communication doesn't require initiating connections. They reuse them because they centralize requests to serve different users (this is even explicitly called out in the article).
Lets say your closest datacenter is 50ms away, and the closest PoP/CDN node is 10ms. Just initiating the image download through HTTPS all of sudden is 500ms vs 100ms.
Sure, PoP/CDN might need to go to the datacenter to fetch the image (and only if the content is not cached there already) but that only happens once before it gets cached, and there's still a lot of ms to use on that to make the tradeoff worth it.
You don't need FTL tech. You just need good private fiber. Geographic distance is hardly ever the limiting factor on the public internet; yet it often is for private networks.
Typically, the CDN nodes will be within your ISPs network, typically on the way to the PoP (or at the PoP). You're unlikely to see a significantly different round trip time if your request is proxied through the CDN vs going to the remote datacenter directly. But, even if your total round trip time is a little longer (10ms in your hypothetical), you get benefits from local TCP and/or TLS termination.
At ~ 100ms latency, your effective bandwidth is usually limited more by round trip than actual connection bandwidth, because of congestion control / slow start. A 1 MB image is going to be somewhere around 700 packets with ~ 1500 MTU. Assuming standard congestion control, the initial congestion window is 10 packets, for each packet acked, the server can send two packets. Assuming you and the server and the path between have infinite bandwidth and a large enough receive window, the client wound receive 10 packets at t=100ms after the initial request, then 20 additional (total 30) at t=200ms, ... 320 additional (630 total) at t=600ms, and the remainder at t=700ms. If your effective bandwidth is less than about 40 Mbps, you're going to hit congestion in the 6th bunch of packets, but any connection speed above that and your 1 MB transfer is going to take 700 ms. If you've got a much smaller MTU, you might need more round trips; and if you've got a larger MTU, you could end up with less round trips, but word on the street is inter-AS jumbo packets is rare, and Path MTU Detection isn't great, so lots of servers force an effective max MTU of 1500 or less, because it's easier to send smaller packets to everyone than to fix path mtu issues.
But if we have your steps of 10 ms to the CDN, 10ms to the PoP, 90 ms to the datacenter, and let's assume the transit time between you and the CDN and you and the PoP is symmetric, we get
Your request starts at t=0, CDN to PoP request starts at t=5 ms, PoP to DC request starts at t=10ms, then the PoP to DC request takes 7 round trips = 630 ms for all data, and you get that 10ms later at 640 ms. Your first byte of response is delayed by 10 ms, but it's worth while because time to last byte decreases by 60 ms. If the transit times between hops are asymmetric, the times that the client gets data don't change, but the math makes my head spin more.
If you change it up, and say you're 50 ms round trip to the CDN, which is 50 ms round trip to the PoP, which is 50 ms to the DC, first byte jumps to t=150 ms, but time to last byte would be 7 * 50 ms + 100 ms = 450 ms.
And, if the PoP -> DC connection happens over a warm socket, with an appropriate congestion window and receive window, the transfer from the DC to the PoP can happen much faster. Certainly, you could potentially have a warm connection to the data center from the client as well, but most services don't want to have millions or billions of client connections to all their datacenters, so running them through PoPs or CDNs can be pretty handy.
There's some handwaving here (I ignored processing time at each hop, but it's usually pretty low), but it's really helpful to process congestion control closer to the user. Any lost packets can be resent sooner for much quicker recovery if managed locally as well. And for things that are cachable, read-through caching with a CDN -> PoP -> Datacenter approach makes a lot of sense to reduce demand on the Datacenter, and benefiting from likely locality of reference --- people in the same area / on the same ISP are likely to fetch images that others in their area have fetched.
I almost wonder if this is preparation for them launching their own public cloud. Anyone from Meta care to comment?
If Meta can pull off the public cloud correctly, I’d trust them greatly - they’ve shown significant engineering and product competence till now, even if they could use some more consistent and stable UI.
To the side I have had GCP work but it’s been isolated and as if I were moonlighting.
Anyway in that context, you don't need trust when you have a sufficiently large legal department
I mean, Netflix and Apple even use public clouds quite heavily.
Their infrastructure is cloudy, but it's built around mostly a single customer and assumes the infrastructure software people and the application software people communicate deeply and continously. Running on a public cloud isn't that similar, at least as a small customer.
Could they pivot towards being a cloud service? Probably, but they'd need to do a lot of work to make their platform viable and to earn trust of potential customers, and they'd be entering a crowded market; there's already 6 S&P100 companies in Cloud (Amazon, Google, Microsoft, Oracle, IBM, Salesforce), and tons of smaller players.
IMHO, given their revenues and profit margins, there's no reason to do all the work it would take to offer cloud services too. Unless there's some opportunistic large customer deal made. They also might also need to renegotiate their content node agreements if they use them to serve cloud customer traffic, and that's a long process.
I get this vibe whenever I use the AWS or GCS dashboards, yet here we are!
Think of it as the difference between OCI and AWS.
Meta would be unlikely to launch a public cloud uncompetitive with Amazon's feature set.
The internal tools heavily depend on other internal tools, and none of them were written with customers other than Facebook in mind.
The infra is not isolated enough to be anywhere near ready for public offerings.
Sure there is process isolation, and its hard to break out of the jail/container/whatever you want to call the unit of execution. But its just not ready for public consumption.
The methods for monitoring, creating, deploying and scaling the "service(s)" are just too intertwined with internal access. Whilst there is sorta fine grained control on which other services you have access to, its nowhere like AWS et al.
the _other_ thing is that everything needs to be compiled to fit the platform. Its not run in VMs, its bare metal processes, with some rather fancy shims to isolate away libc (don't ask me more, I know that its there and some of the reasons why, but the mechanics are a mystery to me).
That platform that you compile to is a movingish target.
So no, meta isn't going to host thirdparty stuff, mainly because meta doesn't really have enough capacity for what it want to do now, let alone add more consumers.
IIRC, graphql is a means of papering over a bunch of legacy APIs. They removed foreign keys from mysql using it as a column store db, a vestige of the original LAMP stack still on PHP.
I don't think Meta infrastructural choices are applicable to most folk.
What does serverless land your average dev? A high AWS bill. Elastic managed Kubernetes stack? A higher bill.
Did you know that you can use YAML and provision actual cloud provider resources with boring tech? Welcome to Ansible. There is no need to recreate Linux network stack when you have the Linux network stack, and it actually works!
Quite a lot of hacky gak is required when you run node.js as a production public facing web service. A statically compiled binary won't invent novel code execution paths 4 days into a memory leaking runtime bender.
Boring tech is boring, I guess, even if it's new and shiny. Facebook creates tech to mitigate the pathologies their past continuously present.
Remember when they hacked a running Android Dalvik machine because their organizational constraints were such that they could never remove code or delete unused classes?
https://engineering.fb.com/2013/03/04/android/under-the-hood...
Facebook seems like a place where they do amazing engineering to temporarily stave off the disastrous consequences of their previous feat of amazing engineering.
Anyone using Ansible for cloud infrastructure management is not to be taken seriously. It's among the worst tools for the job - not (always) idempotent, no state tracking, slow, very limited in the resources it can manage, very lacking templating, fun stuff like "state: absent", running, and then having to remove the corresponding lines to delete, etc etc. You're literally better of bash scripting the cloud provider's CLI than using Ansible. Terraform/OpenTofu, Pulumi/tfcdk if you hate your future self are just clearly so much better.
I was making a point about provisioning VPSes instead of trying to wrangle postgresql restores inside kubectl or equivalent, of how your cloud provider is already provisioning a single physical server via a hypervisor.
I was making a point that facecrook overengineering is about them being boxed into corners, about how very little of big techs solutions translate to real world usage in the web industry i am very much taken seriously in for over 30 years.
You read 'ansible recommended', which I could also argue with you about, but I shan't.
You made a clueless comment which I tired to (constructively) dismiss. If you don't want to be treated as clueless, don't recommend the equivalent of using a hammer to peel potatoes.
Your comment feels like it's not actually engaging with the contents of this article. It's not that Meta is creating bespoke technologies only out of fear of breaking past code. Their entire methodology of innovation is highly iterative and grounded in feedback through practical demonstration of results. You say that "Facebook creates tech to mitigate the pathologies their past continuously present", as if that is a bad approach, but considering Meta's success, I think it would be wise to seriously reconsider that position.
This isn't specific to Meta/Facebook in any way. A very large percentage of big tech companies use MySQL without foreign key constraints, because they're problematic at scale.
To be clear, foreign key constraints are not used, but it's still very much a relational use-case / workload.
> using it as a column store db
That's nonsensical. Column stores are for analytics, whereas MySQL is used for OLTP. Both InnoDB and MyRocks are row-oriented storage engines.
Maybe you actually meant "key-value store", which is a more common claim, but that's still completely wrong. The query pattern for Facebook's social graph relies heavily on range scans over indexes, which isn't a concept supported by key-value stores. And there are many MySQL workloads at Facebook outside of the social graph database tier, with extremely varied use of MySQL functionality.
There's also a good paper on the approaches to configuration management here: https://research.facebook.com/publications/holistic-configur....
https://blog.cloudflare.com/how-we-use-hashicorp-nomad/
https://blog.cloudflare.com/cloudflare-deployment-in-guam/
https://blog.cloudflare.com/behind-the-scenes-with-stream-li...
Get off Meta. Stay off Meta. Don't work for them. To integrate them. Don't share links to them.
> Planetary
> What's the math behind?
Seriously though, this is painful to watch. Reminds me of those articles/videos where people have to have their jobs "invest" or software/hardware partners in that website and all of a sudden this person is a "Person of Interest".
Meh
> software/hardware partners in that website
> all of a sudden this person is a "Person of Interest"
Would you mind rephrasing these points? It's unclear what they mean or are referring to (e.g. what does it mean when a job "invests"?)