Why WhatsApp Only Needs 50 Engineers for Its 900M Users (2015)
wired.com
wired.com
It's scope that increases headcount, not users.
WhatsApp could scale to 900M users because the app they build didn't do a lot. It was for sharing plaintext messages between groups of people. That's quite hard at scale but you can solve it with 50 people as they demonstrated. When they started adding more features, that's when they needed more people.
If you keep the scope of your app small you don't need many people to build it.
I guess that theoretically it's possible to have a productive very large organisation but it's really hard to do that compared with the simplicity of a smaller team.
Fred Brooks - The Mythical Man-Month (1975)
Architecture increases headcount.
Bad quality of product and operations increase headcount.
Skill level of employees increase headcount.
It’s very similar to a proper algorithm cs some naive way of solving things. You can throw a few million on hardware to it, and still only get mediocre results, or solve something in a way that actually works
You’d think that teams closer to the companies core would be more competent and higher value… but often you find groups of thirty pushing inconsequential metrics. Alternately, you’ll find small groups on the fringe that appear to be doing nothing, but occasionally spit out the next huge product.
For these reasons, investors have historically penalized large conglomerates for inefficiency. The growth of big tech through the ‘10s may have been due in part to negative interest rates which valued discounted future free cash flow at approximately infinite dollars.
In a high interest rate environment, one would expect cash flow to be redirected to investors who will then invest the money. If big tech executive pay drops, then there is limited reason for tech executives not to move into startups.
Enterprise sales, personalized support, marketing programs, etc. Oh, and you probably have a legal team now. And you need to enable those sales reps and manage them. Oh, there's HR now too. And a docs team becomes more important...
That kind of transformation only comes from scope or the things the GP pointed.
It adds up. Sure they probably have fat to trim, but twitter gives us a great case study how often the massive headcount is focused on revenue generation activities the user isn’t privy to.
Does it? Twitter the product has been fine (barring clumsy design changes driven by Musk) after the massive downsizing, and they still make billions. I was under the impression a lot of the lost revenue was because of Musk's caustic image that made brands pull their campaigns from the platform.
Long story short they were bought by some holding company that completely neglected the platform and stopped any kind of moderation, DMCA notice or anything else. The result? It became a huge magnet for trading child porn.
While listening to this podcast I thought of how Elon gutted twitter to the bone and probably fired entire teams designed to deal with this kind of content…
Not that companies never have bloat, but it's obvious from your comment you don't know what DropBox actually does these days. Have you seen the admin pane for enterprise users? There is a huge amount there alone that goes into permissions and role management, auditing, compliance requirements, etc.
One simple example - how many engineers do you think it took to get all the compliance frameworks (e.g. SOC II, HIPAA, etc.) implemented and maintained? My guess is that that is one of a hundred requirements you didn't even consider.
I don't think any sane person would argue that DropBox going into the enterprise was a bad business decision.
In the small company I may something like "flip this windows setting and reboot the server". They do, we test it, and the call is done.
In enterprise the different levels of complexity spiral out of control very quickly. Is it the standard windows setting issue, or is it the enterprise endpoint protection causing the problem? Ok, lets follow the windows setting issue. Do we have access control the windows settings on this machine? No, we'll need admin access. Actually its controlled by a GPO. Now we need to raise a ticket with the active directory team and may need a custom change. Custom changes need a compliance and security officer to review. Because this change also requires a reboot a impact report needs filed before implementation and reboot. Do we have to test this in the customers Dev/UAT/QA system before implementation? And if this is a new issue will I have to write up the issue for the documentation team?
And after that I'll have spent 10 hours of on the call time with the customer, mostly waiting on them. And next year when sales comes to them with an even higher price tag for the same product they pay the bill because they know when they call we'll stay with them until the problem is solved.
This is obviously untrue if you want to offer and kind of support at all.
I also strongly doubt it's true you can scale up your application arbitrarily without needing more people, though I'll readily believe that it's a limited increase compared to what will happen if you have 20 product managers and designers trying to find ways to justify their salaries.
There's traditional web search.
Maps.
Images.
Video (which is essentially YouTube integrated into Google Search).
Each of these products would be Fortune 250 companies on their own.
Obviously, it's all powered by the Ads team / business.
But that's two majorly different parts - Google's 1P ads (including video), and the display network (3P ads).
None of this is useful without a gigantic analytics org.
And that's just "search".
Cloud & G-suite would each also be Fortune 100 businesses.
maybe years ago, this is not the case now.
100% And adding a way to accept payments (which WhatsApp never really did) and have functionality for businesses is an increase in scope. So, 50 engineers is terribly understaffed for an economically sustainable business.
For example, Facebook, Twitter, Uber, etc. have hundreds of engineers who literally just work on how the company gets paid. And no, this isn't just "well they can just insert Stripe.js into their app" - it's way more complicated than that.
if they hadn't sold to FB, FB or another company would have run them out of business by doing the same loss-leading advertising and subsidies for their own product that they did for WA.
free is not a guaranteed winner if you have three trillion dollar companies paying people to use their competitive products and they control the platforms that your app runs on.
People want to share videos and now even uncompressed photos. Considering Facebook paid a lot of $ for the service they can't let it slide into obscurity.
It lead to performance issues at almost all of our large customers.....
See, before that point customers would just export the XML statistics and generate their own reports, but now with the reports being good, they pushed report generation back to our product. Report generation had always been coupled with another service before that point so a team had to be split off to separate these services and API controllers so they could run standalone. This required a fast feedback loop with the customers to ensure we understood their storage requirements and the ability to access secondary SQL datastores instead of the primary. Now we have a much more scalable, but far more complicated product so we had to train the support team handle these issues.
Fallacy: https://wikipedia.org/wiki/Sunk_cost
It is hard to learn to avoid biased thinking in our own lives! I think that cutting one’s losses is a very important skill to learn.
This is 100% true. Where scope creeps in is someone realizes that adding feature X will increase usage by enough that it's ROI-positive, even with a team of 10 supporting it.
But then, he’s built 4 mega yachts [1] and spent $400 million+ on mansions [2].
1- https://www.superyachtfan.com/yacht/mogambo/owner/
2- https://www.dirt.com/gallery/moguls/tech/jan-koum-house-beve...
I don't mean this in a disrespectful manner - building and selling WhatsApp was an astounding feat.
I'd definitely want to pursue some hobbies, but I wouldn't expect any of them to make money or change the world. I could at least donate to nonprofits for the latter, though.
Jeepers. I am guessing your other priorities in your life make it so you need to live in an unsafe area?
I am middle aged, and I have never lived anywhere in my hometown Christchurch that had an active security system. In more that one place the front door or back door was never locked. Probably because I just have never owned much worth stealing! Correction: my parents had a security system, but they stopped using it and it stopped working maybe 15 years ago?
I currently live in a particularly safe little area of my city, although I wouldn’t mind having a secret lockup for the few things that can’t be insured. Most rural or small-town New Zealand is even safer than living in my city.
Above said, I do believe thievery is increasing in Christchurch. My friends in Woolston had a small shed ram-raided (failed because the shed builder went overboard on big posts), and those friends have needed to add security cameras to help prevent thieves casing their home. Anecdotally theft in Lyttelton is on the rise.
It may not be so much you live in an area with danger, but the danger comes to your safe area looking for money (ransom).
And not just you, but those around you:
* https://en.wikipedia.org/wiki/J._Paul_Getty#Kidnapping_of_gr...
See also "Lottery winner 'regretted scooping £11m' when brother hired hitman to assassinate him":
* https://www.dailystar.co.uk/news/weird-news/lottery-winner-r...
These are outrageously unlikely in my country, and would guess are extremely unlikely even in the USA.
If you are worried about extreme events, then a security system is hardly going to help you. When you win the lottery, you can move or buy a good security system.
The best example I've seen of this is chamillionaire. He had a string of 3-5 hits but diversified and has stayed wealthy while many of the people in a similar vein are broke.
You could say "he does what he likes", but it doesn't sound smart nor grounded to me.
Not that it makes any financial sense whatsoever, but the scale of wealth is unfathomable. Money no longer makes sense.
Money is power and power comes with moral responsibility. I think that's why the Gates' started that foundation.
For many people it's easier to feel good about having so much if you can claim that you're using it for good.
I don't know if I agree with that. Are they actually qualified to be wielding that much power? Maybe it's better to just let them spend it on personal luxury or passive philanthropy than to have them start twiddling society's levers just because they have a strong urge to "do something".
Whatever you do with so much dough is gonna have consequences, there's no staying neutral.
Think of it as make-work income for inner city shipyard workers, metal workers, and countless other legacy professions.
At the end of the day, the guy doesn't have the money any more, they all do. So they're happy, and he has a boat.
Note that I'm not saying Gates is a good guy and doesn't use his foundation to exert influence. But, saying he started a foundation to transfer assets while dodging estate tax reeks of "they just write it off...". Wouldn't he have been better off handing over that $66 billion to his children and paying 18% to 40% as taxes?
1- https://www.gatesfoundation.org/about/foundation-fact-sheet
In a world where mere 10s of dollars can fund a medical intervention that change s the life of a poor child, having 4 megayachts is obscene.
[1] http://highscalability.com/blog/2014/2/26/the-whatsapp-archi...
Jan Koum was more concerned about getting his shiny BMW dinged and more famous for badly parking the said BMW. (just kidding, Jan, please don't bash my Prius!)
Both were very good engineers who managed to create something amazing with a laser-like focus on the product.
To address a couple things elsewhere in the thread. WhatsApp circa acquisition was not text only messaging. We had multimedia messaging (video, images, audio), I think since before my time. I don't remember when end to end encryption happened (we were working on it for a long time, and it went through a lot of not user visible testing before we announced it and started turning off plain text messaging), but before that, the multimedia servers also did transcoding; post encryption, we had to do transcoding client side. If you think back to the variety of client platforms supported at the time, video codec support was all over the place.
We did have infrastructure to receive payments (apple and google in app payments and paypal), although requiring payment was very selective. I don't think there's published details on that, so I won't go into specifics, but you can't really require payment in places where $1/year is a significant burden or where it's difficult to pay a US based country; and you don't want to require payment in places where people will nope out and use any of the many alternatives. Otoh, my spouse paid for the 5? year plan just from prompting, without even asking me if she could get it free, and she usually doesn't do phone base payments without asking me for help.
On profit and loss, I understand why GAAP include stock based compensation, but it's weird to say there was a giant loss because of it. From what I've seen in public numbers, revenue was a bit more than expenses, and that's what we were told internally as well. After acquisition, I had much less visibility into the company financials; I did see daily verification expenses as part of my job, and sometimes saw our SoftLayer bills, but not our headcount expenses or any accounting for server resources in Facebook Infrastructure.
Real time voice calls launched
As to why things went right.
Limited scope and clear product vision helped. You could answer almost any design question by opening the platform SMS/MMS app; our app should look as much like that as possible, because that's what users know and expect. Dedicated platform teams buidling in the platform SDK was the right choice for that, although it does leave the real issue where it's hard to transfer your user data between platforms because each platform designed their own local storage databases.
Largely experienced workforce, with great autonomy and responsibility. On the server side, different services were mostly independent and often managed with a team of one or two; we'd do some amount of cross-training, but while I was on vacation, nobody did verification server development, just emergency fixes; and similar for other teams of one. But since I didn't need to coordinate with others, I could push verification server changes multiple times a day as needed (sometimes a few times a day). Chat always had the largest team, if nothing else, mostly all the separate services had to do something with chat, too. Although originally things were very separate, running things some things through chat made state management and authentication simpler.
It greatly helped that Erlang is the right fit for a chat service, and that Jan and Brian choice ejabberd to start with when people were using the early WhatsApp as a cludge to chat (originally, WhatsApp was just a short public text 'status' you could see for your contacts, they pivoted to chat later). About half the early server team had used FreeBSD at Yahoo, and zero had any experience with Erlang.
It helped that the SoftLayer servers (mostly SuperMicro, although some Lenovo post IBM acquisition) were very stable. This fed into the stability of FreeBSD and the operability of Erlang. Whenever we shut down chat servers, we'd find some clients with chat connections open for 45 days (mostly Nokia S60, which had a stable networking stack, but no push services, so we had to stay connected). Ocassionally, we'd need to do BEAM updates or FreeBSD kernel updates to address issues, but mostly servers were running for months or years uninterrupted. This is only possible with quality software and hardware. Both FreeBSD and BEAM/Erlang/OTP are quite approachable for local patching as well. They don't have a lot of churn, so patches don't need a lot of changes between releases, and things are well organized. We didn't have a ton of patches, but we did run things towards the limits. Not quite the same limits that Netflix explores though. We never did more than 2x10G at Softlayer, but chat didn't need that much bandwidth (ran out of CPU first), and for the most part MMS would be real close to disk bandwidth limits before network limits; and more MMS servers gave us larger storage capacity as well as more network and more CPU (TLS isn't free). Edit: I've heard from SoftLayer that by not using most of their services (including their load balancers, ugh), we stayed in the sweet spot of stability; we did have issues with LAN stability from time to time though; I used to joke that we were their network monitoring team, but they did beef things up there towards the end of our time at SoftLayer, I started getting their incident alerts before we noticed and reported problems.
We didn't use any sort of service orchestration until Facebook. I'm old and grumpy now, but I hate all these layers of stuff that hides things. We never needed to split a physical server into multiple jobs, so just running FreeBSD on bare metal was good enough. Erlang's dist and pg2 with some augmentation here and there worked for finding the current active servers.
Another thing is WhatsApp fits an offline first model quite well. Chats and MMS queue easily, so does contact synchronization and status updates. Twitter and Facebook feeds where it's loaded interactively and has to (more or less) find your contacts and then get all their recent posts and then sort them with a short deadline is hard. We didn't really have to solve a problem like that, because those processes were background processes and didn't have the same response time needs (some other design decisions also make it easier). Pushing all the state to the client makes things easier for the server, IMHO it also makes the user experience better. And we didn't need to store things forever on the server, which reduces cost.
Gosh, that was a wall of text. Hope it's useful.
Edit: oh yeah, we all enjoyed being under the radar. There was a lot of distraction at Yahoo and later at Facebook being constantly in the news for this or that thing that was out of engineering's control for the most part. I rembember most of the US coverage of the merger announcement having to explain what WhatsApp even was.
toast0, if you ever wrote a book about your experience & lessons learned in how to build a great product from the technical side - I'd definitely buy it for myself and gift it to friends. Just say'n :)
And it would be payback for the anxiety I got when I heard that Terry Gross was going to interview me later in the day on Fresh Air and I wws totally unprepared. Luckily, it was already recorded and they played a clip and it wasn't me. ;)
For user facing product design my advice is simple: think about what I would do... and then do something else ;) but I'll accept my advice/thoughts for user invisible stuff could be useful.
> About half the early server team had used FreeBSD at Yahoo, and zero had any experience with Erlang.
Can you say anything more about the choice of Erlang and experience with it? It seems pretty remarkable that you built such a successful business / service with no experience of Erlang.
Erlang is basically a perfect fit for a chat server. ejabberd is out there and there's several high profile use cases where someone needed a chat server, picked ejabberd and became smitten with Erlang. Facebook did it before WhatsApp, but they rebuilt chat in C++ because they couldn't hire Erlang people? and the two people working on it wanted to have some time to do other things. Riot Games did it more recently.
The nice thing about Erlang is it's really easy to make changes to your system; you don't need to move traffic and restart a server, you can just hotload the code changes. So, they started with ejabberd, and by the time I joined it was completely different.
Erlang also has tremendous observability. Each Erlang process has detailed information available (heap size, message queue length, reduction count (cpu use, kind of), you can view the message queue if you want, etc). The BEAM is also observable, it's even got DTrace hooks. There's a way to build where you get stats on lock contention (lock counting), which pretty quickly helps you find the root of scaling issues. Once you find lock contention, almost always the answer is find a way to distribute the work over more locks so there's less contention; same thing you do when a process's message queue backs up.
It is a small community, so if you're blessed with needing to scale, you are a bit on your own, but you probably could reach out to Erlang Solutions, if you get really stuck. Erlang will scale much farther than the community thinks though. WhatsApp routinely ran mnesia databases with way more data than people thought possible (just don't use disc_only_copies for anything other than the schema table) and dist clusters way larger than people thought possible (pg2 and global locks in general got iffy at times, but I think the new pg is better, and we didn't use global locks except through pg2). There is a general trend where there's no explicit limit on things, but the system will not work well if you go over the implicit limit, and you won't have much guidance.
I hear WhatsApp now uses proper Erlang application packaging, wkth releases and relups, etc. But when I was there, we just did hotloading and it was good enough.
I dont know why but I cant stop laughing when I read this line.
I'm not sure about that. These framework projects sound like internal projects developed to meet internal needs and improve internal processes. The GraphQL concept and use case in particular, as well as API management tools, is something I've seen reinvented internally by companies to be able to simplify processes. You just hear about GraphQL because Facebook decided to release it publicly. Everyone else has been more or less reinventing the wheel to meet the same needs.
Hiring an engineer has a significant opportunity cost.
You're right that in the short term you aim to hire where the return is highest, but in the long term that means they're right that you'd want to keep hiring as long as you get a positive return.
You want to spend it on whatever has the highest return, which will frequently be on an expense other than hiring, and when it is on hiring, it will frequently be on hiring roles other than engineers.
There isn't really an opportunity cost to hiring more people when money is no object. More likely to be the opposite.
Why do you think they increased their headcount so much in the last few years? Simply because they could.
The fact there is finite budget and headcount doesn't necessarily mean anything. If you're limited by the number of quality employees you can recruit and onboard, budget and headcount are effectively meaningless because they're not limited by money.
How to apply it wisely is a different, much wider topic.
The question really is, how can you increase shareholder value? If the answer requires engineers then they are hired. They are paid less then the value they add. When that stops being true people like Elon Musk can increase value by reducing headcount.
That’s not the part that gives insight to the question “how can it take X engineers to run service Y?” The insight is that employees can add more value than they cost even if they aren’t working on the most visible core product that the company is known for.
> a firm maximizes profit by producing that quantity of output where marginal revenue equals marginal costs.
This is a pretty fundamental result in microeconomics. The profit maximization section of the wiki article explains it quite well: https://en.m.wikipedia.org/wiki/Marginal_product_of_labor
The usual disclaimers about underlying assumptions apply of course.
Adding yet another zero to your user headcount is costly and a true engineering challenge.
Also, organizational inefficiencies are very costly and require more and more people to fill more and more layers of cintributoand managers alike. But mostly managers.
Edit: Third, product space exploration should not be forgotten.
And demonstrably 50 engineers was sufficient for handling stuff at scale.
but you could add another 0 to the message count. You could add another 0 to the image count.
That seems highly unlikely unless they launched a new product, or this was a gradual shift over many years.
Matrix as a protocol provides the building blocks for both messengers and collaboration tools (and more exotic stuff like spatial collaboration apps). Element is an example of a collaboration app (although Element X is being built to feel more like a messenger which happens to also support collaboration use cases - more like Telegram, but encrypted and open and decentralised). FluffyChat is an example of more of a messenger use case on Matrix. Thirdroom.io is an example of virtual worlds on Matrix.
So: I’d say that WA and Discord are very different products. But you could build both on the same underlying primitives (eg a stack of IP, HTTP, TLS, Matrix)
I remember the first time I saw the Facebook Ads platform, the Events Manager etc - it's a highly complex system that changes all the time. Many more engineers are working on that than on maintaining the ability to Reply to a post feature.
WhatsApp, until Meta, had none of these things.
The bit that you get wrong is assuming that Facebook or Twitter are simple platforms, and that their main and most complex features is allowing users to post, search and read messages. This is entirely wrong as these platforms are actually ad-selling platforms and thus include region-specific services, localized ranking, all sorts of data harvesting, etc etc.
Also, as companies grow they start to feel the need to implement efficient development processes, and thus start to allocate resources to stuff like putting together their own build farms, CI/CD infrastructure, deployment auditing, etc etc etc. They also develop things like front-end and back-end frameworks. It's not like a random team at Facebook just signs up for gitlab and does all it's business from there.
They don't need to release the actual code, but releasing "these are the inputs we use, these are the high level functions of the pipeline, these are the outputs" would be enough.
And the fact that pipelines that drive how people consume an increasingly large majority of their content have become this convoluted, should get you annoyed at social media companies: not people who don't understand how bad its gotten
The cost of their actions is easier to frame if you know the details of their implementation.
If people just naturally like using FB that's fine. If FB is hiring experts for behavioral studies in order hypertune a pipeline that exploits blind spots in the way the human brain works... that's a lot less fine. And if the inputs are wider in scope than the general public realizes, it becomes even less ok.
People will always have free will and most would rather we don't try and override that by just making all social media illegal or something... but there's an inherent asymmetry in the resources social media companies have in dissecting the psyche vs individuals have countering that.
Forcing them to show their cards is one way to thrust the results of that asymmetry into the spotlight and better equip individuals who often don't understand the extent to which they're being manipulated.
You do understand that these are ML models where there is no way to explain their behaviour in a layman way other than what is bleedingly obvious.
That it will try and show you more content that you and the people you follow might like.
They've created a monster.
Controversial, hateful stuff was always going to rise to the top. The difference between now and then is we have pretty much unlimited bandwidth covering the globe to spread the message. If you made FB disappear you're going to see the exact same behavior in mastadon servers, just hoping the scale is limited by lower audience sizes.
ML is perfectly explainable, and all this mysticism around what are just weightings of features is silly.
What all recommendation platforms could be required to provide is which are their most important variables for determining what a user sees, and sensitivity analysis.
You're either commenting on something you are oblivious about, or are pulling a red herring just to be contrarian.
Concepts like revealed preferences [1] are nearly a century old, not to mention the fact that the whole consumer behavior field as a whole has plenty of quantitative methods to infer "what people like".
Facebook and Twitter support a feature where quite literally users specify in no ambiguous term that they like something. Theirnl platform is so effective that their problem is filtering out bad actors.
Then there's also engagement. You can get a pretty good idea of what people are interested in by tracking what they read.
You might try to nitpick about the true meaning of "like" but you'd need to be really determined to intentionally miss the point.
When shown an outrage-inducing tweet are you being showing "what you like" or something quite different? Eg., what is most likey to addict you to the platform, what elicits the highest emotional response. Do people trapped gambling their lives away in casions "like" being there?
The idea that social networks are engaged in a sincere project to estimate people's preferences is on the face of it absurd.
And requires such superstitions as "revealed preferences" which are often little more than an observation that people do what their environment most easily affords.
To call that an exercise of preference, to say that people like it, is a great misunderstanding of human nature.
People are much better than what they happen to do in any given environment, and social media companies bare extreme levels of responsibility for creating environments of so-called "revealed preferences" where, apparently people want eating disorders, conspiracy theories, twitter rage sessions, and so on.
It really doesn't. You're trying to debate the semantics of the keyword "like", but that discussion is meaningless and irrelevant. The only objective fact about these services is that they use the data provided by users to infer the preferences they reveal. These services then use this information to meet the company's goals, such as meeting the expectations of paying customers. These paying customers include marketing companies, and their expectations include driving up user engagement and associate their programs with specific type of content.
You might complain about whether these goals are aligned with what you believe are some user's goals, but you'd be missing the fact that users of services like Facebook or Twitter are the product and not the customers. You're not paying for that click, the marketing company is. If customers pay for that click then this clearly means these services do a good job determining what users "like".
So you agree that social media recommendation services are not optimising for situations their users like. So the conclusion of your disagreeable grandstanding is the one I offered in my initial comment.
> The only objective fact about these services is that ... infer the preferences they reveal.
They force them (/create environments) into choices where their actions are advantageous to an ad. company, yes.
The theory of "revealed preferences" is highly contested, and its highly unlikely any such things exist. People's apparent preferences are functions of their environments, and are largely created by them, not "revealed".
The entire tradition of neoclassical economics begins from a methodological premise that closed-form mathematical functions ought represent key terms of interest (eg., preferences). However that assumption is so laughably false it's ridiculous, all social systems -- including preferences --- have path dependence (hysteresis) and are parameterised on such a large number of variables that no such "formula" can be written.
So the least objective, most pseudoscientific description of a situation in which a person enters a casino and gambles their life away is that, somehow a magic casino machine is "revelling" the Preferences of The Human Mind.
In sum, my reply to the comment stands unaltered. Social media recommendation systems are not optimising for what their users like,.
It's not that simple. These features are also social. When you like a tweet on twitter, you send a notification to the author, and other people might see the tweet with something like "alice and bob liked this tweet". So you might decide to like or not to like a tweet for all these reasons.
Youtube doesn't consider likes at all in the algorithm, only watchtime, even though likes are private in youtube, because it's too vulnerable to manipulation, to people telling you to "like and subscribe", etc..
ML is applied to very specific tasks that are pipelined in an intentional and introspectable way.
And even at the level of a given model while you might struggle to explain behavior depending on how pedantically you define that, you still know what your reward function was.
-
This is like saying "There's no point in knowing the source code if you don't have the schematic and microcode for your CPU to know which transistors each line of code is changing the state of"
And an army of observers will be happy to interpret for the less technical.
Sounds like a conspiracy theorist's dream.
It would most likely take several teams to push out what you are asking for. You’d need lawyers, support staff, engineers to create a platform other teams can hook into in order to fulfill these requests, you’d need ways to strip training data of any PII, you’d need a team to write the front end for support staff and the external users. Etc… that’s a non trivial amount of headcount!
You're raving about how much work it would be to build in introspectability into a system that has an extremely large effect on the population, as if that's not justified by the fact we'd be gaining introspectability into a system that has an extremely large effect on the population
You're also overselling it, no one is saying we need every single api request to fetch a feed to also hit some endpoint stating "this is why post id X was shown to user Y". ML again, can be introspected at a higher level than all this seems to imply.
OpenAI can't tell you exactly why ChatGPT gives a specific string of text but they can (and have) shown how it works at a high level: https://openai.com/blog/instruction-following/
You seem to be pushing the line that, because it's difficult to explain why these systems work how they do, and the responsibility is spread so far and wide, we shouldn't bother. It's the literal opposite: it's exactly why its so needed. Even internally these things should be wanted, and I'd be truly shocked if finely tuned advertising giants don't already have a high degree of these understandings in poorly named suborgs that no one pays much attention to (not because cloak and daggers, but because engineers tend to focus on more glamourous things than "introspection").
I'm not sure you thought this problem through. Even if you can specify the inputs, the high level functions mean nothing and the output can and very often is determined by implementation details and the current state of the whole service.
Say for example that personalized recency search of tweets takes into account bad actors being banned/shadow banned/downanked. How exactly do you expect to specify what output you expect from your inputs and your high-level description of your pipeline when a specific spammer is downranked?
I'm sure you can setup a pretty little diagram, full of colors and nice rounded edges, but it would be completely meaningless as state matters, and it's practically impossible to specify what state a global or regional system has in place.
then the follow-up argument should be a much needed (IMO) "simplify this crap so users and regulators come to know how much neutrality and facts they get served".
I understand enough ML to know how and why such systems can't be introspected and why their output is inscrutable to humans, but that's where I draw the line personally: a key component of your business, or one that affects many users, shouldn't be permitted to be so complex that its outcome can't be explained based on a finite number of documented variables and logical branches. We can even think about having a standard to describe and document ML systems (and I would totally look that up because I have the right to know how my keyboard suggests words, why this search result that seems unusual is in fact a sponsored placement, or why this autonomous vehicle chose to run over the pedestrian so not to kill the baby on board...).
And this is not meant to hinder research and development in AI, but to protect people and businesses from potentially dangerous blackboxes, and bring back accountability and ethics in a field that knows very little of it.
There are plenty of effective pharmaceuticals for which we don’t entirely understand the mechanism (e.g. lithium). I think it’s a good goal to be able to understand mechanisms, but if you can effectively show that a solution works to solve a problem through methodical trials, we shouldn’t disallow it’s use.
The issue with Facebook is not complexity or black-box-ness. It’s that the goals of the company and how it makes a profit are fundamentally misaligned with a healthy society.
hmmmm
> We can even think about having a standard to describe and document ML systems (and I would totally look that up because I have the right to know how my keyboard suggests words, why this search result that seems unusual is in fact a sponsored placement, or why this autonomous vehicle chose to run over the pedestrian so not to kill the baby on board...).
Yea you know nothing about ML...
You don't think that any business should be able to use random forests, let alone anything neural network based?
You have managed to make these all sound incredibly inefficient ;P.
I'm not even sure if it uses react/angular or something else
[1]: https://www.cnbc.com/2023/01/20/twitter-is-down-to-fewer-tha...
In a 2022 AWS conference they announced that S3 is currently powered by 235 microservices[0]. It started with 8.
I think it's a great case of something that looks easy from the outside but has a mega iceberg of complexity under the hood to operate at the scale and reliability that S3 has. It's also a good example of if you don't need that scale you can go with a much easier solution of literally uploading a single file to 1 server and being done with it.
Twitter and WhatsApp are both way more complicated than used to be both in terms of features and usage. They also make way more money, which both takes more people and supports more people.
Twitter's main competitor is something that was built by volunteers over the course of a few years. Mastodon of course has its limitations but in terms of features it's actually doing most of the important things that Twitter does. That project has hundreds of competitors but most of the work gets done by a handful of people.
A lot of these big companies accumulate people almost as fast as they accumulate wealth. It's a weird dynamic where they basically recruit themselves into being slow, bureaucratic, and ultimately ineffective. Once you have thousands of people, change and taking risks become very hard .
That's why the startup model is so popular. Startups make things happen that big companies struggle to make happen.
It is possible that the companies could be much smaller and cheaper to run if they dropped the ad-selling and the associated data recording and analysis infrastructure. But it seems to me that currently insufficiently many people are willing to pay for something like that.
1000000000/100 = 10^7
This is a feature of software as an economic good which, particularly when further leveraged in open source ecosystems, makes it fundamentally different that any other example. Imagine 50 "bricks and mortar" engineers trying to provide anything to 900M users.
When you need to start making money, you suddenly need to a lot more engineers. Now you need developers to make sales tools, compliance, moderation, finance, A/B test engineers for growth ideas, niche features for big customers, engineers to bring down cloud costs, etc.
It doesn't make their engineering less impressive but I think we need to understand the full picture.
At the end of the day, WhatsApp needed to make money to keep their service going.
you posted the same comment two times already.
same response
They only spent $9.9 million over $10.2 million in revenues for operations.
WhatsApp’s goal is still growth, rather than monetization. Mark Zuckerberg and WhatsApp CEO Jan Koum said when the acquisition was made in February that ads aren’t the right way to earn money on messaging
They were investing on growth, an investment that lead to a $16 billion acquisition from Facebook.
I would love to replicate their same "unsustainable business"
Are you really that upset about GAAP accounting and how non-cash, stock-based compensation can lead to paper "losses"?
My point remains. They created impressive technology. But 900m users and 50 engineers needed more context. It's not like they shipped a 900 million iPhone-like product with 50 engineers. They made $200k/engineer which is pretty bad compared to more successful companies. When they needed to make more money, they would have needed a lot more engineers.
They never spent investor money from their Series A-on. After Sequoia completed the Series B, Jan and Brian forwarded them a copy of the bank account statement before the Series B financing, showing the original invested funds untouched. They were acquired less than a year later.
Whatsapp was a money-printer! It deserves that reputation.
* ability to isolate a fairly universal need that can be solved in full by software. as soon as you have to customize to meet cultural/political fragmentation you start losing orders of mangitude. once there is a need for a "human in the loop" to provide person-to-person support the "magic" is pretty much gone
* ability to deploy across vast numbers of devices. this universal canvas is generally not the case. exceptions being e.g. web standards or having an oligopoly in mobile devices
still, despite all those caveats, software is a completely different game at its heart. in principle it scales to infinity and it becomes a whack-a-mole game to figure out which factor will land things back to reality
Few hundred or thousands of users can run happily off one server. Then you can just... buy a bigger box for long time, maybe not need more for hundreds of thousands of users.
And once you put all of the effort to get highly available loadbalancing, a bunch of app servers, maybe use some distributed database or add a lot of caching for your current one you can, again, just add more and more nodes for a long time till you get to "next step".
Of course, make your app in Rails with people that have no clue about architecture and some N+1 queries and you might need to get to the scaling steps much earlier
> Imagine 50 "bricks and mortar" engineers trying to provide anything to 900M users.
They are not providing any support to the users tho. Hell, I'm pretty sure some of the best selling cars and motorcycles in history(...being probably beetle and honda super cub) had less than 50 engineers
So you're a front-end guy?
Just because WhatsApp scaled to almost half a billion users with a small engineering team doesn't mean that's the standard, or even achievable, for almost all teams.
The problem is if you decide that you will need 1000 person org in 5 years, you will have 2000 person org in 5 years.
If you challenge yourself to have a 100 person org, you might end up with 1000 people anyways, but at least you're giving yourself a fair chance.
Way too often I see engineering for headcount rather than scale. Building out systems thinking "we'll hire X experts" instead of "we're Y experts, so let's use Y` that's inline with our in house skillset"
It wasn't a sustainable business.
Essentially, due to WhatsApp’s quickly rising valuation, it used share-based compensation to attract top talent. Eventually, the $22 billion acquisition by Facebook would largely make the “expenses” of issuing that stock moot. This wasn’t cash that WhatsApp was burning, but paper money it was doling out.
WhatsApp might be so straight to optimum that no visible properties of engineering are extractable, besides that the engineers made no errors in their judgments.
Making no errors is something you want to replicate, you just don’t know how. Seen like that, the example has no value.
Personally I think that showing that it is possible had some value per se
It's hard to claim, though, that WhatsApp was perfectly engineered when its financial model was unsustainable.
Trying to assess the performance of the engineering team with the money the company generated would lead you to use let's say crappy php web framework everywhere, because that's what most businesses that make profit out of the internet are running for their CRM systems.
Doesn't make any sense.
Trying to assess the performance of the engineering team with the money the
company generated
But is that wrong? And if so, why? We're hired, as engineers, so that the business can make money. Insofar as our excellent engineering makes the business money, it is valuable. When David Hennemeier Hanson wrote the Rails framework for Ruby, his justification for using it was not that it was more elegant than Java, although it certainly was. It was that his convention over configuration approach allowed small teams of developers to launch websites quickly, by having the framework assume that they were doing things in "sensible" ways. would lead you to use let's say crappy php web framework
everywhere
Maybe we should use more "boring" technology like PHP and Java. Has the continued replacement of Java with PHP, PHP with Rails, Rails with Node.js actually benefited our customers? Or has it benefited programmers' desire for novelty at the expense of our customers? At work, one of my co-workers proposed writing a new front-end component in Vue.js, rather than React, because React was "old" and "showing its limitations". I pushed back, arguing that all the rest of our code was in React, and I wasn't sure that he'd necessarily made the case for Vue providing tangible improvements in maintainability or speed of development. We went back and forth, but in the end he won out, and now we have a bit of Vue in our codebase. Has this improved anything for anyone, other than the programmer who now gets to say on his resume that he's implemented production code in Vue.js?EDIT: That's not to say that I'm against all innovation. There are certain classes of innovation that have definitely brought tangible improvements in the speed of development and reliability of the software developed. Memory management is a big one. Going from (legacy) C++ to Java is a huge step forward, just by freeing the programmer from having to worry about the most common sources of memory leaks. Likewise, strong type systems are another advance, that is just coming into the mainstream with Rust (and, to a lesser extent, Typescript). Lisp-like languages (although little used) offer another step forward in productivity by allowing the programmer to write functions that can inspect the internals of other functions as data.
But I feel like that's a different category of thing than trying to argue the merits of React versus Vue or Python versus Javascript.
And then your computer wails in agony as another opened chrome tab drains 6GB of RAM.
So do those tabs Chrome is probably loading.
And yes, in 2040, Chrome will be using 60GB of RAM for "similar" stuff (superficially similar).
In terms of actual revenue per employee, Twitter (pre-Elon) actually beat Whatsapp (pre-Facebook).
A better measure of success might be employees/revenue.
In 2014, Whatsapp made $10m and lost $140m. That's only $200k/engineer.[0]
By comparison, Apple is at 2.3m/employee. Meta is 1.6m/employee. Twitter is 680k/employee. These include non-engineers of course.
When you need to start making money, you suddenly need to a lot more engineers. Now you need developers to make sales tools, compliance, moderation, finance, A/B test engineers for growth ideas, niche features for big customers, engineers to bring down cloud costs, etc.
[0]https://www.theverge.com/2014/10/28/7085905/facebooks-prized...
I don't think it is commonly held up as a standard. Too few people are aware of the magnitude of its success for it to play that role, IMO. If anything its story is underrated and understated. Instagram's, as well.
Idolizing anything is a mistake, but treating them as a case study or example for lean engineering feels entirely appropriate -- who else would you point to?
But most of all, it was created much much later, after all the patterns and techniques have already being ironed out by the first wave of chat systems.
I don't see how that has an edge over the cloud sync functionality though.
Cloud sync, having thousands and millions of people in a group, group topics, 4GB file sending limit are more impressive than some closed source implementation of E2EE we can never even verify.
> all the java phones whatsapp does ( not sure if whatsapp still do though).
It doesn't anymore.
Telegram works even without an app, through their website (On phones too, yes) so I'd say it's more accessible and having open source clients and APIs means that you can create a Telegram client for literally any platform.
This is, imho what makes whatsapp uniquely hard to scale.
The article is from 2015. Close to irrelevant, alas.
People were always fascinated by Erlang (or rather the underlying BEAM) but it never really broke into the mainstream, probably because it is so different.
The hard part is making this a successful business. Because the minute you try and monetise it with advertising you will quickly find out that building a competitive ad platform is where the real challenge is.
No, you can't.
If it really is trivial today, it was in 2015: cloud providers already existed, smartphones already existed, programming tools were already more than mature, anybody could run a few instances of ejabberd and build an app to talk to them.
WhatsApp image upload used a simple PHP page...
It wasn't exactly rocket science, it was simply very well executed.
Doing a crappy chat app, yes sure that's easy. But one that works as reliably, with the general performance of whatsapp ? Even with today's tech, that's super hard.
so things become irrelevant after 7 years now?
will this comment be irrelevant in 7 years?
This doesn't ring true for me, especially in an era where services so easily communicate via http
Whatsapp raised 250,000 seed. Now seeds are over a few millions at least.
That in turn makes companies spend like crazy and not focus on efficiency and so on.
Essentially, they have the same revenue & product as Asana, but with only 1/20th the employees.
Such a feature hasn't been changed and now I know why
Who is "everyone"? Anyone that searches your name, or all your contacts? (I've never used WhatsApp.)
1. People in a conversation with you (including a group conversation you are added to, which I believe you can limit to require your consent) can see your phone number
2. People who already have your number in their contacts can see that your number is registered with WhatsApp
WhatsApp only uploads phone numbers from your contacts. No other information. This is all documented here: https://faq.whatsapp.com/1191526044909364/
https://zerodha.tech/blog/hello-world/
Tldr: How 30 member tech team formed over seven years built India’s largest stock broker.