The bit that you get wrong is assuming that Facebook or Twitter are simple platforms, and that their main and most complex features is allowing users to post, search and read messages. This is entirely wrong as these platforms are actually ad-selling platforms and thus include region-specific services, localized ranking, all sorts of data harvesting, etc etc.
Also, as companies grow they start to feel the need to implement efficient development processes, and thus start to allocate resources to stuff like putting together their own build farms, CI/CD infrastructure, deployment auditing, etc etc etc. They also develop things like front-end and back-end frameworks. It's not like a random team at Facebook just signs up for gitlab and does all it's business from there.
They don't need to release the actual code, but releasing "these are the inputs we use, these are the high level functions of the pipeline, these are the outputs" would be enough.
And the fact that pipelines that drive how people consume an increasingly large majority of their content have become this convoluted, should get you annoyed at social media companies: not people who don't understand how bad its gotten
The cost of their actions is easier to frame if you know the details of their implementation.
If people just naturally like using FB that's fine. If FB is hiring experts for behavioral studies in order hypertune a pipeline that exploits blind spots in the way the human brain works... that's a lot less fine. And if the inputs are wider in scope than the general public realizes, it becomes even less ok.
People will always have free will and most would rather we don't try and override that by just making all social media illegal or something... but there's an inherent asymmetry in the resources social media companies have in dissecting the psyche vs individuals have countering that.
Forcing them to show their cards is one way to thrust the results of that asymmetry into the spotlight and better equip individuals who often don't understand the extent to which they're being manipulated.
You do understand that these are ML models where there is no way to explain their behaviour in a layman way other than what is bleedingly obvious.
That it will try and show you more content that you and the people you follow might like.
They've created a monster.
Controversial, hateful stuff was always going to rise to the top. The difference between now and then is we have pretty much unlimited bandwidth covering the globe to spread the message. If you made FB disappear you're going to see the exact same behavior in mastadon servers, just hoping the scale is limited by lower audience sizes.
ML is perfectly explainable, and all this mysticism around what are just weightings of features is silly.
What all recommendation platforms could be required to provide is which are their most important variables for determining what a user sees, and sensitivity analysis.
You're either commenting on something you are oblivious about, or are pulling a red herring just to be contrarian.
Concepts like revealed preferences [1] are nearly a century old, not to mention the fact that the whole consumer behavior field as a whole has plenty of quantitative methods to infer "what people like".
Facebook and Twitter support a feature where quite literally users specify in no ambiguous term that they like something. Theirnl platform is so effective that their problem is filtering out bad actors.
Then there's also engagement. You can get a pretty good idea of what people are interested in by tracking what they read.
You might try to nitpick about the true meaning of "like" but you'd need to be really determined to intentionally miss the point.
When shown an outrage-inducing tweet are you being showing "what you like" or something quite different? Eg., what is most likey to addict you to the platform, what elicits the highest emotional response. Do people trapped gambling their lives away in casions "like" being there?
The idea that social networks are engaged in a sincere project to estimate people's preferences is on the face of it absurd.
And requires such superstitions as "revealed preferences" which are often little more than an observation that people do what their environment most easily affords.
To call that an exercise of preference, to say that people like it, is a great misunderstanding of human nature.
People are much better than what they happen to do in any given environment, and social media companies bare extreme levels of responsibility for creating environments of so-called "revealed preferences" where, apparently people want eating disorders, conspiracy theories, twitter rage sessions, and so on.
It really doesn't. You're trying to debate the semantics of the keyword "like", but that discussion is meaningless and irrelevant. The only objective fact about these services is that they use the data provided by users to infer the preferences they reveal. These services then use this information to meet the company's goals, such as meeting the expectations of paying customers. These paying customers include marketing companies, and their expectations include driving up user engagement and associate their programs with specific type of content.
You might complain about whether these goals are aligned with what you believe are some user's goals, but you'd be missing the fact that users of services like Facebook or Twitter are the product and not the customers. You're not paying for that click, the marketing company is. If customers pay for that click then this clearly means these services do a good job determining what users "like".
So you agree that social media recommendation services are not optimising for situations their users like. So the conclusion of your disagreeable grandstanding is the one I offered in my initial comment.
> The only objective fact about these services is that ... infer the preferences they reveal.
They force them (/create environments) into choices where their actions are advantageous to an ad. company, yes.
The theory of "revealed preferences" is highly contested, and its highly unlikely any such things exist. People's apparent preferences are functions of their environments, and are largely created by them, not "revealed".
The entire tradition of neoclassical economics begins from a methodological premise that closed-form mathematical functions ought represent key terms of interest (eg., preferences). However that assumption is so laughably false it's ridiculous, all social systems -- including preferences --- have path dependence (hysteresis) and are parameterised on such a large number of variables that no such "formula" can be written.
So the least objective, most pseudoscientific description of a situation in which a person enters a casino and gambles their life away is that, somehow a magic casino machine is "revelling" the Preferences of The Human Mind.
In sum, my reply to the comment stands unaltered. Social media recommendation systems are not optimising for what their users like,.
It's not that simple. These features are also social. When you like a tweet on twitter, you send a notification to the author, and other people might see the tweet with something like "alice and bob liked this tweet". So you might decide to like or not to like a tweet for all these reasons.
Youtube doesn't consider likes at all in the algorithm, only watchtime, even though likes are private in youtube, because it's too vulnerable to manipulation, to people telling you to "like and subscribe", etc..
ML is applied to very specific tasks that are pipelined in an intentional and introspectable way.
And even at the level of a given model while you might struggle to explain behavior depending on how pedantically you define that, you still know what your reward function was.
-
This is like saying "There's no point in knowing the source code if you don't have the schematic and microcode for your CPU to know which transistors each line of code is changing the state of"
And an army of observers will be happy to interpret for the less technical.
Sounds like a conspiracy theorist's dream.
It would most likely take several teams to push out what you are asking for. You’d need lawyers, support staff, engineers to create a platform other teams can hook into in order to fulfill these requests, you’d need ways to strip training data of any PII, you’d need a team to write the front end for support staff and the external users. Etc… that’s a non trivial amount of headcount!
You're raving about how much work it would be to build in introspectability into a system that has an extremely large effect on the population, as if that's not justified by the fact we'd be gaining introspectability into a system that has an extremely large effect on the population
You're also overselling it, no one is saying we need every single api request to fetch a feed to also hit some endpoint stating "this is why post id X was shown to user Y". ML again, can be introspected at a higher level than all this seems to imply.
OpenAI can't tell you exactly why ChatGPT gives a specific string of text but they can (and have) shown how it works at a high level: https://openai.com/blog/instruction-following/
You seem to be pushing the line that, because it's difficult to explain why these systems work how they do, and the responsibility is spread so far and wide, we shouldn't bother. It's the literal opposite: it's exactly why its so needed. Even internally these things should be wanted, and I'd be truly shocked if finely tuned advertising giants don't already have a high degree of these understandings in poorly named suborgs that no one pays much attention to (not because cloak and daggers, but because engineers tend to focus on more glamourous things than "introspection").
I'm not sure you thought this problem through. Even if you can specify the inputs, the high level functions mean nothing and the output can and very often is determined by implementation details and the current state of the whole service.
Say for example that personalized recency search of tweets takes into account bad actors being banned/shadow banned/downanked. How exactly do you expect to specify what output you expect from your inputs and your high-level description of your pipeline when a specific spammer is downranked?
I'm sure you can setup a pretty little diagram, full of colors and nice rounded edges, but it would be completely meaningless as state matters, and it's practically impossible to specify what state a global or regional system has in place.
then the follow-up argument should be a much needed (IMO) "simplify this crap so users and regulators come to know how much neutrality and facts they get served".
I understand enough ML to know how and why such systems can't be introspected and why their output is inscrutable to humans, but that's where I draw the line personally: a key component of your business, or one that affects many users, shouldn't be permitted to be so complex that its outcome can't be explained based on a finite number of documented variables and logical branches. We can even think about having a standard to describe and document ML systems (and I would totally look that up because I have the right to know how my keyboard suggests words, why this search result that seems unusual is in fact a sponsored placement, or why this autonomous vehicle chose to run over the pedestrian so not to kill the baby on board...).
And this is not meant to hinder research and development in AI, but to protect people and businesses from potentially dangerous blackboxes, and bring back accountability and ethics in a field that knows very little of it.
There are plenty of effective pharmaceuticals for which we don’t entirely understand the mechanism (e.g. lithium). I think it’s a good goal to be able to understand mechanisms, but if you can effectively show that a solution works to solve a problem through methodical trials, we shouldn’t disallow it’s use.
The issue with Facebook is not complexity or black-box-ness. It’s that the goals of the company and how it makes a profit are fundamentally misaligned with a healthy society.
hmmmm
> We can even think about having a standard to describe and document ML systems (and I would totally look that up because I have the right to know how my keyboard suggests words, why this search result that seems unusual is in fact a sponsored placement, or why this autonomous vehicle chose to run over the pedestrian so not to kill the baby on board...).
Yea you know nothing about ML...
You don't think that any business should be able to use random forests, let alone anything neural network based?
You have managed to make these all sound incredibly inefficient ;P.
Hiring an engineer has a significant opportunity cost.
You're right that in the short term you aim to hire where the return is highest, but in the long term that means they're right that you'd want to keep hiring as long as you get a positive return.
You want to spend it on whatever has the highest return, which will frequently be on an expense other than hiring, and when it is on hiring, it will frequently be on hiring roles other than engineers.
There isn't really an opportunity cost to hiring more people when money is no object. More likely to be the opposite.
Why do you think they increased their headcount so much in the last few years? Simply because they could.
The fact there is finite budget and headcount doesn't necessarily mean anything. If you're limited by the number of quality employees you can recruit and onboard, budget and headcount are effectively meaningless because they're not limited by money.
How to apply it wisely is a different, much wider topic.
The question really is, how can you increase shareholder value? If the answer requires engineers then they are hired. They are paid less then the value they add. When that stops being true people like Elon Musk can increase value by reducing headcount.
That’s not the part that gives insight to the question “how can it take X engineers to run service Y?” The insight is that employees can add more value than they cost even if they aren’t working on the most visible core product that the company is known for.
> a firm maximizes profit by producing that quantity of output where marginal revenue equals marginal costs.
This is a pretty fundamental result in microeconomics. The profit maximization section of the wiki article explains it quite well: https://en.m.wikipedia.org/wiki/Marginal_product_of_labor
The usual disclaimers about underlying assumptions apply of course.
In a 2022 AWS conference they announced that S3 is currently powered by 235 microservices[0]. It started with 8.
I think it's a great case of something that looks easy from the outside but has a mega iceberg of complexity under the hood to operate at the scale and reliability that S3 has. It's also a good example of if you don't need that scale you can go with a much easier solution of literally uploading a single file to 1 server and being done with it.
Twitter and WhatsApp are both way more complicated than used to be both in terms of features and usage. They also make way more money, which both takes more people and supports more people.
It is possible that the companies could be much smaller and cheaper to run if they dropped the ad-selling and the associated data recording and analysis infrastructure. But it seems to me that currently insufficiently many people are willing to pay for something like that.
Twitter's main competitor is something that was built by volunteers over the course of a few years. Mastodon of course has its limitations but in terms of features it's actually doing most of the important things that Twitter does. That project has hundreds of competitors but most of the work gets done by a handful of people.
A lot of these big companies accumulate people almost as fast as they accumulate wealth. It's a weird dynamic where they basically recruit themselves into being slow, bureaucratic, and ultimately ineffective. Once you have thousands of people, change and taking risks become very hard .
That's why the startup model is so popular. Startups make things happen that big companies struggle to make happen.
I'm not sure about that. These framework projects sound like internal projects developed to meet internal needs and improve internal processes. The GraphQL concept and use case in particular, as well as API management tools, is something I've seen reinvented internally by companies to be able to simplify processes. You just hear about GraphQL because Facebook decided to release it publicly. Everyone else has been more or less reinventing the wheel to meet the same needs.
I'm not even sure if it uses react/angular or something else
[1]: https://www.cnbc.com/2023/01/20/twitter-is-down-to-fewer-tha...
Adding yet another zero to your user headcount is costly and a true engineering challenge.
Also, organizational inefficiencies are very costly and require more and more people to fill more and more layers of cintributoand managers alike. But mostly managers.
Edit: Third, product space exploration should not be forgotten.
And demonstrably 50 engineers was sufficient for handling stuff at scale.
but you could add another 0 to the message count. You could add another 0 to the image count.
That seems highly unlikely unless they launched a new product, or this was a gradual shift over many years.
Matrix as a protocol provides the building blocks for both messengers and collaboration tools (and more exotic stuff like spatial collaboration apps). Element is an example of a collaboration app (although Element X is being built to feel more like a messenger which happens to also support collaboration use cases - more like Telegram, but encrypted and open and decentralised). FluffyChat is an example of more of a messenger use case on Matrix. Thirdroom.io is an example of virtual worlds on Matrix.
So: I’d say that WA and Discord are very different products. But you could build both on the same underlying primitives (eg a stack of IP, HTTP, TLS, Matrix)
I remember the first time I saw the Facebook Ads platform, the Events Manager etc - it's a highly complex system that changes all the time. Many more engineers are working on that than on maintaining the ability to Reply to a post feature.
WhatsApp, until Meta, had none of these things.