What would open sourcing the Twitter algorithm actually look like?
transitivebullsh.it
transitivebullsh.it
And by giving more visibility to polarizing content it incentivises others to create more polarising content.
The algorithm therefore is a positive feedback loop of polarization. It doesn't just show what we are, but amplifies our worst (or tribalistic) tendencies.
Relevant Howard Stern reference: https://www.youtube.com/watch?v=9G6xu-J_Dmc
i.e. If the algorithm had more insight into a given person’s emotional (and intellectual?) response, the outcomes could be better than the status quo.
[1] https://open.spotify.com/episode/3DBNkrAY5sCIX2sqHpZ567?si=M...
But you understand that technology have social consequences? Discussing them is vital part of discussing technology. "Moving fast and breaking things" is fun and all, but when you endangering peoples lives and wellbeing, maybe you should think twice about what you are doing.
Talking about technology as though it has any moral imperative other than "make more money" in the West is naive at best. At worst it is not only boring to anyone over 21, but it's counter productive to improving the world. Our newest toys are there to distract from the woman behind the curtain.
In short: let me read what I want and save your moralizing for the people who actually make decisions. None of which are technologists (any more).
You mean it's slave to people controling capital? Because capital by itself isn't sentient or capable of making decisions.
>Talking about technology as though it has any moral imperative other than "make more money" in the West is naive at best.
Technology have moral consequences, not a "moral imperative". Technology doesn't care if it counts number of Jews, carrots or cans of Zyklon-b. But people should, and normal people do. Tehnology shouldn't be an excuse to do whatever you want, because "computer said so". Someone programed/tauhgt that computer. Someone is responsible for it.
>At worst it is not only boring to anyone over 21, but it's counter productive to improving the world.
Capital, corporations and billionaires are human inventions, so does kings, slavery, white supremacy, misogyny. Kings didn't stopped to be kings just because everyone asked them nicely. It took time but in the end, kings were no more. Because in the end they depend on ordinary people to do their biddings. Jeff Bezos doesn't check if drivers reached their quota of urin filled bottles, people (directly or indirectly) do this.
>In short: let me read what I want and save your moralizing for the people who actually make decisions. None of which are technologists (any more).
Not talking about problematic technologies, decisions made by people in power doesn't make them go away. Pretending bad things didn't happen is delusional, but just talking doesn't help either. Actions are also needed, even if they are inconvenient or dangerous. There are many ways change can be brought, protests, civil disobedience and more. But pretending everything is ok, will make every call to action fail, even before it started.
Yes that is why I started of this whole thread with: > "The algorithm" is irredeemable.
My personal opinion is formed by tweets that where not about the social or societal consequences of software development but totally unrelated events from people I didn't follow.
We are independently blaming every social media platform for what is, fundamentally, a human problem.
Second, part of the problem is diving into the delusion by picking a side. One such side is men vs women.
I would argue that almost the entire problem with the internet(including Twitter) is advertising. All the problems stem from it.
Getting rid of the "trending" block, with something else, like "trending tag within people you follow" would get rid of this normalisation of extreme behavior.
Then you need moderation.
Seeking engagement is not necessarly evil.
GitHub increased engagement, and it lead to more cooperation in the open source software, that's a plus for everyone.
The issue of a lot of social networks, is they want engagement, even if it cause addiction.
I think twitter social network "model" is salvagable in something healty, but it needs a lot of changes.
I'd be happy to pay $1 or $2/mo for this directly from Twitter, but until that's an option I don't feel remotely bad removing my sliver of ad revenue from their bottom line.
> Seeking engagement is not necessarly evil. GitHub increased engagement [...]
At this point it really depends on what you mean by "engagement". GitHub's definition of "engagement" is definitely not the same as Twitter's.
Technically, Twitter's goal is not "engagement" but "time spent looking at ads". It just so happens that "engagement" is a very good proxy for that. In Twitter's case, increasing "engagement" is definitely evil.
I also don't believe GitHub's goal is to increase engagement. Their goal is ultimately to deliver features useful to their paying customers - companies that build software, where as Twitter's (and any other ad-supported service's) goal is to serve more ad impressions (or at least make advertisers believe they do).
I don't think that's any better than the current scheme; infact, it could be worse becasue the current ad system has centralized control which can check on extremist echo-chambers. I suspect Patreon-style funding will increase polarization[1], with the "content-creators" optimizing for what brings in the most money, not necessarily their thoughtful ideas.
1. With a side of increased brigading, coordinated harassment,doxxing and swatting; but that'll be the price of "free" speech.
It's nothing magical. It's just a big system trying to find the things it thinks you are most likely to find engaging.
There are similar algorithms driving advertisement and one of the ads that has been targeted at me lately is a nicotine source two years after I kicked my nicotine habit. https://en.wikipedia.org/wiki/Snus The intention is clear, to make me pick it up again. Given this is simply a personal anecdote but it is eerily accurate sometimes.
At least with facebook, but I assume with others, they are using engagement as a proxy for delivery value to you. That is the goal they are trying to maximize for.
So they want to maximize how much pleasure you get from the service, so they maximize engagement, but they don't necessarily maximize revenue, because that would hurt engagement.
Twitter's scale and real-time nature make it a difficult beast. Their network graph contains hundreds of millions of nodes and billions of edges.
And it's constantly updating. So any graph ML algorithms you want to use have to deal an underlying graph that's eventually consistent at best — and oftentimes very sparse in terms of feature availability.
If yes, how can I contact you? Twitter?
Building it today would be much faster because there are actually proper libraries/programs for doing this, rather than my inefficient vanilla Python implementation.
For transparency, it might be nice for Twitter to continually release their algorithm from a certain time period ago, say 6 months
I see the tweets of everyone BUT the ones I follow. Hence I have a chrome plug in that cleans up their junk into a usable state.
I also use https://github.com/giuseppeg/refined-twitter-lite
I sincerely hope I never learn the names of the people at Twitter responsible for this user-abusing product decision.
Unfortunately for Twitter I will keep using the browser plug-in which also hides their ads (maybe they get paid for the display, even if hidden)
1. I use Tweetdeck on PC (browser).
2. I also prefer to add people to lists rather than following them.
So on Tweetdeck I have multiple columns open at once, the first column is the people I follow (Tweetdeck presents them in chronological order, no messing about with similar users or liked tweets) and the rest of my columns are my lists. (F1, MotoGP, NBA, Space/Rocketry, Tech).
I mainly use Twitter on PC but whenever I use it on mobile it's to look through my lists. The timeline is pretty much useless on all apps.
I wonder why platforms wouldn't just offer different algorithms. e.g. historical, maximise-time-on-screen or exploration. Most of the users would be nudged into what the platform wants anyway, and they will be able to claim moral superiority over other platforms that doesn't offer this. Heck, you could even charge for this feature.
TL;DR they used to have an open API and various third-party clients, each of which had their own flavor. BUT this made it really hard for twitter to guarantee a good UX and even harder for twitter to implement an ad-driven business model.
Aside from that, thanks for the link. That is a great post!
Since I don't see open sourcing the actual feed algorithm to be something that would practically happen any time soon (which makes me a sad panda).
It's why government regulation is probably the only way to fix the issue, since you need companies to act in a way that's counter to their own interests (but would lead to a better experience for users.)
https://writings.stephenwolfram.com/2019/06/testifying-at-th...
I got the inspiration from this article https://towardsdatascience.com/i-created-my-own-youtube-algo...
And was writing FE for it in Electron/React
SELECT
t.*
FROM
tweets t, followers f
WHERE
t.author_id = f.author_id AND f.user_id = ?
ORDER BY rand()
there its open source nowPeople, please, I am begging you - turn those back on. It's the only way we can stay up-to-date with your amazing content.
Anyway, this is what I came to say. A great post, full of insights. We have discovered quite a few of them ourselves, while building https://murmel.social, and I envy the author a little for publishing the knowledge first. Next time ;)
https://transitivebullsh.it/feed.xml
Great call!
What I'm really interested in, however, is help with answering the following questions:
What would open sourcing the Twitter algorithm actually look like?
Would it be possible to abstract away all of the engineering complexity that it takes to run a global, production system like Twitter and produce an OSS spec or API that is actually useful?
Would it be possible to produce meaningful results without access to Twitter’s full data set?
What does meaningful even mean here? How would we define success?
What would need to happen in order to make this a reality?
What are some practical proposals that would help improve the status quo?
What are the actual requirements that you're looking to satisfy by doing this?
These are exactly the types of questions I'm trying to get more people to think about — and the reason I wrote the article above.
It's not well-defined, but the goal imho should be to allow more transparency and optionality around Twitter's main feed UX.
Open sourcing twitter's algorithmic feed would be a monumental task that's honestly unlikely to ever come to fruition, Elon or no Elon. BUT twitter could certainly do more to move in that direction.
> What are the actual requirements that you're looking to satisfy by doing this?
If it improves transparency around how twitter's algorithm works.. if it introduces more optionality into twitter's main UX (and isn't just a chrome extension, for instance), if it's coming from twitter itself and not a third-party.. if the effects of different factors on relevancy signals are made more transparent (like being able to play around with a simulated tweet under different conditions and seeing the resulting relevancy scores for a set of example users).. there's just so much that twitter could do here to improve transparency around how their feed works. And I'm generally in favor of anything what will move the status quo.
In Elon's words “Civilizational risk is decreased the more we can increase the trust in Twitter as a public platform.”
like tweet was displayed because 50% interest in dogs 20% friends liked it 30% generally popular
Given that Twitter/Facebook/Reddit/Tiktok are the behemoths in the room that all have the power to sway public opinion by either promoting or suppressing information, I think they should be regulated so that any topic that was promoted or suppressed by the algorithm or manual human intervention should show itself on the content just like an ad does. "This content has been quarantined by [Twitter Staff / Automated Review] and is only available to those with the direct url." would be a perfect example of such a message.
From there, users should have control over if they want to see such quarantined content or if they want a "safe" experience that Twitter curates for them.
The scary thing is, this isn't a conspiracy myth. Back in 2014 FB got caught at experimenting at how algorithm changes affected the mood of users [1].
[1] https://www.forbes.com/sites/kashmirhill/2014/06/28/facebook...
Even something as simple as a keyword search, ie "hunter biden laptop", would require searching tens of millions of strings per hour. Distributed across thousands of servers, integrated with whatever dev ops set up they have. It just seems like a huge pain, for little practical gain.
Their scale is so huge that I doubt twitter has much manual control over what people see, even if they wanted it. It seems much easier to just multiply weights
I have five columns: personal account, replies to it, side project, replies to that, messages. Only thing I miss from the old OSX Twitter client is the ability to copy and paste in images.
Hell, whose to say that some people aren't chosen for long term "experiments" with different behavioral/belief modification goals? That's what I would do.
Even if there were some sort of open source spec like the W3C specs, for instance, Twitter would still be a black box that could choose to conform to the spec or not choose to.
What we need is something more enforceable. So either you do something completely green field like https://blueskyweb.xyz/, or you spend $43B & roll up your sleeves to try and fix the problems with the current platform.
Note that everyone has access to the reverse chronological feed called "Latest tweets". See here for how to switch: https://twitter.com/TwitterSupport/status/144797117918529536...?
Also agreed that twitter pro would be a great place for twitter to add more options.
[1] https://twitter.com/johnkrausphotos/status/15172153497243525...
[2] https://www.washingtonpost.com/news/monkey-cage/wp/2016/05/1...
[3] https://twitter.com/Alwaleed_Talal/status/151461595698675712...
Twitter would be so much better.
People are obsessed with personalization algorithms and disappointed when this involves giving people control to personalize. As though personal control isn’t magical enough (in comparison to the magic of the algorithm).
I tried to reference as many of the official sources as possible (in addition to drawing on my past experienced at Amazon + Facebook).
thanks for calling out scale and engineering problems btw! sometimes i think people think we're sitting around all day throwing darts at a wall of pictures of conservatives to pick who to ban next.
The Coca-Cola formula seems like a pretty good analogy, honestly. Most people can't tell the difference between different cola recipes, and there probably isn't anything in the recipe itself that's responsible for the company's success. Not as much as marketing, and simply being in the right place at the right time, anyway.
Source: https://www.washingtonpost.com/technology/2022/04/16/elon-mu....
With that said, the timeline is somewhat secret sauce but mostly it’s just a glop of various parameters that aren’t very interesting to share. Like, how are you going to even describe it to someone? It’s easy to talk about changing a button color, but “we promote retweets 10% more now” is boring and not really something that most people care about.
In reinforcement learning there is a notion of exploration and exploitation. All of supervised ML can in a way be seen as exploitation.
Incorporating ideas of exploration in a clever way into the standard recommendation algos, I think could be the solution to a lot of the mentioned problems.
The goal here is to increase transparency and optionality around twitter's core feed. There will always be difficult, contentious human problems with a platform like twitter — but instead of writing it off as "laughable", there's a lot that can be done to improve the status quo. And imho that's a very impactful goal worth striving for.
Twitter uses algorithmic feeds to present its users with an ideologically driven view of whatever is being discussed, this can easily be seen by reading the first comments on contentious issues where they nearly always push some 'progressive' comments to the front, even when those comments are clearly less popular with users than whatever other comments follow. Comments critical of 'progressive' issues often require one or more extra clicks to be made visible. Sometimes this leads to the first page being devoid of comments, requiring a click to access whatever comments are present where those comments are nearly invariably critical of the 'progressive' issue being discussed. This does not happen when there are positive comments on those 'progressive' issues since these are presented directly on the first page.
TikTok and Youtube also use algorithmic feeds to channel users but their algorithms seem to be tailored to keep users on the site/app for as long as possible.
As you pointed out, not everything is algorithmic. There are definitely human factors at play, but without a more solid understanding of how twitter's algorithmic feed actually works, the discussion will remain hand wavy at best.
Just my 2 cents :)
Popular non-'progressive' senders - Trump and now Musk being the poster child of such - probably get their own comment manager who makes sure there is always some vitriol waiting no matter the message. This does not scale but the Pareto principle makes that it does not have to since there are not that many of such. The same may be true for popular 'progressive' senders where comment managers may make sure there are supportive comments but this seems to be less clear-cut.
In the secretive context of the current non-profitable public Twitter model, yes, I agree, they cannot do anything useful.
Not sure why you got a downvote though, so have my upvote.
Though I personally think throwing out any form of algorithmic feed wouldn't happen. There are just too many business and UX reasons it makes sense, which is why every major social platform uses them (https://www.socialmediatoday.com/social-networks/truth-about...).
Someone with a short following list can see the drama if they look at the huge list of replies that a popular account might get. So much cruft.
'The Algorithm' is not that important to Twitter.
What makes it work, is a semi-coordinated mass of users on a reliable platform.
If that platform could be distributed, and, there was a critical mass, it could work, maybe.
But 'open sourcing' is really not the issue at all. There isn't a line of code at Twitter that's valuable on it's own to anyone.
You could hand over the Source Code to Trump & Co. for their deluded 'Truth Net' or whatever and it wouldn't make a difference.
A good thought exercise is to try to imagine what would happen if an absolutist of any kind got their way in every way possible, and still hated the results.
Really committing to free speech absolutism requires a very strong stomach very few people truly have. Back when Voat was a thing it was such a cess-pit that when r/TheDonald tried to move there, they couldn't take it. Voat disdained moderation, and as a result filtered for the kind of person who can look at a feed full of conspiracies, the most open and clear hate, fascism, racism and antisemitism, and every kind of porn that's banned on Reddit, and not even flinch at it. The denizens of r/TheDonald were used to setting rules on their turf, and were relentlessly mocked for it until they crawled back to Reddit.
Voat eventually died. I think probably because in the end even the creators grew disillusioned with the community they had created, or because they realized it was only a matter of time before they got into serious legal trouble.
``` Elon’s motivation is clear and consistent with his modus operandi. It’s the same reason he’s working so hard to build a sustainable colony on Mars, why he’s devoted resources to understanding the potential dangers of AI super-intelligence, and why he’s so insistent on combatting climate change. His guiding motivation is to improve humanity’s chances at a positive future. ```
That's certainly one interpretation.
What I was trying to do was explain my motivation for digging deep into this topic. It has nothing to do with ideological polarization, though it may be a bit speculative.. this is mainly based on Tim Urban's excellent series of articles on why Elon Musk chooses to work on the problems he does. https://waitbutwhy.com/2017/03/elon-musk-post-series.html