Twitter Should Open Up the Algorithm
every.to
every.to
That's not an algorithm, that's product strategy & management.
The net suggestion is to demand more, and start describing solutions, today, to get data visibility without compromising other things, to help offset the inevitable objections that will come once people are forced to demand it when they realize the code didn’t really solve the fundamental problem.
A social network complaining about hyperbolism and outrageous claims being made on it's platform is like complaining people are breathing too much air.
Twitter is not a place grownups should be participating.
The early philosophical developments around freedom of speech didn’t foresee this but now that it is here, we have to ask the question of what ought we want if we are to uphold the principle that people should be free to speak unpopular ideas without censorship or fear of excommunication, given that a key ingredient in progress, revelation of the truth, and finding compromise to avoid violence.
Twitter is of course a private company, and can do what they want. But what those who wish to see free speech principles upheld say is we ought to want those in twitters position to be as lenient as possible so as to not stifle the free exchange of ideas. And, perhaps failing that, we ought to want to see another mechanism that is less susceptible to widespread censorship and overreach. Too many people seem counter someone explaining that they feel the status quo is undesirable with an is/ought fallacy - it’s our job in liberal democracies to continually raise and debate issues regarding basic freedoms and the ability for our society to continue to evolve according to shared principles.
https://blog.openmined.org/announcing-our-partnership-with-t...
And also other way around, data depends on algorithms and tools that created them. You can't have insightful understanding of data without knowing how and why they were created.
> “One of the things I believe Twitter should do is open source the algorithm and make any changes to people’s tweets — you know, if they’re emphasized or de-emphasized — that action should be made apparent so anyone can see that action’s been taken. So there’s no sort of behind-the-scenes manipulation either algorithmically or manually.” Musk also added later, “The code should be on GitHub so people can look through it.”" (CNBC interview)
If I could pick something to regulate about recommendation algorithms, probably a big part of it would be mandating consumer controls over the algorithm. Complete transparency might be hard when sometimes even the programmers of these systems don't have the ability to see why they recommend what they do, but there should definitely also be mandatory oversight by consumer protection agencies that have full ability to audit the code and report on it.
Fortunately this has not turned out to be the case in reality. If there are flaws in the system they will probably be easier to find and fix with more eyes on it.
In content moderation vs spammers, you have spammers adapting to your tools in real time while posting content in the same way real users do, and the balance between "don't want to accidentally ban real users (who post stuff very similar to the bots)" vs "does it matter if someone gets past your defences a few times an hour" is massively different too.
The problem with google search isn't that people are exploiting 0-days that the developers aren't aware of. The problem is that the majority of the internet is becoming nearly identical pages that do the bare minimum to get ranked by search engines. Making the algorithm public will only exacerbate that
You assert that more eyes on the flaws would make the issue worse rather than better, but i disagree.
Further, compromise of one box typically does not render all of them running that OS useless and harmful.
Especially not if the top choices were fundamentally different from each other.
It may go so far as to make it significantly more difficult for bad actors to operate.
Let's stop the Communistic mentality around keeping things hidden. Please read about Soviet Russia, about Siberia and how Russians had no idea about the miseries inflected on them for decades.
1. Source?
2. Thought experiment: if said subjective human intervention was recorded and codified into an algorithm it wouldn't be as contentious?
“After close review of recent Tweets from the @realDonaldTrump account and the context around them we have permanently suspended the account due to the risk of further incitement of violence,”
Heck, even create an alogrithm store where people could create and purchase different ways to view their feed.
It doesn't seem that difficult to add these multiple sort and filter options, but maybe it's more complex than I imagine.
Read only audit is an overly naive way to think about understanding complex algorithms. In an audit they would find an enable of large DNN transformer models with thousands of layers and thousands of features often with their own transformation trees. There are entire CS departments dedicated to researching tools to understand complex nonlinear models. You definitely can't do with a read only code audit. You can _barely_ do it with full access to the model, full access to the input data and being able to retrain and rerun the model on that data.
Engineers don't exist in a vacuum, the same engineers who write the algos are here on HN. Show HN the Twitter source code and a lot of engineers will understand it.
What is more likely is that the engineers are currently bound by NDAs and can't say what influences the algo.
The fact that we calculate precisely what users see on their devices on servers is a result of the architectural constraints of the time, and more importantly, the ad model of monetizing social networks.
If social networks were not monetized in this way, there would be far more power allocated to clients and APIs.
This isn't to say user devices can do everything: they can't. But can easily be given significant power to filter, reorder, and request different content — and with more advanced engineering, allow users to parameterize feeds.
The reason we can't have this is that allowing this degree of user choice undermines the ads model.
That would involve giving clients access to information that clients probably shouldn't have. eg: If a part of the weighting for recommendation is that people you follow who regularly DM other people you follow should be weighted higher, doing it client side would allow you to see other people's private DM information.
I do worry about the political implications of it though. People choosing to subscribe to only their world view would create an impenetrable echo chamber. At least now, there is some crossover. If people chose algorithms that avoided the other side of a debate, it would make things worse. Furthermore, you could have outside influences tricking people into certain algorithms to secure a populace with their beliefs.
1) Remove the recommendation part. Go back to a simpler version of Twitter. 2) Still give the option for a recommendation system but, somehow, open the code, maybe some version of the data and publish documentation (like papers or something) detailing the training process of the current running version.
than build a recommender system marketplace? Also, in what sense would this be different from allowing third party Twitter clients + opening a ritcher API?
The ultimate idea of using ML is to "automatically" build the recommender system the user likes (measuring this with some particular metric like online time or retention) the most and also automatically adapt it as his/her preferences change. The problem to me is more the metrics chosen to be optimized.
However, I believe that in the end, and in order to be profitable, user retention and time on the platform will still be pursued. It doesn't seem like an easy fix to me.
Regarding the "free speech" part, I'm not an expert, but I'd say (and after having watched the TED interview) that countries' legislations will considerably constraint this.
I love the idea of a true free platform tho
Just the data and code needed to make the “algo” scale to what it is means there is no good way to “open it up”.
But let’s say, and why not, that it actually was opened.
I can see it now, everyone and their mother would be recommending changes to it. People would want it at the extremes, or tweaked to just not recommend their pet peeves.
And if they did not get their way, they would go off in a huff and threaten to join another service.
Years ago I ran a big enough service that had millions and millions of users every month. A member of the military came on and demanded we allow him to do something that was against TOS.
He complained that this was what he fought in Iraq for. So he could write anything he wanted anywhere.
America is certainly not free and it’s time we stop giving our content to services like Twitter and Facebook if we believe that.
I think the solution to this is to let people switch between different recommendation algorithms, a stream of chronological tweets, most liked among recent posted, one powered by ML, etc. I don't see any reason the implementations behind this, or even Twitter more broadly, could not be open source. There are many other open source projects, like Signal that manage this just fine. And for a while, even Reddit used to be open source. It's definitely do-able.
if (tweet.from in bad_users): score *= 0.8
Do they have the levers in place to manipulate results or is it a clear objective scoring system that shows what you see on your feed and search results? I believe they have something like this. To what extent its being manipulated is a separate question.
This approach only works if auditability is mandated to also be a property of the model, which typically reduces model complexity and accuracy too. We do this with credit scoring models semi-successfully.
Is "security through on obscurity" necessary for automated content moderation?
Would an open algorithm be trigially gamed by spammers as they can now test offline exactly how their posts will be ranked/promoted?
My gut says yes but I'm not an expert in this area. Curious if anyone has a theory or idea on if an open moderation algorithm could work.
SpamAssassin exists and is open to moderate success. But is that just because it's use is not widespread enough to bother to test your spam mails against it?
If every email account in the world was covered by SpamAssassin, what would spam look like, and how much would make it through.
I want to read stuff from people I follow in more or less chronological order. Sure if something has a 1k likes and was posted a few hours ago and I hadn't seen it, show me that. There are simple formulas for time based rank that are out there.
But I regularly see twitter suppressing posts from people I follow. I won't see someone tweet for a few weeks and I check their account and I see they've been tweeting this whole time. It's wrong and annoying to me as a user.
I think it would be better if it was an informed argument.
However, open sourcing any and all manual interventions over the algorithm + the guidelines used for evaluation and/or labeling (if any is done), would help to build a little bit of trust.
Not that much though, but it would be a start.
What would you say about opening up every poster who’s been blocked and exactly how and for reason (or keywords) they’ve been blocked?
How about opening up what keywords trigger mail to go to spam vs inbox for email providers? It’s going to be very valuable for someone to know how spam filtering works if their delivery rate doubles!
Stumbled on this idea on https://twitter.com/nbashaw/status/1515054551371378688
Soybean oil, sweet relish (cucumber, glucose-fructose, sugar, vinegar, salt, xanthan gum, calcium chloride, natural flavour), water, vinegar, egg yolk, onion powder, spices, salt, propylene glycol alginate, colour, sugar, garlic powder, hydrolyzed (corn, soy, wheat) proteins. CONTAINS: Soy, Wheat, Egg, Mustard.
-- https://www.mcdonalds.com/ca/downloads/IngredientslistCA_EN....
And there are strict rules about how that ingredients list must be constructed:
"Ingredients must be declared by their common name in descending order of their proportion by weight of a prepackaged product. The order must be the order or percentage of the ingredients before they are combined to form the prepackaged product. In other words, based on what was added to the mixing bowl"
and on and on for many thousands of words, https://inspection.canada.ca/food-label-requirements/labelli...
No, that's not the full recipe, but it is information a company might not want to disclose, but is required to disclose because it affects the health and safety of consumers.
Another way of thinking about it, is that if someone could see what features were most important for the ranking on some site, then they could start to optimize for those, breaking the usefulness of that feature. One obvious example of this is "Please remember to like, comment, and subscribe" on YouTube.
"The choice of which algorithm to use (or not) should be open to everyone" - Jack Dorsey
On the contrary, Jack knows what he's talking about and wants this because it would allow him to abdicate responsibility for Twitter platform moderation.
I don't care about the actual algorithm.
Chronological presentation is less useful for serving advertisements though, so I don’t expect it to show up very often.
A screen full of this feed would just be tweets made at the moment the data for the screen refresh was fetched.
Is that what you or anyone wants?
It's not a global feed of every twitter in existence, it's a feed of accounts that you proactively choose to follow.
We have seen this with Facebook as well with Facebook charging 3 times higher for ads from the opposition parties in a bid to influence the elections: https://www.aljazeera.com/economy/2022/3/16/facebook-charged...
Not sure why this is getting downvoted. It's not crazy. Even Jack is pushing for it: "The choice of which algorithm to use (or not) should be open to everyone"
edit: complexity lol
I don't think "marketplace of algorithms" necessarily mean people need to be able to make arbitrary database queries either, but seems besides the point.
Having their own DSL for sorting timelines can also solve that problem nicely and within whatever performance requirements they would end up with.
You mean Turing-complete?
1. Absolutely incredibly complex architecture. The recommendation engines are not just single libraries that can be easily open sourced or transformed to a plugin system.
2. Recommendation engines exist to push up relevant metrics. These are either clear wins for users (like reported satisfaction), mixed wins for users and businesses (content engagement), or clear wins for businesses (ad engagement). Most businesses aren't thrilled about subbing in systems that degrade metrics.
How does an algorithm "that attempts to prioritize nuanced conversations about important topics" work? (Or "to find mind-expanding threads", for "savage dunks" or "thirst traps of hot new snax"? -- other examples from the post.)
I suppose you spend some time with existing tweets and ML and develop a model that can produce a score of some sort on these concepts for a tweet, and then run every tweet through the model and present the high-scoring ones. Of course, you can't just look at individual tweet, which in isolation don't mean much, but also at how it fits into a conversation. (For that matter, I'd be interested to see what a model of a conversation is.)
Sounds expensive and quite possibly not accurate enough to be worthwhile.
It also seems like a major strategic decision to give third parties every tweet by everyone. It seems like the business changes from being what twitter is now to a message routing backend, and these third parties become what twitter used to be. That's a fundamental shift, that probably devalues the company by an order of magnitude, since they would be dissolving the valuable thing they have -- the social network they have.
Just doesn't make any sense to me.
if(tweet.author.affiliation == "democrat"){
promoteTweet(tweet);
} else {
shadowBan(tweet.author);
}Related: yes, I do support interoperability requirements between platforms. No, that still doesn't mean you get to blast your opinion all over the internet without hitting a roadbump every now and then.
Now, many people want them to stop doing that. Which they decline, since they KNOW what will be happening in that case.
So, if you think you can do a better platform, while disregarding the (minimal) lessons learned from Twitter (or Reddit, or...), go ahead! You will fail, not because "criticism of the original platform is off-limits", but because it's a well-known anti-pattern.
You’re right about human judgement but that’s not the topic. The central point, I repeat emphatically, is about transparency, not governance.
Twitter can continue exactly the same way but just be transparent. The intense pushback is because they’ve holed themselves into an untenable position? Not sure why people are so against transparency. Maybe they lied in congressional testimonies?
No it can't. If the algorithm was transparent the only thing you'd see is spammers who have put tons of resources into figuring out the exactly optimal way to maximize engagement. Grassroots engagement would be impossible.
Also, Twitter's spam control has been objectively bad.
https://twitter.com/paulg/status/1487022342630957062?lang=en
People think that the entire platform has been hijacked by left-wing / progressives and the reason for lack of transparency is more insiduos than "spam". For example, being liable for what they told Congress.
Public knowledge of what "X" is doesn't really help, I think, other than to aid spammers? And a requirement to "talk to a human" upon hitting X would surely immediately degrade into "Google has reviewed your appeal and has determined that the infinite block of your account remains in effect. There is no further appeal"?
That secret sauce is ripe for manipulation and extremely powerful.
Against combating spam - I mean, isn't this how something gets stronger? HN has a strong view that open source software is more secure because it gets hardened through exposure, not through obfuscation.
Where did this sentiment originate from? I never heard of it before and all of a sudden in the last few years I hear so many people parroting it. Why is it that all these people were silent for so long and now they're yelling in unison about how bad Twitter as a megaphone is?
Before for that, other “movements” that had bridged the gap between real and online worlds were celebrated.
Arab Spring circa 2012, is a particular good example.
I don’t want them arbitrarily hidden from my timeline by an algorithm. Twitter offers a chronological timeline but has repeatedly reset my user preference for it. If there weren’t third party applications that respected my preference I definitely would not be using it anymore.
It will be so satisfying for Musk to buy Twitter, open it up completely, and then be able to use this argument in reverse.