Which search engine is the least censored?
michaelsuede.substack.com
michaelsuede.substack.com
I understand there's probably a reason others do it - maybe that that's what's needed to serve a general audience, they're optimising the results for my 70 yr old dad and 10 yr old niece. But that does mean that using them is a frustrating experience for me half of the time, of quoting everything and micromanaging the query terms so that they understand that I do mean this specific thing rather than this overgeneralised idea they seem to parse from it. Yandex was a breath of fresh air in that respect.
If you do a search, click tools, click the "more" dropdown, and choose "Verbatim" then the results are actually pretty close to "don't fucking try to read my mind".
They claim the following about Verbatim (although this info is from 2011, and it may be different now)
----
This verbatim search removes personalized, corrected, suggested, related, and non-inclusive results.
On the verbatim page, users will see only results that:
Include all their search terms.
Match their exact spelling.
Use the same tense (e.g., “is” and “was” will be seen as distinct).
Use the same verb form (e.g., “swimming” and “swim” will be seen as distinct).
Use the same plural vs singular form (e.g., “hat” and “hats” will be seen as distinct).
It's under Tools -> "All results" dropdown for me.
I tried to see if it was possible to bookmark to get a verbatim search directly, and while I couldn't decipher the obfuscated query they seem to be using these days, the overall query length was much shorter than with the original non-verbatim search. I wonder if that has anything to do with the:
> This verbatim search removes personalized, corrected, suggested, related, and non-inclusive results.
claim. The original search had a very long 64 bit encoded parameter, which perhaps had some context information - it seems plausible it had some typing history information (perhaps taken from their Google Docs code), maybe how fast I typed (for error probability), some kind of personally identifying ID about my current session perhaps. Looking at these query strings and how opaque they are, feels a bit creepy that all this information is being sent from my computer without my knowledge or significant control. Makes me glad I rarely use Google these days.
Russia thousands arrested -- should return five results about arrests at recent war protests in Russia
Russian casualties in ukraine -- should return five results listing figures in the thousands.
Maybe you should do a pull request!
Something like ublock lists and p2p sharing in closed communities could provide semantic indexing and gradual manual curation, so with a few thousand people you might get novelty + quality in an engine. If you stick to a small, curated list of sites, you won't get site level novelty, and quality will eventually suffer relative to searches that scan over new things.
Even then, consider reddit - build a reddit search that's as good as Google. That updates once an hour.
Sharing bookmarks isn't feasible when you want the functionality of a general purpose search engine.
"Censored" is definitely the wrong word here; it implies that there is some objective pool of information that we could all be getting to if it weren't for the search engines getting it wrong, and that's just not it.
It's definitely much closer to -- imagine 5 hypothetical bookstores owned by different people. Each has to make decisions on what to choose to carry based on their business interests, and so they do. But there's nothing I would call "censorship" about that.
And if we're deciding it is, then we have to be much more serious and intelligent in what to do about it. I'm thinking something like the EPA or FDA for search engines. If they have to tell us whats in the ketchup, then they have to tell us what's in the "search engine algorithms."
The most widely recognized example is when major search engines delisted pictures and results for the Tiananmen Square massacre last year.
The bookstore analogy would be if you came in and asked for a book on the massacre, and the owner said "sorry, I don't know what you are talking about" when they have a whole shelf in stock.
1. I think that when it comes to "search engine censorship", most people know it's about search engines downranking/hiding certain results, not them literally going around burning books or blocking/taking down websites.
2. The definition for "censorship" is "The use of state or group power to control freedom of expression or press, such as passing laws to prevent media from being published or propagated". When it comes to the internet, getting downranked/hidden from search engines does a pretty good job at preventing such information from being "propagated", even though you could still access it via the deep web. As a thought experiment, suppose the government banned 1984 from being sold/read anywhere, save for you going to the library of congress in person. You can theoretically still read it, but at great difficulty. It's certainly more work than reading harry potter or whatever. Would you say such measures are not "censorship"?
Shouldn't an "objective" search engine return results that are true (that there was no systemic election fraud) instead of results that fit his personal biases?
If I type "the daily stormer" (a very unique name which corresponds to something very specific, i.e. the website for the far-right online newspaper by that name), I expect to find that thing in the first 5 results. That's not the case with Google; it's just not there at all. You have articles from other papers talking about it, you've got a wikipedia page talking about it, you've got sites like the ADL and the SPLC talking about it, but not the actual thing itself. Bing, Yahoo, Yandex all return it as the first result.
When I “search” for something, I am searching for it. I mean, it’s the definition of the word “search” for crying out loud. Somehow over time this idea of a “search engine” got perverted into something that either tries to predict what I don’t know I really want or what someone else thinks I need to see.
Your claim that the search engine should show results which indicate that there was no systematic election fraud, is a wishful thinking that you don’t want people to believe in it, not a good decision making regarding what search engines should display.
What should it do? Article titled "There is no proof that the earth is flat" also satisfies the query completely from a keyword perspective, although not from a semantic. But if there is no actual proof that earth is flat, any article claiming there is would satisfy the query only from a keyword perspective but technically not semantically. It is a good question.
It I mean it would be nice if it surfaced whatever actual report you’re talking about, but maybe it’s just not the best query for that? I’m sure I could find it by clicking through any of the links
Reading that post should tell you how smoothly you're being manipulated.
Now of course I know its been purposely deranked. But couldn’t google claim its not censored if it exists but buried in results?
I want to know what legal basis we have for asserting certain sites must have the highest rank on google for certain searches. I’m assuming there isn’t one.
Really, you consider Putin's stated reasons to be the actual reasons for the invasion? And the linked article in that paragraph is just a pile of misinformation (even calling it subjective would be a complement).
This article is biased, cherry-picked nonsense. Is anyone actually looking at the results of his example queries and evaluating the quality themselves before responding?
To go with your example: the first three results a Google search returns for me on ‘Putin Invasion Speech’ are two NY Times articles analyzing the speech, and then a full translated transcript from Bloomberg. It’s hardly hidden.
On the author’s own example, “Why did Russia invade Ukraine? 2022”, the first article is from the BBC and tries to lay out the full historical context - and mentions Putin’s speech.
The rest are from mainstream news sources of various descriptions, again laying out historical context.
The author may not _like_ that these are trusted, mainstream, popular sources - but that’s what Google has _always_ strived to return. Pagerank was never an ‘unbiased’ thing that searched solely for page content - and, wow, can you imagine the toxic SEO nightmare we’d be in if it was…
Because that's not what I'm searching for. Even if we all agree all the Russian sites are propaganda, I still want to be able to search through it. It's a political decision by Google to filter out these results. It's not algorithmic because no matter how precise your search query is, you will never get results from certain forbidden domains. Google just directly removed rt.com from results recently, so even if you search for "rt.com" you will not see it. I'm sure they have similar filtering based on certain keywords.
Did you try this out?
I specifically didn't call out the complaint about missing results when searching for direct quotes from an article (event thought they are also blatant misinformation), because that indeed is an issue. If I search for some pronoun X then I expect to find a known website with the name X. But if I search for a general term, then the only reasonable way to thing that a search engine can return is websites that fit the consensus view around that topic.
Take a look at who the author of this "comparison" of search engines is. His next post is describing Zelensky as a neo-Nazi. He claims Putin will exit Ukraine once these "Nazis" have been removed. He says he trusts Putin more than Biden and that the US was responsible for the invasion.
https://www.factcheck.org/2022/03/social-media-posts-misrepr...
https://www.usatoday.com/story/news/factcheck/2022/02/25/fac...
https://www.bloomberg.com/news/articles/2022-03-08/china-pus...