Also curious to know the extent to which Google de-lists things?
Long term, perhaps a decentralised search engine could get around de-listing and provide a more reliable and rigorous search experience.
Also curious to know the extent to which Google de-lists things?
Long term, perhaps a decentralised search engine could get around de-listing and provide a more reliable and rigorous search experience.
https://twitter.com/yegg/status/1501716484761997318?s=20&t=9...
Before someone chimes in, yes I understand the humanitarian perspective. That's not my point. My point is that DDG is not neutral, and is politically biased.
Ultimately what I'm getting at is: There's no market for DDG. Use Google for biased searches, and use other search engines that are not biased(which excludes DDG) for unbiased searches.
What are the use cases for DDG?
Maybe i was only the few that cared more about the filter bubble angle and less the "we're selling bottled privacy™" angle [1] and am not interested now that `yegg` has clearly reneged upon the former
[0] https://techcrunch.com/2011/06/20/duckduckgo-to-google-bing-...
[1] https://pictobar.tumblr.com/post/63785124046/the-banality-of...
It's not to me, when was the last time you used it? I haven't used Google in around 4 years now and I'm getting on just fine, majority of searches answered on first page.
If anything I now toggle between Duckduck and Ecosia as Ecosia still isn't 100% up to scratch (frequent 500 errors, slow, results are bad) but I like the idea of my searches planting trees.
When I want unbiased searches I've been using Kagi but more are popping up. they approach search differently so it's useful when google feels to "sanitized" for certain searches
I'd expect that all sites wanting to draw traffic would attempt to grab the reins of the search engine to point toward themselves, and the result would be search results ordered by rein-grabbing power.
Not that centralized search engines are immune to this; they're almost as vulnerable (seeing as sponsored search results exist) but the maintainer at least has to balance that with the utility of the search engine overall, to prevent the search engine from falling out of favour.
With a decentralized engine, parties that have deeply invested in manipulating the results will still want the engine to be popular too, but I'm not sure how you resolve the prisoners dilemma there as a whole.
> I'd expect that all sites wanting to draw traffic would attempt to grab the reins of the search engine to point toward themselves, and the result would be search results ordered by rein-grabbing power.
I would venture to say that a combination of allow-lists and block-lists from trusted parties, ranked using some kind of distributed web-of-trust system would work reasonably well.
The basic idea is when you rate something, your client also looks up in the DHT other people who have rated the same content with similar ratings. Your client then pulls the latest ratings collections from those people, and computes the cosign distance between your ratings and their ratings (over the intersection of content that both of you have rated). Periodically, your client signs and publishes an updated ratings document, where the rating for other raters is the cosign distance. The cosign distance, the size of the ratings intersection set, and maybe some other factors go into deciding which raters get published out in your ratings update.
When you query for the rating for a given piece of content, your client grabs the list of ratings for that content from the DHT. It then pulls the latest ratings published by those raters, computes cosign distance, and then does something similar to Djikstra's shortest-path algorithm to recursively search the DHT using these cosign distances as weights. In general, the DHT wouldn't have many signatures stored under the content's hash, but by recursively following the graph of other raters, your client hopefully finds other raters that rate things similarly to you and have rated this content. The path weight to a given rater is the product of cosign distances, and so by using a priority queue for querying, you get something close to a breadth-first search of the ratings graph. Once your client has accumulated enough weight of ratings for the given content, it stops and shows you the weighted average of the ratings (and maybe the weighted std. dev. is displayed as a confidence score to power users who have enabled it).
Presumably, the UI for the ratings system maps 0 to 5 stars to 0.0 to 1.0 (probably not linearly, more likely the client locally keeps a histogram of the user's ratings and then maps the star rating back to a percentile rating), and the "spam" button rates the content as -1.0.
The tricks come down to the metrics used for how the DHT decides priorities for cache eviction of the per-content ratings and also the per-rater ratings. You don't want spammers or other censors to be able to easily force cache eviction. Getting cache eviction metrics right is the key to having the system scale well while also preventing spammers/censors from evicting the most useful sets of ratings.
Microsoft (and any other big company) has many competing interests other than just being helpful to users.
Are the main ones i) reducing competition and ii) managing their reputation?
If so, the case for a decentralised search engine got stronger.
I don’t think anyone wants to use a search engine that never delists anything. Ransomeware, Markov chain junk, plagiarism. A search engine that never delists anything is useless.
The problem is when delisting is used against the end-user’s interests.
They use Bing as an index, and it was Bing who de-listed it.