People rely too much on other people's infra and services, which can be decommissioned anytime. The Google Graveyard is real.
People rely too much on other people's infra and services, which can be decommissioned anytime. The Google Graveyard is real.
I just searched for “stackoverflow” and the first result was this: https://www.perl.com/tags/stackoverflow/
The actual Stackoverflow site was ranked way down, below some weird twitter accounts.
It’s cool though, and really fast
You are absolutely right, it is the hardest part!
I feel like Google-style "search" has made people really dumb and unable to help themselves.
Even with PageRank result prioritisation is highly subject to gaming. Raw keyword search is far more so (keyword stuffing and other shenanigans), moreso as any given search engine begins to become popular and catch the attention of publishers.
Google now applies other additional ordering factors as well. And of course has come to dominate SERP results with paid, advertised, listings, which are all but impossible to discern from "organic" search results.
(I've not used Google Web Search as my primary tool for well over a decade, and probably only run a few searches per month. DDG is my primary, though I'll look at a few others including Kagi and Marginalia, though those rarely.)
<https://en.wikipedia.org/wiki/PageRank>
"The anatomy of a large-scale hypertextual Web search engine" (1998) <http://infolab.stanford.edu/pub/papers/google.pdf> (PDF)
Early (1990s) search engines: <https://en.wikipedia.org/wiki/Search_engine#1990s:_Birth_of_...>.
Fair play to them though, it enabled them to build a massive business.
I publish exports of the ones Marginalia is aware of[1] if you want to play with integrating them.
[1] https://downloads.marginalia.nu/exports/ grab 'atags-25-04-20.parquet'
"Affiliation" is a tricky term itself. Content farms were popular in the aughts (though they seem to have largely subsided), firms such as Claria and Gator. There are chumboxes (Outbrain, Taboola), and of course affiliate links (e.g., to Amazon or other shopping sites). SEO manipulation is its own whole universe.
(I'm sure you know far more about this than I do, I'm mostly talking at other readers, and maybe hoping to glean some more wisdom from you ;-)
I've also seen some benefit fingerpinting the network traffic the websites make using a headless browser, to identify which ad networks they load. Very few spam sites have no ads, since there wouldn't be any economy in that.
e.g. https://marginalia-search.com/site/www.salon.com?view=traffi...
The full data set of DOM samples + recorded network traffic are in an enormous sqlite file (400GB+), and I haven't yet worked out any way of distributing the data yet. Though it's in the back of my mind as something I'd like to solve.
I'd also suspect that there are networks / links which are more likely signs of low-value content than others. Off the top of my head, crypto, MLM, known scam/fraud sites, and perhaps share links to certain social networks might be negative indicators.
Have a lil' data explorer for this: https://explore2.marginalia.nu/
Quite a lot of dead links in the dataset, but it's still useful.
It’s also why it is so hard to compete with Google. You guys are talking about techniques for analyzing the corpus of the search index. Google does that and has a direct view into how millions of people interact with it.
The Chrome iOS app still knows every url visited, duration, scroll depth, etc.
There is a native Chrome app on iOS. It gets all the same url visit data as Chrome on other platforms.
Apple blocks 3rd party renderers and JS engines on iOS to protect its App Store from competition that might deliver software and content through other channels that they don't take a cut of.
Indexing is a nice compact CS problem; not completely simple for huge datasets like the entire internet, but well-formed. Ranking is the thing that makes a search engine valuable. Especially when faced with people trying to game it with SEO.
amazing, for real.
everything i’ve read and heard about the good internet is that it was good because sooooo many of the people did stuff for exactly that, fun.
i’ve spent some time reading through some of the old email lists from earlier internet folks, they predicted exactly what weve turned this into. reading the resistance against early adoption of cookies is incredible to see how prescient some of those people were. truly incredible.
keep having fun with it, i think it’s our only way out of whatever this thing is we have now.
Can you talk a bit about your stack? The about page mentions grep but I'd assume it's a bit more complex than having a large volume and running grep over it ;)
Is it some sort of custom database or did you keep it simple? Do you also run a crawler?
There's a reason Google became so popular as quickly as it did. It's even harder to compete in this space nowadays, as the volume of junk and SEO spam is many orders of magnitude worse as a percentage of the corpus than it was back then.
It's driven by my own personal nostalgia for the early Internet, and to find interesting hidden corners of the Internet that are becoming increasingly hard to find on Google after you wade through all of the sponsored results and spam in the first few pages...
[0] https://www.site.uottawa.ca/~stan/csi5389/readings/google.pd...
It is intended, that the page currently shows a link to the wordpress login?
https://github.com/rumca-js/Internet-Places-Database
Demo for most important ones https://rumca-js.github.io/search
> I am trying to say is many servers ignore "Accept-Language"
I wouldn't have expected that to be a hard rule, more like if there are multiple pages to return to have a factor, which one the user most likely wants.