A new approach to domain ranking
marginalia.nu
marginalia.nu
(hold my beer as I DDOS my own website by offering multi-gigabyte downloads on the HN front page ;-)
(I did a quick `apt search` to see if something like that is already available, but didn't find anything.)
This way I wasn't just able to get much better results searching than with PageRank but there was another benefit byproduct of this approach in that you could cluster the results and choose a distinct separate cluster for each subsequent result. With google you would not just get bad results at number 1 but results 1-20 would be near duplicates of just a few distinct efforts.
Unfortunately I was a terrible software engineer back then and had much to learn about making a product.
A dense (bitmap) representation of matrix wouldn't fit in memory, would require about a PB of RAM unless my napkin math is off. The cardinality of this dataset is in the 100s of millions.
(An additional detail is I'm actually using a tiny fixed width bloom filter to make it go even faster)
This makes me feel in the old open web again.
I have actually been developing something like that, but it does more, including down ranking certain categories of sites that contain unnecessary filler, such as some recipe sites.
You can poke around in the result valuation code here: https://github.com/MarginaliaSearch/MarginaliaSearch/blob/ma...
https://explore2.marginalia.nu/search?domain=news.ycombinato...
This is such a great idea, often when I find a small blog or site I want more of it! This is the perfect tool to discover that. It’s a clear and straightforward idea in retrospect, as all really great ideas tend to be!
https://web.archive.org/web/20230217165734/https://www.margi...
The only thing different is the domain names, and those character strings themselves are more than 42% similar.
I am newbie in SEO. I would grately appreciate if marginalia provided clean readme about it, about their algorithm.
At marginalia search front page we have access to search keywords, page algorithm is important enough to be at least discussed on layman terms.
How to optimize page, so it could have a high ranking?
I undestand this could be in the code documentation, but I have not yet checked it, sorry.
> This new approach seems remarkably resistant to existing pagerank manipulation techniques
I am writing my own web scraper. That is why I am in fact interested in this topic at all.
To distinguish poor pages from better I check HTML pages. I think all scrapers need to do that. I rank pages higher if they contain valid titles, og: fields, etc. Etc.
There is nothing wrong with checking it and asking for what can I do to make my site more scrap friendly.
Thanks,
So to manipulate the algorithm, you'd need to find an important website, and then find a way of making changes to all the websites that link to that website to add a link to your own website.
This could be helpful in the short-term, but I'm skeptical long-term as it'll become just as gamed.
This algorithm uses the same method to calculate an eigenvector in an embedding space based on the similarity of the incident vectors of the link graph.
Regardless, well done!