Marginalia – A search engine that prioritizes non-commercial content
marginalia-search.com
marginalia-search.com
Initially Marginalia used an interesting variant of PageRank discussed in the original paper, called Personal Pagerank.[1] Currently pages are ranked with BM25.
I think Personalized PageRank is still used for a new feature of Marginalia which is ranking pages based on similarity. I think this is already integrated into the website but there used to only be this testing page: https://explore2.marginalia.nu/
In any case I have a lot of respect for the creator. Marginalia has seen a lot of growth and it's been interesting reading the blogposts.[2]
[1]: https://www.marginalia.nu/log/26-personalized-pagerank/ [2]: https://www.marginalia.nu/log/
on the contrary i thought any essay, paper or report written for uni or school can be considered for publication in some form. not everything is worth publishing of course, but if it is, talk to your supervisors about it.
Explore2 and the website discovery tools now built into the search engine are using cosine similarity of the incident link vectors. I wrote a blog post about the technique called "Creepy Website Similarity" available :-) https://www.marginalia.nu/log/69-creepy-website-similarity/
* I couldn't find the random button. But, I used the old site to find the correct URL. [0]
* The random URL on the old site is broken. [1]
It's still the same search engine :-)
(Amazing what you built, btw)
If I ever do find one that's appropriate I might set up a CNAME record to point to marginalia-search.com
I've found over the years plenty of <= 5 character domains that have a meaningful name just from using their search filters.
(I'm not affiliated in anyway to them)
How often does yours (re)crawl its indexed sites?
It's more of a way to find interesting things on the "small web".
Please keep that in mind when you check this out.
(I'm not the creator)
That is correct.
But I used[1] to find that in certain niches (Linux and open source software, history) I'd often get significantly better results with Marginalia than with Google or DDG.
This seems to be related both to
- the input stupidifier that seems to be in use with mainstream search engines (if I search for anything unusual it assumes I misspelled or mis-remembered and replace my search with something I didn't search for before sending it to the backend)
- and mainstream engines going out of their way to prefer corporate media over original, authorative content
> It's more of a way to find interesting things on the "small web".
Yes, he keeps saying that, but it was[2] still not just interesting but useful for me.
[1][2]: past tense because I switched to Kagi almost 3 years ago and now I don't have to maintain all the hacks (like separate search engines for separate niche topics) that I used to have. Full disclaimer: I know certain other people frequently want to contest this, saying they get great results with Google or bad results with Kagi, to which all I have to say is I have documented history of Google consistently failing even simple queries going more than a decade back and if Google works for you, more power to you, but it hasn't work reliably for me for over a decade and I am fed up.
If you have spare dollars but not time, you can also contribute to the war chest: https://about.marginalia-search.com/article/supporting/
<https://duckduckgo.com/bangs?q=marginalia>
I've submitted it as a suggestion.
To the author, thank you so much for this trip to the glory days of the Internet.
Marginalia seems to work okay-ish if you want to learn about something, but not if you want to find something.
If you want to read what people think about Scott Alexander's Substack, you will get some decent results, but not if you want to find the newsletter itself.
Phrase matching in Marginalia Search - https://news.ycombinator.com/item?id=41696046 - Sept 2024 (24 comments)
Marginalia: 3 Years - https://news.ycombinator.com/item?id=39501061 - Feb 2024 (44 comments)
Interview with Viktor Lofgren from Marginalia Search - https://news.ycombinator.com/item?id=38470832 - Nov 2023 (21 comments)
Moving Marginalia to a new server - https://news.ycombinator.com/item?id=37800753 - Oct 2023 (39 comments)
Marginalia.nu API - https://news.ycombinator.com/item?id=35871186 - May 2023 (22 comments)
Marginalia: DIY search engine that focuses on non-commercial content - https://news.ycombinator.com/item?id=35611923 - April 2023 (193 comments)
Marginalia Search has received an NLNet grant - https://news.ycombinator.com/item?id=34945541 - Feb 2023 (17 comments)
A Theoretical Justification (2021) - https://news.ycombinator.com/item?id=32586273 - Aug 2022 (22 comments)
The Evolution of Marginalia's Crawling - https://news.ycombinator.com/item?id=32565052 - Aug 2022 (22 comments)
Marginalia Goes Open Source - https://news.ycombinator.com/item?id=31536626 - May 2022 (72 comments)
Uncertain Future for Marginalia Search - https://news.ycombinator.com/item?id=31200319 - April 2022 (37 comments)
Marginalia Search: 1 Year - https://news.ycombinator.com/item?id=30823481 - March 2022 (29 comments)
Show HN: Marginalia – Exploration Mode - https://news.ycombinator.com/item?id=30047455 - Jan 2022 (53 comments)
A search engine that favors text-heavy sites and punishes modern web design - https://news.ycombinator.com/item?id=28550764 - Sept 2021 (717 comments)
I imagined my ad-free, www-free, firstname-lastname domain would be welcome on marginalia, but it seems deranked just like what Google has inexplicably done since I bought the domain. Despite following best practices and webmaster console recommendations. A squatter owned the domain for some time before 2010, but there's no evidence of interesting nefarious activity.
My first name-lastname GitHub account, linked from my site homepage, has a noteworthy number of open-source contributions to all sorts of projects.
P.s. marginalia_nu: Your Java and architectural design is beautiful, you're a true craftsman. <3 Best in show.
https://github.com/MarginaliaSearch/submit-site-to-marginali...
Completely understand that you'd be overwhelmed by bots, that personally curating sites is beyond human effort, and that you want to be choosy.
But I've a site that is clearly non-commercial, definitely "marginal" interest, full of completely original human-generated content, yet sadly I've never seen a hit from Marginalia in the logs despite trying to get listed.
Are you excluding certain content? Political? Anti-Bigtech? Sites "related to hacking"?
Either way, I'd love to investigate to see what's going on. Mind sharing the domain name? Either here, or if you don't want to doxx yourself, email contact@marginalia-search.com
The internet is kind of like a fire. Most things go away but a few things miraculously survive for fairly unexpected and benign reasons.
> P.s. marginalia_nu: Your Java and architectural design is beautiful, you're a true craftsman. <3 Best in show
Well like any project this size (especially when it's a one man show) it has its warts and the design isn't entirely uncontroversial[1], but I'm pretty happy with it.
[1] Last time I posted about it someone likened it to https://www.youtube.com/watch?v=y8OnoxKotPQ , not entirely unfairly, though it does look like it does for very solid reasons that go beyond "netflix uses microservices", which is usually where things don't work out so well.
All of it feels so human and is a joy to view, in contrast to the raw garbage Google spews at me whenever I bother to use it.
Likely because of the enshitification of google.
But what is most on my mind, is how the enshitification of google might have affected the global economy.....
Allow me to try and explain. Market efficiency is deeply tied to market information. And early google was fantastic at providing accurate and valuable information. It's hard to put into words, and I am not aware of any studies, but it felt like it sped up and/or made the whole economy more efficient.
And today it feels like that has been lost. And I wonder if there could ever be formal accounting to detect how much peak google contributed to.... not so much GDP, but how the economy functions. And compare to today.
Going on for well over a decade we've been putting far more slop on the web than useful content. It really adds up over time.
I also feel this is largely why the GPTs have been so successful. They're if nothing else a lot less annoying than the web search experience has become.
The result is current web search returns static output from smaller models in the form of affiliate marketing listicles. Using a larger LLM directly on your query is simply better. There is a semantic/vector search happening with them such that there might be relevant terms or concepts even in a false answer. That's a step forward towards real information at least. That doesn't happen with the listicles.
Hi again! Lots of respect for your work.
This thing I considered for a while but I am now certain that it cannot be the whole explanation.
Why? Because for me Kagi has consistently outperformed Google to the point where it feels like having Google 2009 back, just with some extra improvements: It finds the things I ask for and prioritise in a good way so I rarely have to look for page 2 or anything.
I'm not opposed to seeing a return of the integration though it probably could do with some more design work.