Marginalia: 3 Years
marginalia.nu
marginalia.nu
Now do that on Marginalia; you find https://www.scilab.org within the first 10 results, which is an open-source numerical solver software, which get's me to code, which gets me to examples to use.
To be nuanced, I could change my search to "open-source Range-Kutta numerical solver examples" or something better, but why? Give me the weird deeply technical stuff first.
Maybe more a HN example; when I wanted to learn about load balancers, just search "load balancers."
Google; lots of SEO crap (with soooo many ads), youtube videos?, and AWS commercials at the top. No idea, where I'm going.
Marginalia; a linux wiki, nginx (official documentation), a couple blogs by professionals on the topic. yeah, there are some things in here that ain't great.
But If I compare apples to apples the first ten results are just so much better I'd say 8/10 in Marginalia for this example are great to good, while the first 10 (things I could click on) are by companies that don't teach me anything or have articles full of ads.
The recipe filter is approaching something I'd want to explore further, to be able to provide contextual information outside of the search query.
For signed in accounts (which is pretty much is ~3bn Android and/or Chrome users), Google can predict what the user might prefer and yet...
We've already constrained "what we're looking for" to be "niche expert pages" further up thread. If we're seeing niche expert pages even for generic search results, that's probably a good indication that the search engine behaves the way RandomWorker is describing
Google was ~ top 60, which for such a generic term seems fine, not much scrolling down
@unpopularop cant find "all quiet on the western front book movie differences". well you couldn't do that with AltaVista either in 1998.
however if you just type "all quiet on the western front" you get a ton of niche obscure sites talking about it. literally someone's personal blog page.
type in 'polytopes' you get a bunch of universities papers and code sites.
"rust generics" - again, its a bunch of mailing list discussions, blogs, rust discussion groups, personal websites, obscure professional discussions.
this IS how it was back in the day.
my only question is how could this possibly be sustainable financially in the long run.
For now I'm funded by grants and donations, got a few years runway that way.
The actual operational cost is like $100/month for colocation + personal expenses so what money comes in lasts a surprisingly long time. In the future, we'll see. There does seem to be a lot of people that want this type of thing to exist though, so the hope is if I polish it even more, further funding will become available from likeminded people, possibly selling API access to other search engines.
Search is notoriously hard to make money from (outside of ads), though not having a lot of expenses seems like a reasonable path to go.
Universities traditionally have done this sort of thing by playing golf and naming buildings, but I'm sure in the 21st century there are other models. (Fwiw $2k/yr is below a typical golf membership)
A project is usually on the road to success when it starts with a disclaimer like "just a hobby, won't be big and professional like gnu".
I think a larger concern is how you'll address the Bus Factor going forward.
I can't speak to how much energy it is to go from code to serving requests, but FWIW the code is AGPLv3 and seems to be updated regularly https://github.com/MarginaliaSearch/MarginaliaSearch/blob/v2...
But the long term goal is that this is something that's relatively easy to operate and extend.
[1] https://www.youtube.com/watch?v=PNwMkenQQ24 (quick install and demo)
Though I think now there's a bit too much reddit and stackexchange and wikipedia stuff in the default filter.
It’s proving a bit harder than anticipated, not because the software can’t handle it, but because the signal to noise ratio of the web isn’t very good; a huge reason why the search engine works relatively well is because of what it doesn’t index.
india test cricket lowest total > None of the results are good or giving an answer
raid calculator > The results are OK but you still have random noise like a Pokemon save/cheat editor page because it contains the word raid
all quiet on the western front movie book differences > 0 results. Like straight up no hits, an empty page
The search engine has no ambitions to provide a knowledge graph at this point. It's for finding documents on the internet, rather than answering questions. Answering questions is a definitely something one might want, but it often comes at the expense of finding documents.
> raid calculator > The results are OK but you still have random noise like a Pokemon save/cheat editor page?
The pokemon result was discussing an application called "raidcalc". Seems like a good match, given the search engine does not profile you at all and has no clue about what your interests are.
> all quiet on the western front movie book differences
Hmm, I think there's an upper bound on the query length you hit. Could probably remove this, it's a pretty old, an artifact from when the query execution didn't deal with long queries well.
--edit--
Hmm, I increased the limit but they're still kinda not very good. Although this is definitely squarely within the realm of what I'm working on next, which is query understanding and execution.
Right now the search engine doesn't really know how group the terms. Like a human being can see that you'd want
|all quiet on the western front| in a sequence, preferrably in the title or appearing a few times, and 'movie', 'book', and 'differences' should be important to the document, but not necessarily appear in that exact order.
The search engine currently looks for either documents where they all appear in proximity, or all individual words have high tf-idf relevance markers. Not great for this query.
It is arguably a better UI than handing a barrage of words and hoping the engine does the sense-making.
If you were thus inclined, https://gitlab.com/glitchtip/glitchtip#glitchtip is the actual open source Sentry implementation which (as far as I know) would enable gluing https://docs.sentry.io/platforms/javascript/user-feedback/#u... to the search results page (that client-side library is still MIT: https://github.com/getsentry/sentry-javascript/blob/7.102.1/... )
For some reason all this makes me reminisce about Fravia's Searchlores [1] which always felt a bit like if Umberto Eco was interested in computers. And the site felt a bit like the library labyrinth from The Name of The Rose where you'd turn some random corner and find something incredible, only to lose it forever later on :D
[0] http://ts.sesse.net/ [1] https://www.biostatisticien.eu/www.searchlores.org/indexo.ht...
That, and the set isn't very large, just a few thousand. If you randomly pick 25 items from a bag of ~3000 total a bunch of times, chances are relativley high that you're going to see repetition.
https://search.marginalia.nu/search?query=anime&profile=vint...
https://search.marginalia.nu/search?query=anime&profile=tild...
Demo key is always under siege though.