Marginalia.nu API
marginalia.nu
marginalia.nu
It's how you use Marginalia if you don't want cloudflare to spy on you. It's also (currently) a tiny bit worse than what you get in the web ui. It doesn't dig as deep into the index to make the API queries snappier, but as a result the results are a bit worse.
I added a temporary API key 'hacker-news-front-page' with a higher rate limit because 'public' gonna give you 503s all day, it's constantly under siege by metasearch engines. That's totally fine but it's not very useful.
e.g. https://api.marginalia.nu/hacker-news-front-page/search/paul...
* hashcash-verification
* manual email-based signup
* domain-based verification
* an ancap dystopia where every query costs an infinitesimal amount of cryptocurrrency
* dunno what else there is, probably more creative solutions
I did add it to another search I run to stop spam and that was very effective.
Maybe the poster needed to summon you. I’ve noticed that whenever the website is mentioned, you appear. No need to do it thrice, either, like other legends.¹
Jokes aside, Marginalia (and your interactions) seem to be well liked on HN, so no surprise someone submits them.
On the design front, I’ll recommend you don’t justify text, especially on the web. Unless carefully done it leads to rivers² and huge whitespace gaps which make text hard to read. Try that page on a phone to see what I mean.
With auto-hyphens I kinda like justified text on the web. It's just that I'd forgotten to set the html language so it didn't work. I literally built and migrated to this new website template yesterday so there's some bugs to iron out.
That said, I'm not sure there is a good solution that doesn't sacrifice some technological sovereignty though. The indie web space has some creative solutions but largely seem to hinge on everyone owning a domain name.
Although on some level I think comment fields tend to promote fairly low quality engagement. The average quality of the emails I get is significantly higher than anything I've ever seen in a comment field. Same with HN discussions.
At least no-one's (yet) asking support type questions through it. :D
I expected this to be about sytse, which is known to show up here on HN if you mention his handle (or gitlab) three times after midnight.
I submitted this entry because your search engine has been interesting to follow. And it’s hacker friendly to have a simple and straightforward API like this!
There's a lot of stuff that could be exposed that would make it much more useful. Tricky part is how to keep it simple.
I have plans to allow customizing the ranking algorithm parameters. Could for example make it possible to add a bias toward recent results (or old results). It could also be fairly trivial to limit the search to a list of domains. But that's probably just going to be a per api key thing, not something that's included in each request.
Isn’t this so incredibly true for most things. Thanks for saying it.
How does one know it's metasearch engines. Which ones are they. Have there been any discussions with them.
If a web search engine is truly noncommercial and one wants as may folks to to use it as possible, is it feasible to let others copy the crawler, the index and front end. It's like a fisherman who gives way his catch. As word spreads, he will be overwhelmed. Unless he teaches others how to do what he does, how to fish.
The problem with Google and all the Google wannabes is that they want all traffic to pass through their servers, for commercial reasons. In the beginning Google claimed their search engine would be in the academic realm. Presumably others could copy it. This turned out to be a lie. Now people cannot not seem to separate the idea of web search engine from potential money printing machine. The fisherman does not want to help anyone else to learn how to fish.
We need a noncommercial web search engine that can be replicated. It does not have to be a Google clone. A great way to stop anyone from working on web search engines is to suggest any new search engine must be as good as or better than Google, i.e., it must be huge, it must be commercial, and it must be a black box. On the contrary, a new noncommercial search engine can have humble beginnings and does not have to be anywhere near Google quality. IMO, there should be noncommercial web search engines with low traffic. They could all be instances of the same crawler, same index and same front end, they could be variations, they could be totally different. But IMO there should not be one set of servers controlled by one entity hoarding 90% of noncommercial search traffic.
SearXNG supports marginalia but uses the demo key. Any instance that enables marginalia support will spam the API with this key by default. I don't really mind, it's just that they need to get a key to be useful. The shared global ratelimit of the public key is (IIRC) 15 queries/minute because it's a demo key that's not intended to actually be used in any applications, just to try out the API. Maybe I should set up a daily rotating key or something to prevent this problem.
I'd be all for having more search engine instances. Problem is you need specialized hardware to run a search engine, even a small one like mine. I've paid $5000 out of my own pocket to build the server I'm using, plus give or take $200/month in network and power.
I'm crazy enough to take that loss. I don't expect many other people will do the same. There's really not much money in running a publicly accessible search engine.
Yeah. Looking over the GitHub issue you raised about it, they didn't seem to really understand the problem before they closed it.
I've just asked a question about it (same thing, different wording), which might work:
https://github.com/MarginaliaSearch/MarginaliaSearch/blob/ma...