SearXNG is a free internet metasearch engine
github.com
github.com
It hooks into your browser to give you an augmented experience. The UI is pretty simple (think 1997 era google but without CSS haha), and we don't do anything super complex with search (but could in future), but it works not bad. Check it out!!!
https://github.com/dosyago/DownloadNet
Oh, it also makes your content (again either everything you browsed or only what you booked) available offline. So if you work on an oil rig, or shipping, or long haul freight, can be a good way to browse as normal but save yer satellite bandwidth!!!
I just tried it out and it seems to be tied to Chrome. Since I use Firefox and Chromium as my daily drivers this does not work for my case. I understand that they probably rely on some Chrome internals to dig through the content, a SOCKS Proxy approach would have worked better and would have no need to switch between a "save" and "serve" mode. But then again I was only scraping the top of it because of the lack of browser support. Will keep an eye on this one though!
One .py file. Only one dependency (urllib3).n with a little love the concept could become a full transparent proxy.
i don't need anything else archiving anything related to my internet browsing except for my human brain. and yes, that's just me...
but how is the shameless plug of this not just therefore off-topic but diametrically-opposed-to-total-personal-privacy tool appropriate here?
I get if you’re not interested, but I imagine people interested in locally hosted search-related solutions, may be.
Your view is probably more personal and hard to support in general given this, and given the comment’s position and votes indicating at least some people are interested.
I totally understand why you wouldn’t want your browsing history archived anywhere. But that is what search engines do somewhat. It’s okay, everyone’s different.
this tool is not relevant here.
> this tool is not relevant here
No it’s relevant. You don’t think self hosted and offline is private?
there is no man in the desert. And no man needs nothing.
Tho I prefer the west coast of Zaire or Suid-Afrika myself.
mental privacy, huh, nickburns? That’s an interesting concept.
you made your plug under the guise of asking if anyone had interest, i offered mine, and now i think we're done here, keepamovin.but if i may just say, privacy as a concept for a truly egalitarian society is something very near and critical in my opinion. marketing, on the other hand, is not.
good day to you, sir.
One of us, apparently.
> SearXNG protects the privacy of its users in multiple ways regardless of the type of the instance (private, public). Removal of private data from search requests comes in three forms:
> 1. removal of private data from requests going to search services
> 2. not forwarding anything from a third party services through search services (e.g. advertisement)
> 3. removal of private data from requests going to the result pages
From: https://docs.searxng.org/own-instance.html#how-does-searxng-...
The docs mention a caveat below at "What are the consequences of using public instances?":
> If someone uses a public instance, they have to trust the administrator of that instance. This means that the user of the public instance does not know whether their requests are logged, aggregated and sent or sold to a third party.
a proxy (or proxies) and how they can shield but one or many of ' your' IP addresses throughout an egress packet's many hops (and from who or what destination it or those addresses can be shielded) is a pretty advanced concept when you think about it.
not to mention that, at this point, bare source IP address is a pretty dilute tracker compared to other current methods of identity profiling or traffic fingerprinting.
nice succint correction on your part regardless.
a few examples of a self-hosted design that would not, include policy-based routing over a VPN with one or multiple tunneled hops, or through another external proxy. (and then there's also that 'onion' routing 'protocol' there—but i'm not clear if/how that integrates with clearnet destinations like publicly-accessible search engines if at all.)
but the idea is not necessarily anonymity so much as privacy by foiling the creation of any even somewhat accurate marketing/data profile derived from 'your search.'
https://felladrin-minisearch.hf.space/EDIT: Looks like there's already an open issue:
I understand Kagi is generally reputable, but I like the idea of a self-hosted alternative where you're in full control.
My guess is that if they are found to do so, then they open themselves up to lawsuits. Not collecting data isn't merely a perk - it's practically the reason Kagi exists.
If you don't have any logs you can just always say the princess is in another castle, since you can't provide data that doesn't exist.
If on the other hand you do have the requested information, you need to determine the validity of the request, and then extract the data; or refuse to comply and possibly put yourself at legal risk. For a smaller business that's probably a can of worms you'd rather avoid opening.
The intent of SearXNG is to be stateless (with no sessions on the server) and to work without JavaScript.
However, this approach limits certain features because of the restricted size of cookies (and other forms of browser storage require JavaScript).
I'm picturing more of an instance-wide configuration of domain blocks for a private, single-user, self-hosted instance. But I understand this may not be the intended use of the project.
https://github.com/searxng/searxng/blob/f1a148f53e9fbd10e95b... https://github.com/searxng/searxng/discussions/970
Kagi has this feature built in and it is a good user experience.
You can also use the uBlacklist browser plugin. My problem with that is that is slows everything down. I am not certain but I think all the works is done after the search is complete. That it filter the actual result. The two above limit it from ever being part of the result.
If you delegate queries to e.g. google or bing at that rate, you'll be ip blocked in a heartbeat.
btw thank you for Marginalia! The spirit of the small web is very important to me.
Like I don't mind automated access to my search engine, I even offer a public API to the effect, that you can in fact hook into SearXNG. What I mind is when one jabroni with a botnet decides their search traffic is more important than everyone else's and grabs all the compute for himself via a sybil attack.
Honestly, I just use Kagi. Though I need to find some way to limit my searches to 300 per month.
although existing searx instances have been run for years and they don't seem to be dropping like flies...
For personal use, you can run it directly on your machine or access over VPN. Queries to upstream search engines can be forwarded over proxies or VPNs as you see fit. Some work fine over tor and some can go over commercial or DIY tunnels.
Config is in a git repo I give access to if requested. One of the technical users modified it to keep pretty minimal logs. I guess they are trusting me to actually use that config but trust is pretty high in the group so not really an issue.
If you run it locally, and only you use it, then you won't get blocked - a given search engine will see about the same number of requests as if you used it directly.
Add a few house members and you'll still be fine.
(I ran the original searx for a year or two locally - no issues at all).
I use this all the time. A downside is that sometimes you land on an instance that doesn't provide any results or gives you really poor ones. This has been happening less frequently recently.
Running it on a machine that also does NAT for many other machines helps to prevent getting blocked by upstream search engines like DuckDuckGo. It'd be good if access to certain upstream search engines could be sent through, say, a proxy set up elsewhere to prevent this very, very common problem, if you can't run it from an IP used for other things.
I'd like to figure out how to have a mode where my search is 100% literal - where every word I type must be in the search results exactly as I type them. Perhaps that's the equivalent of putting a "+" in front of each word, and putting each word in quotes? It's annoying that my words are constantly getting changed for me because there aren't many results, which I expressly don't want.
Like mentioned elsewhere, I want to be able to explicitly exclude certain domains. I get that SearXNG wants to be stateless, but I could either configure a separate URL for it or simply configure it for all searches. For instance, if I search for a PDF manual for something, I never, ever, ever want to see anything from "manualslib.com" and sites like it.
Other than these things which'd be nice to address, I'd say running SearXNG and encouraging people to use it instead of Google has worked quite well :)
Anyway, if I'm looking for some topic I believe google would be known to filter heavily, or something esoteric, I take a look at presearch to get a second opinion. I'd also love to see archive.org do something similar, archive.org has an amazing collection of data, poorly indexed and poorly searchable.
Apparently Dogpile still exists, didn't expect that: https://en.wikipedia.org/wiki/Dogpile
I'm not sure if this person is still active on HN, but I'm really curious about the results.
There's quite some similarity between the CH and the X sound in English.
But, as this is HN probably someone with a PhD in comparative phonetics will explain why this is a common and infuriating misunderstanding of layfolken.
Hah. FWIW, it's a fork of searX. https://github.com/searx/searx
https://spanish.stackexchange.com/questions/16203/use-of-x-i....
Error! Engines cannot retrieve results:
brave ( Suspended: too many requests )
google ( Suspended: too many requests )
qwant ( server API error )
My bet is that Google will become "Google TV" and search won't be possible. They will just show you what they want. They'll probably frame it as "AI knows what you want to see".
Maybe they should ban Google instead of TikTok (I don't use either though).
- name: yacy
engine: yacy
categories: general
search_type: text
base_url: https://yacy.searchlab.eu
shortcut: ya
disabled: true
# required if you aren't using HTTPS for your local yacy instance
# https://docs.searxng.org/dev/engines/online/yacy.html
# enable_http: true
# timeout: 3.0
# search_mode: 'global'
Change 'disabled' to 'false' and point it at whatever YaCy instance you want to use. It can use the 'general' and 'images' categories.