Searx – Privacy-respecting metasearch engine
sagrista.info
sagrista.info
Will add SearX now. It seems to provide reasonably good results.
Update: It's on (under 'More Engines').
It would be very interesting if it examined and compared results.
Why? What you described appears to be the safest place on the web.
A lot of servers still use HTTP, for various reasons. There are also some clients that can't use HTTPS.
If the same server has two domains associated with it, does it count twice? Now consider a loadbalancer that points to virtual servers on the same machine. How about subdomains?
In reality as far as privacy goes, the matters are on average opposite to your claim. Most sites that will put your privacy at risk today are using https - I am talking about the vast majority of the commercially operated web today. I know my privacy is much better respected on a plain text (no javascript) site using http then on [insert a top 10k most popular site here] using https.
And for security, if I am not performing for example shopping or entering my billing details anywhere on the site, I do not see how a http site can compromise my security.
I actually prefer deploying http sites for simple test projects where speed is imperative because they are also faster - there is no SSL handshake needed to connect.
This can be prevented by (a) using TLS1.3 to ensure the server certificate that is sent is encrypted and (b) taking steps to plug the SNI leak; popular browsers send SNI to every website, even when it is not required.
Of course, I'm fortunate enough to live in a place where MITM attacks are virtually non-existent, aside from WiFi portals and maybe ISP banners (which I've never experienced.)
I don’t know where you live, but i feel like this is more common and insidious than you think. For instance, in the UK Vodafone (or Three, I don’t remember exactly) would break 100s of our sites by injecting js and tracking pixels into the markup.
Now, with behavioural targeting slowly dying, ad tech businesses talking about fingerprinting as valid alternatives, and contextual targeting on the rise, I can guarantee you that the situation is going to get worse.
When I click on "Try out the YaCy Demo Peer", I get "502 Bad Gateway".
http://www.jaruzel.com/textfiles/Old%20Web%20Info/Internet%2...
Google basically killed almost off of them off :(
It would be great to see some proper competition in the search space, especially around specialist search engines.
Holy Moly, are there really over 100 million podcast episodes out there?
Listen Notes was started in early 2017 as a side project, when there were ~23 million episodes.
I remembered seeing the number of web pages indexed by Google in early 1998 was ~25 million, then I thought that "ok, 23 million episodes might justify the existence of a podcast search engine" :)
nona.de
I agree that the creator could make that a little more clear somewhere on the page.
The comment is not wholly conspiratorial, considering the CIA owned Swiss crypto company: Crypto AG [1]
It's within the realm of possibility that most of these privacy services could be owned by 3-letter agencies or small enough to be coerced into cooperation.
[1] https://www.scmp.com/news/world/europe/article/3050193/crypt...
They can just download the Searx source code; modify it as they see fit, and make it available on a server someplace.
Can you prove that searx.be isn't run by a "3 letter agency"? Can you prove that the source code running at searx.be is the same as on Github?
The point being --- unless you have full access to the server, open source means nothing with regard to privacy and security of any service. It actually means less than nothing --- it means it is super easy to build into a honeypot.
To mitigate server-side code from identifying you, you can consume an instance from Tor. Of course, you could try to do that with any other search engine, but most of the other search engines either block exit nodes or provide incomplete functionality if you disable JS.
It's not perfect, but it may be good enough depending on your threat model.
Personally, I just use a VPN with the "lite" version of DuckDuckGo --- no JS.
And even if they are? As a Canadian or someone who isnt in the USA. What exactly is the point? Wouldn't this effectively be the safest host? CIA/NSA wont be selling your private infos. They wont be sending me to a blacksite because i look at python documentation and youtube chill music.
Why not? They used to sell cocaine after all and your info is probably rather less risky.
It will reveal the operation busting any potential for catching terrorists.
Sort of like how most anti-tracking browser extensions eventually turn out to actually be tracking extensions. Or like how used car dealers that have a name like "honest bob's cheap luxury cars" often turn out to neither be honest, cheap nor luxurious.
Ghostery was more widely reported as Audacity adding telemetry. Everyone who cared knew long before to leave or uninstall it.
Hosts blocking is reliable and I’ve never had a single malicious one with the wide assortment I used. PiHole hasn’t been hijacked either and I think it’s unreasonable to think that no group can make mistakes, faltering can’t ever happen, I really don’t think Adblock Plus was that bad.
If the market wasn’t saturated with methods to block, I would have stuck with them if they were remorseful.
-Sent from my not private Apple device I’ll still use since it’s got a huge userbase on messaging in the US
It's also wise to do due diligence on any company/service where you are revealing sensitive personal information. Traffic coming from Google in 2006, for sensitive medical search queries was a catalysts for us going public in 2006 on our strict no-tracking policy and we maintained that position.
We have yet to be contacted by authorities, but you'll have to trust us on that one for now. Since we don't log any personal or identifying data at all, we would have nothing to share [0]. You can read about our investors on our blog.
Building and maintaining a search engine with independent infrastructure has a huge challenge and has meant building proprietary IP over many years. Since we refuse to use techniques used in growth hacking such as analytics from you know who, and all tools involving any tracking, marketing is a bigger challenge than it is for companies without strong principles. It has been a mammoth effort, by mostly our founder whose story you can read here [1].
[0] https://www.mojeek.com/about/privacy/ [1] https://blog.mojeek.com/2021/03/to-track-or-not-to-track.htm...
The original comment was in reference to DDG proudly making claims of not getting requests from .gov and marketing themselves as a company who "cannot see what you search for".
If I can't avoid my data being collected, I will still try my best to make it as worthless as possible just out of spite.
Using SearX hosted by someone else is marginally better, but now you have to trust the owner of the server, which is probably not what you want for privacy-centered search engine.
Is the concern that the latter's IP isn't behind a NAT, and therefore is more unique? If so, I think that's the least concerning of the identifying datapoints that a search engine has access to -- my browser metadata is far more identifying. With SearX, that information doesn't get forwarded (IIUC).
You just reveal your search queries to the hosting provider if he maliciously intercepts them.
* Share your instance with some friends (though of course it shifts the trust one level down)
* Route outgoing requests through tor / some VPN. In a hosted environment that you configure once for "everywhere", it's more feasible to do fun things like "Google searches over OperaVPN, bing searches through Mullvad, everything else through Tor". You could even change proxies.
Especially with the latter you can kind of "eat the cake and have it too", with some added latency of course.
As a bonus, searx can also be configured to rewrite links (yt->piped, twitter->nitter, imgur->rimgu etc) and remove tracking query parameters.
Either way, what we're talking about here is more akin to an http proxy than scanning/crawling - since every request is explicitly triggered by manual user action as opposed to automated so it shouldn't be any issue.
If you expose and advertise a public unauthenticated frontend and end up with 10ks of users maaaaybe some providers will start talking to you about it but otherwise I wouldn't have any concerns. And if you get to that scale, you may already want to look at proxying through providers like Luminati anyway.
https://www.vultr.com/docs/install-searx-with-nginx-on-ubunt...
https://www.linode.com/content/maximize-your-privacy-with-se...
https://chrome.google.com/webstore/detail/privacy-redirect/p...
To reply to the person under me, you’re always relying on a trust in something unverified and untrustworthy filters like VPNs anyway, you’re either revealing your IP using a wrapper that reveals it instantly, use a site that isn’t a search engine and might be using your data, using a VPN that is based on reputable and assumptions, usually based in another country you won’t visit or know much about aside from random reviewers, or using Tor, losing latency, reasonable speed image search, and still be possibly compromised.
- They may be hosted in the US
- They may be hosted on AWS
- You have no idea if the maintainer of the instance is tracking you
Points 1 and 3 aren't relevant if they aren't recording the data. Companies in other jurisdictions have no magic invulnerability you can trust to their data getting out (legally or illegally) if they're storing it.
Points 2 and 5 are equally true of any open source project unless you run it yourself from source. There are _plenty_ of examples of users getting phished by maliciously built/hosted open source tools
Point 4 is obviously not malicious tracking and a mistake any project could make
At the end of the day though, unless you're going to run everything yourself (which most people aren't) you have to pick who to trust -- some random person running a server somewhere, or a company with hundreds of employees recruited under the premise of working on a privacy-centric search engine who could all turn whistleblower
I'm fully aware of the massive crawling and storage requirements, but opensource projects that can get search right can later 1) be hosted by the powerhouses of the cloud or non-profit parties, or 2) become a fully distributed hosting and crawling effort as in p2p and blockchain.
> YaCy is free software for your own search engine.
maybe they rebranded and don't aspire to be a complete web search engine?
An "email client" does exactly the same thing, connects to different email servers and we do not call it "metaemail".
edit: just realized that with the current hype around metaverse, 'metasearch' will probably be more appropriate for something searching the metaverse in the future.
Does anyone else have experience or comments on Swisscows' search engine? Seems like an interesting company all round.
I like the idea in theory but in practice I have no idea who I'm dealing with. They could be far more open about their processes. I like the idea of a paid browser though.
On iOS using a new app called Hyperweb you can the new Safari extensions to access and create a longer preferences list. https://hyperweb.app/
We really shouldn't have to choose this or that, but should be able to easily use multiple choices in search. You can do that today as explained here, but you'll need to switch browsers. https://blog.mojeek.com/2021/09/multiple-choice-in-search.ht...
TIL. That's just... terrible UX.
Firefox is my developing browser and I do really like it, but Safari my actual browsing browser because it's by far the best browsing browser on Macs.
On LibreWolf browser,
On mouse hover, the correct result urls show but then I right-click one, it shows something like:
https://duckduckgo.com/l/?uddg=<destination_url_here>¬rut...
But this same behaviour cannot be observed on Google/Searx on the same browser. It also even isn't observable with DDG on a FF nightly build on the same system and a FF stable build on a separate one.
Edit: url formatting
EDIT2: this behaviour seems to be limited to my particular setup and even a clean LibreWolf profile seems not to suffer from this issue. I apologise for the misunderstanding.
IIRC DDG uses Microsoft servers now exclusively. Makes sense given the volume of queries they're handling and all dependent on Bing API.
It gives more context to the topic, as in it's not just a link to the search engine itself.
There's a bunch of language-specific blocklists on GitHub focused on StackOverflow/GitHub mirrors and wonky machine translations, but I don't know of any mature curation effort yet.
[1] https://github.com/iorate/uBlacklist#supported-search-engine...
[2] https://github.com/HoneyLuka/uBlacklist/tree/safari-port/saf...
[3] https://apps.apple.com/us/app/ublacklist-for-safari/id154791...
The CEO sold his previous company's data before founding DDG. His previous company (Names DB) was a surveillance capitalist service designed to coerce naive users to submit sensitive information about their friends.
Is that a fair statement? Can someone provide more context?
Gigablast.com - Has been improved recently. private.sh is supposed to be a private proxy for Gigablast, but it has been broken recently
Exalead.com - run by a French defense contractor for some reason
filepursuit.com - search for files only. Need to play around with it more.
PeteyVid.com - multi-platform video search
Wiby.me - focus on "classic" style web sites
You can do that in this fork: https://github.com/searxng/searxng/blob/e839910f4c4b085463f1...
The only way to avoid third parties to run your own server ... but this "metasearch engine" is basically just an aggregation proxy. So every search can still be tracked back to your proxy server by Google, Bing or whoever is providing the actual results.
Searx is provided as a service on NixOS, which makes this all simple to run.
I do the same for Nitter, a Twitter frontend, which supports RSS and behaves well while logged out.
Edit: I really like DDG bang and vim like nav keys tho
If you search, it goes through Startpage (Google results, more privately). If you search with a bang, it goes through Duckduckgo. It's probably close to what you're looking for.
Complementary twitters list are maintained here: https://twitter.com/SearchEngineMap/lists
Also not sure what the criteria for inclusion is, but search.marginalia.nu and teclis.com both have their own indexes.
So indeed they're doing okay privacy wise, but a lot of users feel cheated when they realize their "independent search engine" (DuckDuckGo) is just a Bing portal hosted on Azure.
Let's say we wanted to recreate the web index made by Google. How much cost and engineering would it take?
Estimating the size of the web from worldwidewebsize.com [0], this is estimated at around 50 billion (5010^9). The average web page size looks to be on the order of 1.5 Mb (1.510^6). The nominal cost of hard disk space is about $0.02 / Gb [2].
So, roughly, that's 75 exabytes of data (~7510^15). At a cost of $0.02 / Gb that gives roughly $1.5M just to buy the hardware to store (a significant fraction of?) the web. The Hutter prize exists [3], so maybe there's some confidence that we only need to actually store 1/10 of that, so around $150k in costs.
For perspective, that's 10 multi millionaire silicon valley types donating about $150k each, 100 "engineer types" at $15k each or 1000 to 10,000 pro-active citizens at $1.5k to $150 each (just* for the hard disk space, discounting energy, bandwidth and other operating costs).
If we try to extrapolate lowering hard disk space costs and take the price halving to be about 2.5 years with a current (pessimistic?) cost of $0.02/Gb, that's about 10-15 years before a petabyte scale hard drive is available to the consumer for $1000.
From my perspective, I would ask "why hasn't a decentralized search index been created and/or is in wide use?". My guess is that figuring out a robust enough system that's cheap enough is still out of reach. $150 might not seem like a lot, but you have to convince 10k people to devote energy just to search.
Put another way, when does the landscape change enough so that decentralized search is a viable option? My guess is that when people can store a significant fraction of the web locally for nominal cost is the determining factor. Maybe some great compression and/or AI sentiment analysis can be done to bootstrap and maybe some type of financial incentives can help solve this issue, but my bet these will only provide a light push in the right direction and the needed technology is the underlying cheap disk space.
As a side note, the worldwidewebsize.com [0] shows the number of indexed pages by Google holding pretty constant over a five year period with a sharp decline somewhere in 2020. I wonder if this is the method of estimation or if Google has changed something significant in their back end to alter their search engine and storage.
[0] https://www.worldwidewebsize.com/
[1] https://www.pingdom.com/blog/webpages-are-getting-larger-eve....
[2] https://www.backblaze.com/blog/hard-drive-cost-per-gigabyte/
Society programs us to think privacy is our top concern. Is it?
This doesn't make it any less important, but just means that if your main selling point is "we're the search engine that cares about privacy", then odds are you're not going to get a lot of users.
Privacy is most effective selling point when working with sensitive information.
But feel free to prove me wrong.
EDIT: It was located in a german street light 20km away from any of the users. Just to get the geolocation question out of the way.
It was more of an experiment than anything else there will be a talk about it and other FreiFunk (Open Mesh Network in Germany) related stuff at the next virtual CCC congress.
Note: since the version 1.0, searx stops sending request for 1 day when a CAPTCHA is detected which might help a little.
(I'm really interested by the results of your experiment)