Qwant: The Search Engine That Respects Your Privacy
qwant.com
qwant.com
Not as sexy as the high-minded front organizations that reskin it, but it'd be nice if we could be honest with ourselves about the beasts we were feeding.
From ehat I've seen, search engines like Ecosia and Qwant aren't very clear about what they bring to the table, beyond serving as a proxy for Bing that doesn't store as much data on you.
I don't think there's anything explicitly wrong with that. But I don't think they really merit anybody's attention either. Especially since they don't explicitly advertise what they are.
I've been reseasrching some ideas.
Also, for the older crowd who came from slashdot (and still says 'Micro$oft') it'd be morally unacceptable to use Bing directly. They prefer either a rebrand, or to directly give all their data to Google (maybe because they still see it as fresh and new?)
So close, but no cigar.
You think you are doing something, but I have strong reservations it achieves anything except making you feel better and more in control.
Yes, if you block those. You're also not in a search bubble.
I feel DDG can be compared to Apple in many ways, but here’s it’s important to look at what Apple is doing to its Chinese users by letting the CCP spy on all their iCloud data (sacrifice all the privacy values to obtain access to China’s market). Perhaps DuckDuckGo have also been forced to sacrifice their users’ privacy to join Microsoft Advertising as well? It’s difficult to tell since they’re closed source, and Microsoft Advertising is a private ad network (anyone besides DDG/Ecosia that’s been invited?) that doesn’t make any API documentation publicly available. At least with Bing’s API we can tell that Microsoft just need the search query for it to work, but how about their Microsoft’s advertising API?
But since that country is part of 9-eyes IIRC, I wouldn't place a lot of faith on that
Google > Startpage > Qwant > Bing > DDG
But even if they were 100% identical, it has some privacy unlike 3 out of the other 4.
They give you cloud sync functionality in return.
That’s assuming that the users’ goal is to avoid be tracked rather than just replacing Google with Microsoft.
If Bing themselves wanted to get on the HN first page all they have to do is go all in on privacy.
But reading the French version of Wikipedia (I happen to be French) and the conclusion is much more tenuous. So I guess I'm wrong!
Bing sets cookies. Qwant does not set cookies.
Qwant, i.e., lite.qwant.com, prefixes search result URLs to point to Qwant servers. Qwant redirects www.qwant.com to lite.qwant.com when Javascript is disabled. Bing does not prefix search results.
Bing requires sign-up in order to use their API. Qwant's undocumented API is freely accessible, no sign-up.
Qwant, i.e., lite.qwant.com, requires a User-Agent header. Bing does not require a UA header.
Example of Qwant API
curl "https://api.qwant.com/api/search/videos?q=example&count=150&offset=0&f=xyz&t=xyz&l=en_gb&uiv=xyz"Crooks use search engines to find pages that have exploitable js code or sql injection vulnerabilities. They use them to find unprotected comment sections so they can inject spam into them. They use them to build dossiers on people by scraping public information sites. And all that API use, they never pay a dime for access. Nor will they reveal who they are in order to get access. It just isn't how they operate.
At its peak I had over 2.6 MILLION internet hosts black listed from using the Blekko API. Exactly zero reached out to any of the easily found contact addresses and said "Hey your API seems to be unresponsive" :-)
Lots of botnets were highlighted this way, when one address searches for 'joomla v2.3', and the next IP searches for 'joomla v2.3', page=2, and the next IP searches for 'joomla v2.3', page=3, etc.
It was annoying but an interesting problem. We could implement any arbitrary policy and then watch as the bots adjusted to come in just at that policy limit. We banned entire Ukrainian ISPs (they were a big source at the time) and have VPN providers become the big users. We put in limits per day, I tried a "thermal" system where IPs gained "heat" by queries and "cooled" by idle time. We built a server with a "broken" IP stack that we could send the initial TCP connect to, the server would accept the connection and then never respond. A "black hole" if you will. The trick was we didn't actually keep[ sockets open we just pretended like we it was the other end of a TCP connection. It did everything correctly except complete the connection. That would cause any client using off the shelf IP stacks to hang indefinitely.
It is a game with no ending as one might say.
curl 'https://api.qwant.com/api/search/videos?q=example&count=150&...' -H 'User-Agent: Mozilla/5.0 (X11; Linux x86_64; rv:84.0) Gecko/20100101 Firefox/84.0'
Results seem just as good as Google, PutLocker searchers work, and it's as privacy-respecting as DDG. I've been using it all day and I'm happy with the results, so I thought I'd share.
The company is pure hype, no results.
There are actual results - as a public subsidies scam!
This seems like a weird accusation given that government VC arms like In-Q-Tel or institutions like DARPA invest heavily in Silicon Valley tech firms, not to mention that the national science foundation supported the development of pagerank itself[1]. One of Apple's first investors was the Small Business Investment Company, an investment arm of the federal government. The EU would be stupid to not support the development of an independent technology sector.
[1]https://www.nsf.gov/discoveries/disc_summ.jsp?cntn_id=100660
It's Open-source, it's a search engine (they have their own crawler and their own index) and it's private!!!
I mean: I've developed websites... What the was in my head?!
LOL
THANKS for not letting me look like a dumb for too much long
I’d be curious to hear if anyone self hosted this and searx and how they compare.
So I end up using both public and self-hosted SearX, DDG, and Google as last resort.
Still some privacy benefits in not having physical location tied to searches, and an instance can be shared with friends and family.
It's a hard technical problem, but not one with a concrete barrier to entry. And there are lots of really smart people solving really hard technical problems out in the open. On top of that, much of the recent progress in AI is open-source. Surely that could help?
I guess indexing at that scale takes a lot of hardware. This is the most plausible barrier that I can see, but it still doesn't seem insurmountable.
What if you rethought what a search engine is? Does it really have to cover all of the text on every page on the internet? Could it be more focused?
This just seems like a solvable problem once there are enough smart, motivated people. Which there appear to be.
And of course there are lots of other issues. For example the Web is very hostile when you are not Google Bot (i.e. there are a lot of big sites that will forbid you from crawling their content, unless you are Google, or Bing).
> there are a lot of big sites that will forbid you from crawling their content, unless you are Google, or Bing
If you're talking about robots.txt, that's just a suggestion. It doesn't hold any actual preventative power.
Even if all websites were acting in good faith, this would be a really hard problem. Now add in the fact that there's an $80 billion industry devoted to gaming search results, and it suddenly becomes much harder.
Would it be impossible to build a search engine as good as Google? No. But you could very well spend billions of dollars to match or even slightly surpass Google's performance and still lose. Why? Because unless it's clear and obvious that you're superior to Google, people are still going to use Google because that's what they're used to.
That's why the smart players in the space like DDG aren't competing head-to-head in search, and are instead focusing on areas, like privacy, where Google can't compete. As for the others, I suggest you try Googling "cuil" some time.
(Maybe that’s what the average users are looking for, though.)
My recollection is they started by doing their own crawling.
(I want to say the username was Weinberg?? But honestly can't recall if that's right.)
TBH it's probably also the thought that Google started on, as principally we had manual indexes back then.
Again IIRC, DDG got along way with private indexing and then realised it didn't really gain them much and they instead wanted to focus on being a Google competitor. They did that by focusing on instant answers as a distributing feature and by buying results to get closer to the coverage they needed.
Lucky for them Google made their search a lot worse so DDG's quality caught up.
None I've encountered so far are actually good enough to compete. This one doesn't seem any different.
This one's hosted in France. France is part of Nine Eyes, among other things.
Google long ago stopped returning useful results for my academic research. For example, it never returns Wikipedia articles unless I explicitly use Wikipedia as a keyword, which I now have to do for most of my searchers.
So here we are:
1. They are not Open-Source (as opposite to Searx).
Only the plugins are. Can we trust them when they say to respect our privacy ?
2. They are not TOR friendly.
Full of captchaS !!!
3. They have no onion site.
DDG have one.
4. They run some kind of weird analytics.
Each time we click on a search, the JS code trigger a fetch to `https://api.qwant.com/api/action/url` and include: - our current language - the query we searched - the link we clicked - etc... They were backed by Bing before, is it still the case or are they running their own stats engine? I do not know. If we trigger ourself some false fetches, can we show "twitter" as first result when someone searches for "facebook"?
5. Their lite version is not lite (=/= duckduckgo.com/html/).
They redirect all our clicks to be able to run their analytics without the JS's fetch API (proof: `https://lite.qwant.com/?l=fr&q=hacker+news&t=web` the first link does not offer `https://news.ycombinator.com` but `https://lite.qwant.com/redirect/yFOdE8r1P1LTSLsA9IIBNKaZmDF1...`). Possible attack with `https://lite.qwant.com/redirect/yFOdE8r1P1LTSLsA9IIBNKaZmDF1...` (see the fake "&query") ? Also, They cannot store the config: go to settings, do whatever you want, do your search, switch tab (go to "news" for example) => your settings are reset to default. You may fix it by adding your settings to the link "&l=fr...".
6. Their front-side is broken.
If we visit their site without user-agent, we have an exception in their JS which crash the page (blank page). And their is more.
7. They do not care.
I emailed them maybe 5-6 bugs, they never replied nether fixed them.
8. Their API is sometime weird (just because not documented ?).
I ran a custom front-end without Qwant's analytics and the minimum working request is: `https://api.qwant.com/api/search/web?q=hackernews&count=10&o...`. What is "&uiv=4"? Why can't it be null or 0? What is "&t="? Is it really needed? Why is it not needed everywhere? => `https://api.qwant.com/api/suggest?q=duckduck`.
In sum up, in a customer point of view, Qwant is just a frenchy Google hosted mainly in Europa and allowed in China. Nothing new here, I recommend to stick to DDG for the moment.
Edit: fix minor typo + add proof for QwantLite analytics
* They promised (to the public and their investors, which includes the state funds) that if they used Bing, they were improving they own engine which was handling more and more queries. Studied showed it was not the case, that many searches showed outdated results and repeating entries (to make think they have many search results).
* Their previous boss was a tyrannic Jobs wannabe. He wanted Qwant to become as trendy as Google so he kept launching half-baked products (Qwant maps! Qwant mail! etc). Constant chance of priorities and new projects were devastating for developers (especially when the search engine was still very weak). He was finally pushed out of the company.
I've even tried with different IPs / VPN connections and it still refuses to show results.
Surely I'm not the only user who experiences this?
However, it works if you use a non-English language:
https://lite.qwant.com/?l=de&q=foobar
https://lite.qwant.com/?l=en&q=foobar <-- this won't work (or omitting l=en)
https://www.qwant.com/?q=what%20is%20storage%20tiering&t=web
top 5 are ads, nothing relevant :(
Edit: saw your other message, doesn't work in English because of a bug