Mullvad Leta: A search engine used in the Mullvad Browser
mullvad.net
mullvad.net
On a related note, I am also a happy Kagi customer. It's a paid, privacy-focused search engine that gives you a "magic" session link to allow easily searching from multiple devices. Very happy with the search there. Haven't used Google more than a handful of times for several months!
I'm so used to Subscriptions being just a drain. You "buy" the product, and then you pay just to keep using it. Which can feel, emotionally, a bit unappetizing because i'd rather just purchase it fully. The subscription just feels like a money sink with no added value.
Conversely i've not had that opinion with Kagi. Not only am i happy with the product, but the frequent[1] improvements[2] make me feel like i'm buying something newer and better each month.
Developments on FastGPT, increasing what i get for my dollar, integration of more features in general. I frankly assume i just joined at a good time, because this pace can't keep up.. right lol? Regardless my Kagi subscription has felt like i'm getting more value each month. From other companies i'd feel lucky to get these advancements, and if i did i'd expect it to cost me more. "5 new features? Welp, i guess i get to buy a Pro subscription tier to access it" or w/e, is what i'd expect.
Really can't praise Kagi enough.
[1]: within the last couple months, at least, as i'm new to the product and have only been subscribed for 2 months.
[2]: You can see some here: https://blog.kagi.com/blog
Unfortunately they ended that sort of trial usage with the new payment plans. I'm already wary of starting any new saas payments, and one where I need to worry about how many searches I'm doing per month is a non-starter.
"Next, your request will find its way to our servers hosted on Google Cloud platform, where our main application is running [...]"
I stopped there. I'm not going to subscribe. But I appreciate that at least they were honest.
I get the paranoia, but honestly i'm more paranoid of Kagi themselves than i am of processes running on AWS or GCP.
The percentage of the internet you'd be unable to use if you couldn't use any cloud provider for fears of 5 Eyes-like monitoring is kinda intense for me. So i'm not sure where realistic and excessive paranoia meet with respect to AWS or GCP infra.
Thoughts?
I didn't mention but they also use search results from google.
1) A direct contract with Kagi: Pay less for our services in exchange for user's data ("anonymized"). And google has a lot of services that Kagi may benefit greatly: servers, low latency to google search results, "personalized" google search results without ads and the various sheganigans google uses to promote websites, up to the point where all searches actually comes directly from google.
2) No contract, just get my IP and the search string the Kagi server asks of google at the same time. I don't think there are enough simultaneous searches on Kagi to actually prevent IP-to-search-string pairing. Then do the google-ad-to-IP match on the target sites that I visit.
That's fair, but the GCP was mostly what i was relying to.
Fundamentally do you dislike Kagi's offering more than Mulvad (https://news.ycombinator.com/item?id=36402162) on the privacy front? Clearly the two companies have different goals, so i'm not trying to equate them. Just curious.
Also what search engine _do_ you use then? I feel like all of them would be disqualified?
Yes, they are. Until I find a better one, I use DDG.
I might try Mulvad next. I didn't use VPN before, except to access my own LAN.
Also love that they added reverse image search now
Not currently a Mullvad customer ( I was in the past ) but this looks definitely like a good thing!
What sets Kagi apart, and especially, what makes it different?
And yes, I have used Google only a dozen or so times in the last few months since I went all-in on Kagi on all my devices. The search results are very good.
I too love to pay for things, and thus use a product, rather than be the product.
This way I use whatever free search engine for most searches (like when you type "steam" because you don't remember the url), and Kagi when I actually want more than the simplest first result.
However, if I'm paying for Kagi, it would be because I want to cut off using other search engines. Either I trust DDG and it's privacy and don't think theyre selling my data, and I continue using it, or I don't trust them quite enough (not accusing them of anything, I just don't know where they get their money from, and I know its not me paying them, so naturally just a bit worried), and I pay for Kagi instead, then I trust Kagi with all my data because I'm paying them to not share it. And in that scenario I've deemed DDG "not private enough", which would mean I dont want to keep giving them more info about me.
So in the end, I either keep not using it, or I want to make sure I have enough to cover everything. Which is probably just my own issue.
What did you mean with that that is where Kagi is headed anyways? Will you offer the unlimited package for a price which is a bit lighter on the wallet?
Another example “XYZ-brand motorcycle boots after crash”. I want to know how well they survive an actual crash and the brand is popular enough that I bet there are plenty of images out there. Yet all I get is a bunch of promo images of brand new boots.
Give me a search engine that’ll actually return results I want!
I’d suggest to give it a try, it took less than then 50 search a month limit for me to jump onboard
Or maybe it is because I haven't used high-quality search, and I am blinded by it's true greatness.
I use search engines a TON though (especially for work) so 10$ a month is absolutely worth it for me. I am currently trying to convince my boss to buy it for our whole IT team
I want a search engine that's actually good.
(But of course, ideally they would have something in place to prevent such a leak at all, and perhaps they do somehow?)
> 3.4.1 Note Plaintext search queries in cache database
> Assured recommended hashing search terms before insertion / lookup in the cache database. Since search term cache lookups are only performed with exact matching, this should not affect functionality.
> Mullvad: We are now hashing (and salting) the search terms before they are added to Redis
[0] - https://mullvad.net/en/blog/2023/5/16/security-audit-of-our-...
Sure, If I search for "44 little poney street", then the result itself is cached at Mullvad, and someone needs to search himself/herself for "44 little poney street" by entering precisely this search string to access the cached page.
So I don't see a leak with caching... There are leaks anyway: the search term sent to Google, if someone compromises Mullvad, etc... But not one specific with caching and related to other users.
No "might" about it, that's one of the most important traits of this type of service.
Seems like a security risk.
If so it's like 16 digits. Isn't that 10^16 values? If they had 1 million users, that's still a lot of numbers to test before you find 1 valid one :)
I suck at math, but that's like 999999999 non-existing accounts per valid account? (10^16 - 10^6 - 1)
> Regrettably individuals have frequently used this feature to host undesirable content and malicious services from ports that are forwarded from our VPN servers. This has led to law enforcement contacting us, our IPs getting blacklisted, and hosting providers cancelling us.
https://mullvad.net/en/blog/2023/5/29/removing-the-support-f...
The specific content doesn't really matter, tripping the sensors for enough sites could potentially get entire IP blocks flagged.
They may not also be able to reveal specifics if it is an ongoing investigation.
Prompt: Give me a single sentence technical reasoning a VPN company could use to discontinue port forwarding feature.
GPT4: "Due to the increased security risks and potential for exploitation associated with port forwarding, we have decided to discontinue this feature to enhance the privacy and security of our VPN services."
It's not a secret that a no-log policy also attracts abuse.
https://mullvad.net/de/blog/2023/5/29/removing-the-support-f...
What I get for trusting mullvad I guess.
- Leta is much faster than Startpagw - Startpage offers a lot more of Google's features, eg date range filter, image search, and so on
I would guess that both differences are due to Startpage not doing any caching.
Startpage also has a neat "Anonymous View" feature where they proxy the request for you, acting as your HTTP client. If you trust Startpage, it's probably a pretty good ad-hoc anonymity tool.
It's a caching proxy for Google Search and could well just be Squid.
I assume it also doesn't interact well with Google's location services.
It's a bit like asking if you can install Cortana on Trisquel GNU/Linux.
If you remove that, it's as bad (or worse) than most of its competition.
I remember a time when my wife was trying to look up a fix for Mass Effect on a 21:9 monitor. Terms like "ultrawide mass effect" and such. Google would not stop returning Blizzard help pages on how to configure the resolution for Heroes of the Storm, another game that she played. Not a single page related to the actual search terms. The more we poked at it the more I couldn't believe it. Bing, of course, just did the dumb, obvious, correct thing and returned a bunch of web pages containing the search terms, which were helpful.
Google seems to do this infuriating thing where it reduces search terms to basic "synonyms" (which are often more general than the original word, e.g. "Mass Effect" becomes "Video Game") and then injects personal search history related to the synonym (which is how Heroes of the Storm ends up as part of the query). Most of the time it's just subtly enraging; you know the page you're looking for exists, and you know your search is extremely precise, but Google keeps giving you overly-generalized results with a skew towards your "profile".
Anyway, all that is to say that I feel exactly the opposite of what you feel about the relationship between this Google "feature" and the quality of its results.
I am not expecting it to know my exact GPS location but would be nice if it could at least bring state or country level tailored search results.
But for local results, I'd just prepend $cityname to your search query. Faster, unless you live in Llanduwhatsthattowninwales.
https://en.wikipedia.org/wiki/Taumatawhakatangi%C2%ADhangako...
The difference is that there is less noise from other users as it is limited to Mullvad subscribers, and there is presumably a smaller user base.
Otherwise, there is probably little to no difference, considering that Searxes are not used by many in the same vein.
However, self-hosting is the equivalent of directly using the search engine under your own IP, just without javascript. There is no noise from other users looking up unrelated queries.
If it takes a correlation of the 3 datasets to identify you, then it is better to use 3 different providers.
However, if any one of those datasets is sufficient to ID you, then it is better to choose a single provider.
The Mullvad website and the https://mullvad.net/en/check page show that Mullvad already has tools to detect users of its VPN.
It seems the people this most affected were the ones using VPNs primarily for torrenting, which I've always just used a VPS or dedicated server for. Though, even in that case, it's not like it's impossible to torrent without port forwarding, millions of people do it every day behind their NAT.
It is unfortunate they had to remove the feature, but I have to assume the abuse of the feature was at the level where it was threatening the service as a whole, if I had to choose between Mullvad without port forwarding or no Mullvad at all, I'd obviously choose the former. They also do seem to be refunding people who request it, so it doesn't really seem like any kind of "rug pull" or anything.
It did change though, I’ve been using them since they started but in the past 2ish years their network is very bad, slow, continuous interruptions and disconnects (can’t say it correlates but noticed happened around the time Mozilla VPN started as they use the mullvad backbone), blocked in a lot of regions even in some government websites, anong other issues, the straw was when they stopped port forwarding.
As for IP blocking, I've also rarely encountered that, when I do it's mostly on e-commerce sites, and in those cases I typically find it's just a single exit IP that's blocked and setting up a rule for that domain to tunnel the traffic to a different server (via their SOCKS5 endpoints) fixes it. I can understand how having to do that might be an annoyance to some people, but again for me it's not really a big deal, just a few occasional minor inconveniences in an otherwise good product.
Edit: I should also say I don't really use any services like Netflix or things like that, it's my understanding that streaming sites like that almost universally block Mullvad since they make no effort to mask that their IPs are from datacenters. Again, not an issue for me, but I definitely could understand if that was a deal breaker for some.
For two peers to connect, at least one needs to be reachable by the other. Behind a VPN that requires port forwarding, so if you don't have it you rely entirely on peers that are reachable.
Anyway, I've been a customer for a long time, and will continue to be.
Very much still are.
Some key points:
- Acts as a Google proxy, removes tracking links and caches results
- Only available for Mullvad paid users
- 100 free direct searches a day, unlimited cached searches (further search result pages count towards limit)
- Results cached over all users for 30 days
Also why would I trust them over Google?
- They have strong commitment to open source and have put their finances into that in addition to releasing code.
- They are doing a lot in terms of transparent infrastructure: https://mullvad.net/en/blog/2022/1/12/diskless-infrastructur...
> Also why would I trust them over Google?
For Google your data is the product, for Mullvad you pay for a service.
Mailing cash in an anonymous envelope has a certain charm, but OTOH I have consistently had terrible experiences with the Swedish postal service and that seems to be a widespread opinion.
https://mullvad.net/en/blog/2022/6/20/were-removing-the-opti...
You can see an example of their lack of data retention from a post about when they were raided - there was nothing to find.
https://mullvad.net/en/blog/2023/4/20/mullvad-vpn-was-subjec...
Their blog is a good place if you want to get a sense of what they're like as a company.
This is a cool addition for sure though
I wonder how well the caching actually works for a user base of the size that Mullvad has.
This could be tackled with a different UX, perhaps rather than showing predictive search, instead showing similar queries that are in the cache? I'm not a customer so can't see the product, can any customers give any input as to what the UX is and whether it might be improving their cache hit rate?
The interests being more heterogeneous results in more similar queries, which would increase the proportion of cache hits. Whether this is enough to help make the strategy viable is another matter, but I do think it's worth noting.
I also wonder about the complexity of the queries themselves. The more technical users would probably use more complex combinations of operators, but they're also more likely to search by keyword rather than natural language.
The country dropdown is interesting as far as the cache goes - not selecting a country is meaningful as far as the cache is concerned. My prior "dog" query in the US does not return hits if I don't select a country. Not selecting a country and searching the cache appears to return english results (with a few sample searches).
It's interesting that you can explore the cache with this checkbox. Not sure if there are any privacy concerns with this feature - considering cache searches are "free" you can kind of scrape what other users are searching for, maybe with enough users it doesn't really matter. I suppose there could be rate limiting and such to prevent this kind of attack, but that's just a guess.
It may be useful to have an option to opt-out your search from cache.
I think the profile of their users is less diversified: mostly tech savvy people. "Normies" are using those vpns advertised in YouTube, or not using any at all. This may result in similar interests and lower the number of unique queries.
On the other hand, we may produce more unique queries than other people: who will receive-use the cached "how to fix ValueError on main.py:67"?
A study[1] by wikipedia done with DDG notes it showed up in the top5 results and information module for ~13% of searches with a click-through rate for each at ~8% - so a total of ~16% click-through rate. Granted, that is not a number gained from title searches but the whole articles.
[1]: https://diff.wikimedia.org/2021/09/23/searching-for-wikipedi...
They state that in 2022 they stopped accepting subscription payments because it forces them to store data about their users for long periods of time. Now they only accept one-time payments for monthly memberships.
They really are committed to privacy.
Given, this is assuming everything they say is legit. It's kind of hard to not be jaded these days.