Google maintains the Safe Browsing Lookup API, which has a privacy drawback: "The URLs to be looked up are not hashed so the server knows which URLs the API users have looked up". The Safe Browsing API v2, on the other hand, has the following privacy advantage: "API users exchange data with the server using hashed URLs so the server never knows the actual URLs queried by the clients". The Firefox and Safari browsers use the latter." https://en.wikipedia.org/wiki/Google_Safe_Browsing#Privacy
> Safe Browsing also stores a mandatory preferences cookie on the computer[9] which the US National Security Agency allegedly uses to identify individual computers for purposes of exploitation.[10]
That may or may not be true, but must one be a radical to be concerned?
If you open firefox and browse to a few sites, it will send that cookie. If you then take your computer down to the coffee shop and keep browsing, even if you don't log into anything, it will still send that cookie in the clear.
There are other ways that the NSA can figure out a list of IP addresses you've been using, but this is 1) totally silent, and 2) is common to a lot of systems.
Deleted comment
Also, Firefox hashes the URL and compares the prefix of that hash to a master table downloaded from Google. If the URL matches a prefix in the table, Firefox requests all URLs that begin with that prefix hash. Google is never sent the full hashed URL.
Privacy Theater. No real information is lost since all you need to do is have a database of domains and boom, hash easily reversed.
Not only is that not a valid wikipedia cite, it's not even right. Only hash prefixes are ever sent to Google, and only if it's already tested locally that the hash prefix includes malicious sites.
Humorously this exact same exchange took place in the linked conversation: https://lists.debian.org/debian-devel/2015/07/msg00232.html
edit: wow, it was you. https://en.wikipedia.org/w/index.php?title=Google_Safe_Brows...
Yes, I did that for two reasons:
- I couldn't find a better link for the citation and it seemed like rather important info that should be in the wiki. Maybe someone else would find a better one to replace it with.
- I figured linking to an HN discussion would serve as a great citation, even if I was mistaken about something. Looks like HN didn't disappoint. :)
EDIT: I've updated the wiki text to remove my previous edits and added a mention about the use of hash prefixes.
> Not only is that not a valid wikipedia cite, it's not even right. Only hash prefixes are ever sent to Google, and only if it's already tested locally that the hash prefix includes malicious sites. Humorously this exact same exchange took place in the linked conversation: https://lists.debian.org/debian-devel/2015/07/msg00232.html
Thanks for the link and pointing that out, I stand corrected. I'm curious to know how big the prefix is. Depending on its size this either remains privacy theater or not.
I haven't looked at it in depth for a little while, but according to this it's the first 32 bits of the 256 bit hash:
https://developers.google.com/safe-browsing/developers_guide...
Google Safe Browsing is based on a Bloom filter. The browser downloads the filter in a series of requests when it first starts up, or when the filter is out of date. It also sends a followup request if it finds a hit, but this is rare unless you're actually about to visit a site that's been flagged as unsafe.
And the followup still doesn't send the URL or even a hash of the URL. It sends just a prefix of the hash to download all URL hashes matching that prefix to do the comparison to the actual current URL locally.
I doubt most mainstream users are even aware of safe browsing or how it works. They could very well be ok with it, but that is not obvious to me.
Nor would I say that a desire to disable safe browsing represents a particularly "radical" view.
Quote:
The request body is used to specify what the client has and wants:
* The client optionally specifies the maximum size of the download it wants to retrieve.
* The client specifies which lists it wants to retrieve.
* For each list, the client specifies the chunk numbers it already has.