Edge sends full URLs of pages visited to Microsoft
twitter.com
twitter.com
Most people don't care what's in their food exactly, but that doesn't mean the producer has no responsibility.
"Users don't care" is such an intellectually dishonest argument when users aren't technical enough to understand the subject.
On one hand, a TOS usually describes privacy implications. On the other hand, there's no accountability that a user has understood or even read the TOS. Is it unreasonable to assume people will read the fine print?
Unequivocally, yes. Even if someone hypothetically had enough time to read every TOS they've agreed to I would venture to guess the average user wouldn't be able to comprehend most of the terms[1] anyway. Disclosure documents are written by lawyers (corporate CYA) for lawyers (regulators, plaintiff's counsel, judges).
The law, however, disagrees with me entirely.
[1] Edit to clarify I mean "terms" as in the terms of the agreement
Most users probably don't assume anything, it doesn't even enter the lists of concerns or cares they have or ever would. It's not that they don't understand the subject, they simply do not give a crap one way or the other.
Random family in the US or elsewhere simply does not give a crap about things like this, no matter how much you want them to.
It's not ignorance, and the most intellectually dishonest arguments are really the argument here - "If people knew more about the subject or could know more, they'd care!".
If they knew as much as you did, they still wouldn't care.
Let's say they knew Microsoft can share this information with third parties which might affect their credit worthiness,employability and other very real and immediate aspects of their lives? Let us say the quality and availability of goods and services might change if MS or a third party used this information against them? What if MS can use ML to infer very personal and private aspects of their lives?
By your logic 4th amendment search and seizure restrictions don't matter either because most people assume the government can search them anyways.
There is a reason you have the right to legal representation when accused before the law, you don't understand the law well enough to answer questions and argue on your own. Tech isn't that different, people don't understand it well enough without relying on factual opinions of professionals in the field.
I found myself in the unexpected position of trying to explain why (IMO) it was ok and not a big deal that they did that. I dont know whether I convinced her.
People assume they have privacy when they are alone. Its that simple.
IMO its not ok for any company to change that basic private-when-alone without a very simple, very clear explanation.
"Only 33.3% were concerned about their personal contact information like email address and phone number being shared with third parties, while 10.2% were worried about their friends and contacts being shared."
Can you find me one that says most did?
Because I can't.
Truthfully, if people cared half as much as hacker news thinks they do/should, it would already be a solved issue.
I can easily imagine people not caring about third parties knowing their phone number and email, thats already a lost battle. Everyone is used to giving those out.
Their search history though? their private chats with their friends? the websites they visit when they are alone? Their medical history?
The things they have bought themselves in the last 12 months? their favorite magazines?
Take some time and talk to people, I agree with the studies you found that said people dont care about their contact information, but ask people about other information.....search history is expected to be private, IME, by people who are not tech savvy. The same with their other private actions.
""93% of adults say that being in control of who can get information about them is important; 74% feel this is “very important,” while 19% say it is “somewhat important.” 90% say that controlling what information is collected about them is important—65% think it is “very important” and 25% say it is “somewhat important.”
1) Is it possible to care about something that you don’t understand?
2) Is it possible to care about something whose existence you aren’t aware of?
I personally don’t think anyone can have a meaningful discussion about this issue without answering and agreeing on the answers to at least those two questions I ask above.
Edit: oh, this is in addition to the Mirror stuff. Yea this should never be a thing.
For example they could hash the domain, path and query separately.
The "Google Safe Browsing Update API" (used by Firefox, Chrome, Safari and others) solved this a long time ago. In that protocol, the browser hashes the URL, sends a short prefix of the hash to the server, and receives a list of hashes for the URLs that should be blocked. A huge number of valid URLs all hash to each prefix and the server does not know which one the user has visited. Also, the client caches the list of hash prefixes for which the block list is non-empty, to avoid unnecessary fetches of empty lists, which further improves privacy and reduces response time.
Also, the client doesn't send any kind of user ID token to the server.
From https://privacy.microsoft.com/en-us/privacystatement:
Browser activity, including browsing history and search terms, in Microsoft browsers (Microsoft Edge or Internet Explorer).
I see a lot of this extra functionality being justified in the name of "security", but I don't think the gradual erosion of personal responsibility and agency that results is something which should continue.
"The road to hell is paved with good intentions."
For comparison, it took Red Hat less than two years to open source Ansible Tower.
Epiphany only works on Linux but Falkon is available for Windows as well. Neither works on Macs but Macs have Safari which probably counts as "pure" if you aren't signed into iCloud. iCloud is pretty inoffensive even if you are.
Sure is fast though.
Other than quick site previews, I rarely use it.
This led me to look for a "pure" email client that would _only_ contact my server. Popular apps like Blue Mail and myMail did not work at all behind a firewall that blocked random internet access.
I am happy to report that K-9 mail (open source) was what I was looking for: it never tried to access anything other than my mail server.
I wholeheartedly recommend using NoRoot firewall (or similar) to see just how many random servers your (flashlight) apps talk to (and watch them fail miserably when you block that).
The tl;dr is that
1. Firefox downloads a list of partial hashes every 30 minutes from Google. It's only the first 32 bits of each hash. This can easily be turned off in about:preferences.
2. When you visit a website, Firefox checks the beginning of the URL against the partial hashes stored locally.
3. If the partial hash matches the beginning of the URL, Firefox downloads a list of all of the hashes beginning with the partial hash it matched.
4. Firefox checks the full hash of the URL against the new list of full hashes. If it matches, the site is blocked. If not, you continue on browsing.
When I first found out how it worked a while ago, I was pretty amazed. It goes to great lengths to not know what sites you're visiting.
FWIW, Google Chrome does the same because both Google Chrome and Firefox use the same Safe Browsing protocol v4. The linked post was written a while ago when Firefox still used the protocol v2 but the post is still largely accurate.
Getting you to sign into your computer with a Microsoft account is all about tracking and monitoring everything you do all of the time and selling a record of it to the highest bidder.
Edit: I thought this was a well know fact and if it isn't I might have been to harsh about Google and Chrome.
Edit 2: Thinking about it and searching a bit I conclude that IIRC Google at least used to have access to your browsing history as part of syncing it unencrypted.
"For the transparent proxy to work, it needs .google.com to be added to the URL whitelist to allow all traffic to .google.com. This configuration is not supported because of Chrome security features that are in place, and we recommend that you avoid the use of transparent proxies." https://support.google.com/chrome/a/answer/3504942?hl=en
- Use a web service to help resolve navigation errors
- Use a prediction service to help complete searches and URLs typed in the address bar
and then there's
- Use a web service to help resolve spelling errors
whose description suggests it sends more than URLs.
○ Encrypt synced passwords with your Google username and password
◉ Encrypt synced data with your own sync passphrase
My sync is encrypted so I can't test this for you but I believe if it's not, you can check (and clear) your history here: https://myactivity.google.com/myactivity
Google's privacy policy seems to make it clear that this happens unless encryption is turned on: https://www.google.com/chrome/privacy/index.html#signed-in
Edit: never be afraid to admit your were wrong folks :) So it only gets a “potential list” of sites you visited. Wonder if the operators could aggregate the data enough to deanonymize things
Most browsers/users do, essentially. When a person searches Google for example, results are links to Google that redirect to the target sites. Try it, mouse over a result, and look at the URL. Then a user clicks on one and finds a page that likely has doubleclick.net/adsense/analytics/fonts, which all feed back to Google. Or buttons/pixels/whatever for Facebook. Or both. Or both and 10 more organizations. Then since all mainstream browsers by default send referrer info, and since tracker code is so pervasive, and since browsers are easily fingerprintable, trackers follow you along as you click links going from page to page. Trackers are getting redundant high quality data. They're right there with you as you browse; they see what you see. Although some browsers are easier to configure for privacy, IMO the browser you use is less important than how you use it.
Multiple organizations have the potential to possess a near-complete view of your browsing history.
>curl
curl: try 'curl --help' for more information
>where curl
C:\Windows\System32\curl.exeNot saying MS is not doing anything wrong, but the lesser evil I think.
But what if over the next hour I continue to send hashes with random subject matter sets that include model airplanes each time. Suppose the intersection of the possible subject matters shrinks with each additional hash, quite possibly contains only model airplanes.
I visited 100 URLs in one hour. For each URL there are 100 others with the same hash but independent topics. Lets say there are 10,000 known topics, but every hash I sent has model airplanes associated with it. Now what are the chances I like model planes?
It seems clear to me that this scheme, with logging, reveals a lot in theory. But maybe solving this problem in practice would cost more that the data is worth. For now.
So if you visit 100 URLs in an hour, and using the local hotlist allows your browser to discard all but 1 of those URLs as definitely not bad, that's only one URL prefix checked against Google, not 100 so there's no triangulation.
And the actual number isn't 1-in-100, that's why they picked 4 byte prefixes. I haven't actually checked, but by eyeball from playing with this data for work I would guess 1-in-a-million.
So, you visit 100 URLs about model planes, one of them happens to be a 1-in-a-million match to a possible badware site, your browser sends the 4-byte hash prefix to Google, it gives back the 4-byte prefix of the badware site it was worried about which is different, your browser goes "Phew, good" and nothing happens.
There's just not really an opportunity for tracking here.
1. The truncated hash is checked against local database.
2. If there's a hit, send the truncated hash to Google, and receive a list of full hashes for the truncated hashes.
3. Check if the full hash is in that list. If so, warn.
I'm not certain that the hash is truncated enough, but the general idea is a pretty straightforward trade-off between the amount of information Google needs and the amount of information your computer needs.
[1] https://www.chromium.org/developers/design-documents/safebro...
[2] https://blog.cloudflare.com/validating-leaked-passwords-with...
We've detected that JavaScript is disabled in your browser. Would you like to proceed to legacy Twitter?
And this annoyance persists. Click on a twitter link within the twitter website and you get the message again. And again. It's definitely a dark pattern meant to annoy you into enabling JavaScript.
https://support.mozilla.org/en-US/kb/how-does-phishing-and-m...
Documentation: https://developers.google.com/safe-browsing/v4/urls-hashing
(Disclosure: I work for Google but have nothing to do with the safe browsing API.)
That being said, a four byte / 32 bit hash is enough to almost uniquely identify a website. There number of 32 bit numbers and websites is roughly the same order of magnitude. It's a problem without a good solution because if you create many collisions then you also generate plenty of sites falsely reported as phishing and which admin would want that to happen on their site. If you avoid creation of collisions, you have this identifyability problem.
There is this CRLite proposal [1] using layered bloom filters to stop reliance on web services. Maybe it can be adopted for phishing sites, as well.
Why not send less bits (eg. 24 bit) of the hash?
One possible option is to do what haveibeenpwned does, where you give fewer bits and then locally check. That would be a good improvement to the system's privacy, but you probably want to avoid downloading the hashes of every malicious website that starts with the given 3 bytes (I'd assume the list is quite large) for every page load.
Say I go to https://fakebank.example/security/login and Google has decided all of fakebank.example is a phishing site.
My browser computes [among other things] SHA256('fakebank.example') and then it snips off the first four bytes and compares that to a large dataset it got from Google. It fetches updates to this dataset every few hours. Sure enough the four byte prefix is present in the dataset.
So, we've got an alarm - it calls Google, but it doesn't tell them it's thinking about https://fakebank.example/security/login at all, it just tells them the 4 byte prefix. Google responds with a list of full SHA256 hashes beginning with that prefix that it considers _right now_ to be phishing. The list might be empty (maybe fakebank.example was actually a Greek yoghurt company subject to a PHP 4.x attack, and they upgraded PHP and removed the phishing site so now it's fine) but if it has the entire SHA256 hash we calculated then I get an alert telling me that my browser thinks this is a phishing site and I might want to not visit.
What CRLite does achieves zero for both false positives BUT at quite a price. You need to know absolutely all the things that might ever be in the set before you start.
For CRLite they can almost wave that away by declaring that the set of things that might ever be in the CRL set is the set of logged certificates, so we can get that set from the log servers within 24 hours (the "Maximum Merge Delay" in public certificate transparency logs).
But you can't do that for URLs. The set of possible future phishing URLs has infinite size.
Furthermore the protocol does not declare as URL as blocked just because the 32-bit hash prefix matches. If the prefix matches, the browser downloads the list of full hashes and checks the full hash against that list locally.
It's frustrating to see people jumping to erroneous conclusions about how Safe Browsing works when the spec is publicly available and quite clear. https://developers.google.com/safe-browsing/v4/
Computing 30 hashes of each URL means you are sending a fuzzy hash... and 2^(32*30) is fairly precise...
Since I’m replying to you, I might as well ask, is Google planning on creating a locally cacheable DNS scheme that includes safe browsing information?
Personally identifiable data collection. The goal is obviously data collection.
S-1-5-21-3623811015-3361044348-30300820-1013
The 3 long numbers in the middle are called "Domain or local computer identifier". Further down it says that each of those numbers are encoded as a 32 bit integer. Assuming they're randomly generated, that's 96 bits of entropy, which is more than enough to uniquely identify every computer on the planet[1].[1] There might be a few duplicates due to the birthday paradox.
Same old Microsoft. Second verse same as the first.
While the latter has improved somewhat, only the former has reformed.
Putting their telemetry in your compiled binaries is ... real good? Marketing is the only thing that changed.
Or are you talking about when they enabled generation of ETW events (as in used by YOU in OFFLINE to debug your app.) generation by default?
People like you spreading fud make it difficult to discuss actual cases (like this appears to be)
That's the "new" Microsoft... I'm not sure if there's anything left of the old one anymore.
This behaviour is part of SmartScreen, it's publicly known about, documented, optional. It's explicitly called out somewhere - when you install a new Windows 10 and it takes you through the privacy settings, there's mention of sending browsing data to Microsoft if you agree - https://1re4xlezju-flywheel.netdna-ssl.com/wp-content/upload... - and that's been a warning in the IE first run dialog since before IE 11 - http://www.herongyang.com/Windows/IE_10_SmartScreen_Filter.j...
> When checking web content, data about the content and your device is sent to Microsoft, including the full web address of the content.
https://privacy.microsoft.com/en-us/privacystatement#mainsec...
SmartScreen for Windows Defender will send filenames, hashes and download locations to Microsoft as well.
This is less bad than "like" buttons tracking you around the web in an undocumented and not-optional way with no user benefit.
Is this so different to Google quietly signing you into Chrome when you sign into a Google website, and Chrome sync'ing your browser history to Google by default and sending URLs of pages you visit from the Omnibox?[1] Or "Sends URLs of some pages that you visit to Google, when your security is at risk" from the Chrome options? To Facebook scraping call records and SMS texts from Android devices for years? [2] Ubuntu with opt-out system data gathering in 18.04, and wants to gather data to track "relative popularity of apps"[3], or Steam gathering data about what games you play[4]. Would you believe, Wonderful helpful Google has every email you send or receive through a GMail account or to one, but Evil Microsoft has every email you send through/to an Office 365 account. Wonderful Google and Amazon could access any VM on their clouds, so could evil Microsoft on Azure - except Microsoft built Shielded VMs into Hyper-V to protect the VM from compromised hosts, and to keep secrets in the VM which host administrators cannot access. [5]
Tech companies spy on people for money, isn't as anti Microsoft an idea though, is it?
[1] https://www.google.com/intl/en/chrome/privacy/whitepaper.htm...
[2] https://www.theverge.com/2018/3/25/17160944/facebook-call-hi...
[3] https://www.omgubuntu.co.uk/2018/02/ubuntu-data-collection-o...
[4] https://arstechnica.com/gaming/2018/07/steam-data-leak-revea...
[5] https://docs.microsoft.com/en-us/windows-server/security/gua...
I completely agree with you, this is no worse than many of those things.
"Downvoters should explain" well that "defense" is my explanation. It's not the "same old" behaviour, it's different, what it is the same as .. is other current day, much loved technology companies.
Except for specific ways it's better in Microsoft's favour - this is more explicit and easier to opt-out of than Google Analytics, Doubleclick (owned by Google) or Google+ social media tracking beacons, for one, or "talking to someone who uses gmail" which is super frustrating, for another.