Google Will Soon Let Users Automatically Scrub Location and Web History
buzzfeednews.com
buzzfeednews.com
E.g. I cant store my home location in Google Maps without turning on Web & App activity history. Why? I cant share my current location (not historic, just my latest location - they only need to store one) with my partner unless I turn on Web and Location history. Why?
I am sure a lot of people will now pipe-up and say "Ah well they want your data! That is why!". Kthx bai for that - I get it.
It strikes me as this being a "nudge" more than anything, or perhaps just a cut-corner to get a feature enabled rather than support something custom (e.g. in the case of my current location vs location history, it was probably easier for them to build something that just picks the most-recent record off of the top of your location history rather than create something new for only storing the most recent ping independently from location history along with all the privacy controls, legal reviews, documentation, and UI fiddly bits that go along with supporting something new)
I turned off youtube watch history.
I can no longer see my google play music played counts, because somehow the same permission covers both, even though I would like music history but not youtube history.
The YouTube music experience is kind of odd. It seems that you can only listen to full albums if someone has uploaded them to YouTube as separate songs. Which makes sense, but it means that currently a lot of the music I listen to is nowhere to be found on YouTube music.
And it's absolutely not Android's fault as much it is app developers' faults. Want to use an app that tracks your dreams? Most devs will have you sign up for their terrible service so they can store the information in plaintext on a server.
At least with iOS most devs seem to opt for CloudKit with E2E encryption. Privacy seems more wired into the iOS ecosystem than Android.
The issue is that the developers themselves don't really seem to care as much about privacy etc as iOS developers. This effectively renders the privacy measures mostly meaningless. If I'm writing my dreams into your app and you can read them on your own server, I've already lost all semblance of privacy.
I do see it as a shortcoming of android that i cannot redirect unwanted capabilities into a harmless sandbox version of those.
Not sure if you're trying to provide defensive cover to Google with your comment or not, but it clearly goes to show they are not above these kind of dark UI patterns.
you know why. It's profitable. This information makes them money.
I can understand needing to remember the last query, in order to provide context for the next one. Automatic scrubbing would be helpful...but it's missing an "every 10 minutes" option.
(Today, my solution is a completely separate Google account that only gets used for Google Home/Assistant.)
A Google employee once called me on my personal cell phone when I was a few minutes late to a business meeting. I still don't see how that could've happened unless they'd linked my work (g-suite) account with my personal (gmail) account, and then allowed their employees to access personal data in a totally inappropriate fashion.
Maybe there's some other explanation, but I can't fathom what it would have been.
I in fact do have evidence to believe it's more correct than yours but I have no inclination to discuss it further.
Evidence. Not proof. And it's not secret. Just not something I'm inclined to share here.
> which you refuse to share, but that sounds more like a troll or a bullshitter than somebody with actual evidence. I wonder why.
Because of (1) my privacy and (2) the fact that you're attacking me like this. You're welcome to see me as a troll if you'd prefer that.
I have phone calls and e-mails to a personal g-mail account that make literally no sense to receive _unless_ Google had linked my personal and business accounts in a CRM someplace; an d then given access to personal contact info to some of their staff.
> the fact that you're attacking me like this
How is this an attack? I've done nothing to you that you didn't do to me. Except in my case I revealed the nature of my evidence whereas you revealed literally nothing. You just came up with a dumb and obviously false hypothesis and declared it true.
You're saying I called you a troll?
(This discussion is over from my end.)
There are ways for technical users, but it nudges users to leave all the juicy data to the company, which can then use it in any way it feels (including monetization)
In theory the links are only valid for a few hours. In practice I don't trust that.
Even then, it's not about clamoring, but just having a way to create rules. You just can't do that easily in gmail.
The actual email to non-gmail addresses is just a link to fetch the mail body, which can be validated with SMS.
I would very much like to store my home and work addresses in Google Maps on my phone, but I can't do that without them also storing everywhere I go.
There is no good reason for these two things being tied together other than to make we share my location data. We are talking about storing a few hundred bytes of relatively static data and letting me select them as a shortcut, this is not a complex feature, it's a simple shortcut.
whoops
There's no such thing, at least at scale, as sharing your information without also storing it. Not if you want the system to be reliable. So in order to make Google share your location with others, you have to give them permission to store your location. This should be self-evident enough if you take a moment to think about it.
As for lumping web history together with location history, it turns out you can derive the latter if you know the former. There's been a lot of fuss and breathless reporting about that fact. So authorizing them to know the one without the other is a bit of a fiction, especially in those cases that get written up by tech journalists. Best not to pretend that it's possible to keep them separate.
This new thing is an acknowledgement of the fact that the whole point of using Google services is sharing information with them and asking them to hold on to it, at least for a bit, at least until you've gotten the desired use out of the service... but also of the fact that a certain kind of person, the kind you often find on HN, would sleep better if Google deleted that data after it had served it's immediate purpose.
Sure. I have no problem with that. What I have a problem with is, why is google insisting on storing all the locations that I am not sharing?
I only want to share this current location. So yes, that will have to be stored somewhere. But for the love of everything I don't see why sharing my current location requires saving all my future locations.
You can also imagine the other sub-rosa reasons for this.
So, the question remains, why do they need to record your web history if you want to save your location? You can't derive it that way round.
I have all this stuff turned off, there is no good reason for google to store it, it's a privacy nightmare, esp given automated sharing with the state.
Edit: To clearly demonstrate what I'm talking about: https://imgur.com/a/QW9OxAS
They pin your saved locations to the top of the UI. If I click "Home" it will take me to my exact address. However, if I attempt to search "Home" in the search bar I get
"Turn on your Web & App Activity setting to search for 'home' and other personal places"
No thanks, Google. You don't need to know my entire location history 24/7 to take me to a static address you already have saved.
You can't - or, rather, I can't. It may be a problem that only affects me, I haven't talked to other people about it, but that was exactly my thinking a few months ago. So I activated location services, changed my home address to my new address, checked that it was set right, and disabled it again. And whoops, it's my old address again.
They might've fixed that in the mean time, I haven't checked it recently, I've just given up on using maps.
Why do I need to turn on "Web & App Activity" to store my home or work address?
Because Google wants to inconvenience us into turning it back on, plain and simple.
Which is perhaps generally applicable to "conspiracies" in fact. There are some instances of powerful people making plans; but there are even more instances of results from systemic rewards and punishments, people acting independently with certain interests, from just how the system works. It doesn't make em always great. Systems can be changed.
The only way out for Google is to actually start paying attention to that issue and make a conscious effort. And it'll take time before they'll regain the trust.
In the case mentioned, searching "your locations" appears to be just all on or all off (I have no insider knowledge of Maps). That greatly reduces the surface area for heisenbugs in a high QPS system.
I don't have the exact screenshot, but the popup looks approximately like this [2].
They intentionally wrote the code to make sure that the user doesn't make the wrong choice. And this popup is not really necessary for the user, mostly for Google.
Also, this reminds me of Google's "Amateur Hour" story with Firefox. Every time they make a mistake, it's in the favour of Google, what a coincidence.
[1] https://android.stackexchange.com/questions/115944/how-to-pr...
Deleting it from say my browser, maps, or other history is one thing. Does Google no longer have any of it?
Also the top and bottom bars that leave me with like 1/3 usable space on google's blog is really frustrating to read on: https://blog.google/technology/safety-security/automatically...
They could delete it... and then repopulate that data the next minute pretty easily.
Maybe it is just just the writer's POV, but I do worry they didn't explicitly say it is deleted as far as Google's storing of such data...
Unless they're audited by a source that can be trusted and have the findings made public, I will not believe it either.
Developers could have:
sql = "Select * from Foo Where FirstName = '" + firstname + "'";
All over their code and no one in "compliance" would be any the wiser.Does Google have multiple "Deletion policies" such that deleting data from i.e. your GCP bucket follows one policy, and the "scrubbing" described in this article follows an entirely different policy? If so, do different deletion policies have different processes and different audit trails such that the end "deleted" state is subjective and controlled by the engineering and managerial oversight of the engineering/leadership team of that given product(s)?
From my (naive) opinion, it must be really, really, hard to for example, retrain every ML model that a now deleted datapoint ever touched. Its hard too to believe that, at some high level in Alphabet's org, there is no motivation to have the positive PR of feature(s) like this, but still at essence not delete the parts of the data trail that significantly drive Google's revenue. Do these datapoints significantly impact Google's revenue?
So with that caveat in mind, let me see what I can help answer.
I'm not entirely sure what nuance you're implying when you say "different deletion policies" - while for instance Cloud might have a different timeline or set of triggers for when and what data is deleted, when it happens "deleted" still generally means "deleted". Some products like GSuite have the ability for administrators to say, disable accounts, which removes them from use but doesn't delete the account, but that's transparent to the domain administrator.
It's definitely nontrivial to track data propagation within large systems, but standardizing infrastructure, having central documentation of data handling plans, and having comprehensive privacy reviews for any new functionality that launches helps keep people on the same page.
Edit: oh, and regarding "retraining every model a data point touched" - the easiest way to do this is to just always be regenerating your models on a frequent basis. If you retrain your models once a day or once a week on a fresh snapshot of your data, they'll only ever be that stale.
No offense to you, but I remember when Amazon was releasing their home devices, and many people rang alarm bells in these forums only to be answered by supposed Amazon employees or friends thereof explaining why these devices couldn't possibly been sending data. Well low and behold, they are sending all sorts of data to Amazon. Were those commenters just lying? Were they misinformed? Were they trying to spread disinformation for whatever reason? Perhaps all three..
Here in America, the gig is up. Everyone, even our grandma's, understands that security and privacy is always going to take a back seat to profit. Always.
So again, no offense to you personally, but everything you are saying must be taken with a ginormous grain of salt.
The only answers are either open source or objective third party auditing, or preferably some combination of both. Words from google employees mean nothing.
https://www.usenix.org/conference/srecon18asia/presentation/...
You can see there's a section on privacy and deleted data as well.
Each team has its own policies, because each product is different: at a bare minimum they might be using different storage systems, but it's very likely that their data pipelines are quite different, too. In any case, each team's targets are at least as strict as any published ones, of course.
That's different from your second question about model training/retraining -- there is an answer to that questions elsewhere (but it is taken seriously and training data is also deleted upon deletion of the source data where it contains PII). I don't know this next part, but I suspect models use such vast amounts of data that any arbitrarily small deletion wouldn't have much impact.
I had a similar concern where I was unsure if Google deleted my data and then anonymized it for helping their models anyways. In my opinion, that would be materially different than deleting all of it.
Oh we're all good then!
The entire burden of proof would be on you, vs. a behemoth corp whose bottom line depends on maximizing data collection and retention.
Go to court on your own dime, prove it, and yes, the penalty might have some teeth. Might.
https://www.law.com/therecorder/2018/12/19/facebook-is-being...
this is just naive
https://www.washingtonpost.com/business/technology/google-en...
But I mean if your position is that the NSA has cracked modern encryption technologies, then I guess you better get off the internet. Whether you use Google or not, you're screwed.
Edit: Stray somewhat-related question just in case you'd know, why is it that if I open YouTube in a private browsing window on a computer with freshly cleared cache, it asks me which of my two Gmail accounts I want to log in with? Is it just IP address based plus some browser fingerprinting?
Besides, for that document to work, the authors have to be trust worthy.. in this case the authors are not even close to trustworthy it might as well have been written by the hamburgler.
I bet it's not free to run, but it's cheaper and easier than elsewhere, because Google's infrastructure is built in-house and mostly integrated. I don't envy other companies that want to do the same.
If you're not willing to put even that must trust in your various cloud providers, you aren't the target market. Surely you aren't sharing this information with Google already, right? So the feature isn't for you.
What is worse:
- losing some data of the 0.1% of users who actually care, or
- the regulatory and PR nightmare that would happen if they don't and it is brought to light?
(Disclosure: I work for Google, though not on anything like this. Speaking only for myself, not the company.)
I understand your skepticism but at least in this case, I think your incentives are aligned with Google's. If we don't delete it (within ~30 days I think?) then I'm pretty sure we'd be in violation of GDPR.
Is that to say it's definitely deleted? I can't say that for sure since there bugs are always possible but at least I'm pretty confident the intention is to delete it.
Edit: Nm, I see you're a Google engineer in this domain.
That said, there's also data without PII (anonymized) that is not attached to any specific user, that of course is not subject to this data wipeout as Google wouldn't know what to delete. Such data can be used for general ml models.
Not sure how it's done at Google but anonymization is tricky and it's easier to gather two forms of information from the start, one with PII and one without PII. I think most large companies are doing it like that.
EDIT: Just remembered one thing. At Google, even when building general models data that is rare cannot be used because it could be used to identify specific people. For example: If search query is used by less than X people on a given day then these queries cannot be used for model building.
You can also, of course, read in the data, rewrite it with the deleted portion removed, and write it out again.
(Disclosure: I work at Google, though I don't know anything about how Google handles this)
Google may disassociate your data from your account, but your data lives on.
Maybe it's normal for a startup, but for a billion dollar company, they're not going to roll the dice on 4% revenue fines and IP address geolocation.
(Disclosure: I work at Google, I'm not a lawyer, and this is not legal advice)
[0]https://www.eff.org/es/deeplinks/2019/04/googles-sensorvault...
A lot of people are complaining that if they disable data collection, try their hardest to not contribute to these data models and block the ads with ad blocker then some feature doesn't work for them. So you just want to use Google services completely for free without contributing anything.
Other people complain about permissions that Assistant or other products require. There's always a very good reason why these permissions are required that is not monetary. Each permission request goes through a lot of scrutiny, lawyers reviews, product reviews, approvals, etc. You can ask why disabling some permissions affects features that shouldn't need it. Sometimes it is lawyers fault, sometimes it's just engineers not supporting properly each of 2^{number of permissions} permission combinations.
Not necessarily. I believe that google's services are good enough to use them, and I'd certainly pay a few bucks a month to do so and be their customer - and not be the product that they sell to their customers. I'm also pretty sure that they'd make more money off of me paying them for providing a service to me than they make by selling my data to advertisers (me blocking ads and all that). If the only way of "contributing" however is to give up my privacy, then yes, I don't want to contribute.
A lot of online publications frustrate me with this as well. They're only supported through advertisements, and request that you disable your ad blocker.
LET ME GIVE YOU MONEY! I would pay for the ability to read your site. Right now I have the choice of ad blocking and preventing them from earning much of anything or giving up my privacy to support them.
It's especially ironic given that many of these sites are increasingly writing about the dangers of "surveillance capitalism" and big tech.
There is no 100% security.
Some companies do a lot better job than others at protecting their systems, but it's only a matter of time before they are breached or information is leaked, especially a high value target like Google.
If a powerful nation state wants to breach a corporation, they will.
So, the real question is this: is an affinity for targeted advertising really worth creating the most detailed psychological profiles in history on billions of people if it is inevitable that this information will be compromised and used for more nefarious purposes?
If FB let me clean out all my posts/comments/activity from everything older than a year I might actually start using it again.
- Check the latest user status after 3 months before deleting user history.
- Compare it with the previous 3 months and update the difference ona separate table that are not facing the end user.
- Scrub location and web history from user facing database.
It saddens me now that all google apps on my mobile phone are no longer allowed to access nothing. I try to only grant access if i really need to. This is how it became unfortunately.
The reality is whether they keep the data for 3 months or 18 months doesn't mean anything. It just gives you the illusion that they no longer use your data. I find it hard to believe there's good intentions behind this (other than deleting data AFTER it's used).
I almost wonder if leaving it there has a positive effect, as my life has changed, the locations I go have changed, the products and services I use and the stores I patronize have changed, and Google is left with a version of me that is no longer accurate.
What about those of us without Google accounts? I use pretty aggressive tracker blocking, but wouldn't be surprised if a PREF cookie, ETag, browser fingerprint, or some other tracking mechanism snuck through at some point. My parents don't use any Google services other than search (not logged in), but Google likely has most of their web browsing histories going back years, thanks to ads and analytics.
If I delete my Google cookies, will Google delete my history and not try to re-identify me? It seems like "privacy theater" otherwise.
Just like so many of the privacy features put out by FANG companies, which mysteriously undo themselves upon update roll outs. Funny how I've never seen those roll outs cause the privacy features to automatically turn on, only off.
Still better than nothing, I appreciate the option, but I can't help but being VERY skeptical about this.
The question is, are they deleting that data?
I think they are big enough that, at least for EU users, not deleting that data would land them in a lot of trouble, so I think that they do.
For people concerned with privacy I don’t think Google is trustworthy enough, but these features are good to have for the rest of the world, so I’m glad to see it.
Right now we just have to take Google's word for it, but if they were willing to pinky swear (figuratively speaking), it would let people trust them more, and possibly let them offer new kinds of services.
Or you could just make making false representations to get anything of value from people a tort generally, and in egregious cases a crime as well, rather than making it a privacy-specific regulation, and without making any particular formality on the part of the vendor necessary for consumers to be protected.
I'm taking about a system that would let companies voluntarily increase their legal exposure to specific claims as a proof of commitment. That would allow companies to pick a level of privacy protection they wanted to offer, and market to customers based on that commitment, in an enforceable and credible way.
So does a proposal about a special ceremony which makes the claim binding; you still have to take action when they break it, and prove that they did.
> I'm taking about a system that would let companies voluntarily increase their legal exposure to specific claims as a proof of commitment.
Increased exposure in what specific way? Lowering the standard of proof below preponderance of the evidence? Keeping that standard but allowing statutory minimum damages? Adding a damages multiplier?
If you don't like that approach, there are other ways to make the idea work (or, since this is HN, "well actually" it into a fine powder).
I understood the proposal to be making corporate lying illegal, such that proving damages but not be necessary.
Also, are other "global observers" deleting data extracted from Google?
In practice I assume that data has highly marginal value for them so the small percent of people who bother to use this won't matter much.
I know how bad it is privacy wise but I get so much utility of being able to see all the places I have been on a map, it's so wonderful after travelling or for finding places in my hometown I haven't yet been to.
Source: been trying (off and on) to root a device since 2015. By all accounts, it's not ROMable.
This being a key reason I've all but entirely soured on smartphones and tablets of any description, though Purism are looking interesting.
And, sure, it's deleted from Google. Is it also deleted from ads networks after Google shared the data with them? I doubt about it.
If it has enough information to append to a shadow profile, it has enough information for me to request that the entire shadow profile get deleted.
Even when the data is deleted, I assume the profile stays (as far as I can tell from the meager writing so far), and can continue to be developed with further info, even if all the old info has been deleted.
At the end of the day, the profile is more important than the data that was used to build the profile, so its great for Google - and it doesn't help me that much that my data was deleted.
In addition, Google could save in that profile all sorts of useful metadata (how many emails, from what range of countries, etc.) that might someday be useful.
Other commenters have stated that Google's algorithm currently rebuilds the profile when any data has changed. Obviously this will have been fixed (...to include the fact that the user wants deletion in the profile) before this feature gets rolled out.
1. Only a tiny percentage of users will use this, so the benefits in a legal sense are big compared to the loss.
2. The only loss is that future development will have been able to squeeze more from that info, and the metadata is enough to offset most of that risk.
It would be good policy at Google to do / be ready to actually do the thing so that if called to / required to... they can quickly say "done!".
A CEO sitting in front of congress would certainly like that.