Google will delete new accounts' data by default in 18 months
wired.com
wired.com
I don't think the debate is whether this data can improve your experience. The debate is in how the data is leveraged.
Watching the WWDC 2020 keynote, lost count of how many times the presenters referred to "on device" computation to create the same improved experience you're describing, but without sacrificing privacy.
When Google does it, we know the data is being leveraged not only for the betterment of our experience, but also for the betterment of their advertising products.
> It works like this: your device downloads the current model, improves it by learning from data on your phone, and then summarizes the changes as a small focused update. Only this update to the model is sent to the cloud, using encrypted communication, where it is immediately averaged with other user updates to improve the shared model. All the training data remains on your device, and no individual updates are stored in the cloud.
> https://ai.googleblog.com/2017/04/federated-learning-collabo...
I guess this just doesn't really bother me that much. Honest question - why should it?
I generally filter advertising out of my consciousness. I've clicked on exactly one ad in my life (that mystery box company). Maybe if everything advertised to me was that interesting (to me), it would be a net win for everyone?
Advertising metrics continue to prove that they are wrong.
HN will continue to have "contrarians" who think they are smarter than everyone else, and who reflexively argue contrarian positions, like "privacy is bad, actually", not because those positions are smart, but because the speakers are addicted to being "clever" contrarians.
Then you can get 2% of your ads clicked on.
All that said, I'd be interested in learning if my layman's perspective is garbage (which, let's be honest, it probably is).
I'm in the custom apparel industry, where a lot of purchases are (were - it's dropped off) driven by facebook ads. If Facebook ads were as effective as the HN zeitgeist seems to think, there would be no other kind of commerce.
Practically every commercial interaction (and most social interactions) are "manipulation" of some sort. At the moment we're trying to manipulate each other into adopting the other person's perspective. I don't understand what you find insidious about sales pitches; I'm sure they've been with us almost as long as prostitution.
Because of the impact of advertising on everyone's lives and the society at large.
An incomplete list of issues: http://jacek.zlydach.pl/blog/2019-07-31-ads-as-cancer.html.
https://9to5google.com/2020/06/11/google-photos-explore-map/
In general, I think "privacy as a right" is a better approach for society. If people want to opt out of privacy, they should have that ability, but that should be a choice.
That is not to say your device can’t keep a location history without sending it to someone’s server.
Want to remember your e.g. hotel room number, parking garage floor/zone, or bus stop? Snap. Backup your plane ticket? Snap it. Record the wiring before an upgrade? Snap. To input your dietary nutrition later? Food picture. Part/model number? That's a picture. See behind something? Etc, etc.
The fact you only think smartphone cameras are for "bathroom selfies" (?) says more about you than whomever you're trying to disparage, because you're obviously under-utilizing an extremely valuable tool in your pocket because you've put it into a tiny box and dismissed it.
I don’t need to remember my hotel room number or parking garage zone or bus stop every day. Even when I do I don’t need a photo, sorry, got a pretty good memory, and a digital notebook too. Under-utilize, sure, whatever, it’s my fucking phone and I didn’t ask you to judge. The fact is me and a whole host of people I know who don’t snap photos frequently enough (yeah, been talking to quite some friends about how we’ve basically been spending a fair bit of money on cameras that we don’t use very much) won’t get “the same functionality” as you were claiming, so I hope you’re not on a design panel for something like this where people who don’t conform to your expected usage pattern just get a FU.
Also: who the hell keep those parking garage photos around?
I do! Although, that's because I don't take a lot of pictures, and I look at them even less, so they just accumulate. But it's rarely become a space issue (because I'm not taking many), so whatever. The act of taking the picture seals the memory, so I don't usually have to look.
Storage is cheap. Having the image for some currently-undefined later use can be invaluable. Never using it later is cheaper than the opportunity cost of knowing it was there but now isn't.
When I go to a restaurant I don't like taking pictures of the food, or pictures of the restaurant. I like enjoying the food and the company I'm with - but maybe I'm regressive.
YouTube watch data is used to personalize recommendations, and browsing data is used to personalize ads: https://policies.google.com/privacy#whycollect.
Disclaimer: I work at Google.
https://myactivity.google.com/activitycontrols
The location history setting toggles whether data for My Timeline is kept.
Or the algorithm knows I like David Bowie so recommends me "Changes". The most boring possible recommendation it can give to a David Bowie fan.
I really wish it had a setting to up the daring of the recommendation. It seems designed to just never recommend something that I would really hate but in doing so it never recommends me anything all that interesting.
Edit: Turns out it doesn't actually provide a location history visualization like GMaps does - it only shows "significant locations" - but to my knowledge there's no technical reason it couldn't do so
You don't have to send any data to a central server at all.
I too really like this data, specifically location history. It's come in extremely handy for a number of reasons, whether finding a restaurant I really liked but forgot the name of, finding exact dates and times when I was in certain places (which is handy to know if said place had a COVID-19 outbreak, for example).
sounds like a feature (for them) to me. why bother showing your users new videos when they're perfectly content watching the same old videos over and over again? Better than not showing anything and them leaving your site because there's no more relevant/new videos.
You should be allowed to make trade offs between privacy and convenience where it’s available, but it should be transparent exactly what tradeoff you’re making.
Somewhere along the line, we stopped caring about trying to manage our own information and handed it over to big cloud companies like Google and Facebook. Now we’re seeing the downsides of that tradeoff.
Adding a tag to bookmark: https://i.imgur.com/iVElOWb.png
Searching by tag: https://i.imgur.com/e3UmQDG.png
I don't use any bookmark management addons, so this is definitely default behaviour.
Imagine if that hadn't happened! I think we would have amazing tools for managing our information locally, in a highly automated and intuitive way.
The seeds of these possibilities are everywhere. Search, of course, but also things like tagging and even automatic tagging with with machine learning models. Imagine if Chrome had the same tools for organizing bookmarks as Gmail has for tagging email, plus all of the power of recommendation engines under the user's control.
No, instead Google keeps bookmarks extremely underpowered because they'd rather have people search for everything all the time in order to drive traffic to their ads.
But quickly web designers decided that the conventions were boring or ugly, and changed the default color schemes and removed underlines and borders. The web as a whole began moving towards CSS and JavaScript and more full-fledged web applications.
Putting the disappearance of purple links on cloud companies and Facebook and Google is rather preposterous.
I know because I was there trying to learn programming, or anything, from the web of 1996-2000 and it was a freaking pain, a garbage dump of poor content filled with banner ads and popups and searching on AltaVista and Yahoo! was completely futile. The only useful content I found was not on the web, but in FTP directories. And learning was all you could (barely) do.
People were doing bookmarks because it was rare to find something good, so remembering stuff that wasn't crap was worthwhile.
People tend to glorify the old days. That's because memory starts playing tricks on us.
https://developer.mozilla.org/en-US/docs/Web/CSS/Privacy_and...
I believe we should have control over our own data, and it looks like Google is giving us that.
So I'm glad we can choose to enable/disable this ourselves.
There is no trade off because you can have both! It's just that Google doesn't offer this.
One thing I've noticed is that Google must obviously do a lot of filtering and correction to the data, because my data is nowhere as clean as I've seen on Gmaps, and I haven't found a way to replicate that in any way. My data is totally sufficient for whenever I've needed it, but it's not perfect. The most I'll do is filter out data that is ridiculously out of bounds from what the data directly before and after say, but I wish there were a better solution (and if there is, I haven't found it yet).
I think one big problem with Google collecting all this data is that you don't own it. You can't choose to use this location data to solve your problems in a way that you'd prefer. Instead, you have to hope that Google has covered your needs for the data and if not, tough luck, you're stuck doing a lot of manual work.
If you want to go the extra miles, I remember this talk of Ilya Zverev "Hundred thousand rides a day"[1] seen at FOSDEM 2019.
[1] https://archive.fosdem.org/2019/schedule/event/geo_gpxtraces...
You wouldn't be in the minority if you thought of it this way.
If after an exam I would spontaneously say: “I absolutely did not cheat on question 3 of the test!”, what would you assume I’d have done on questions 1 and 2?
For one that's illegal in the EU at least. Due to GDPR Google cannot use your data for profiling without explicit consent.
Thinking about "what they don't tell you" is how conspiracy theories work. Such claims are either evidence based or they aren't.
And don't get me wrong, if you don't trust Google for whatever reason that's fine, I just find your reasoning odd.
And presumably, EU users will get an extra consent screen going into details. But as far as I know, this doesn't imply these details have to be the part of the document called "privacy policy".
> Thinking about "what they don't tell you" is how conspiracy theories work. Such claims are either evidence based or they aren't.
Absence of evidence is only weak evidence of absence, but it's still evidence. In this case, you could call it "being suspiciously specific", which is a known technique of lying by omission.
Privacy policies don't exist in a vacuum - they're formed by entities strongly incentivized to make money every possible way that's not explicitly illegal[0]. So a suspiciously specific privacy policy can be reasonably considered worrisome, especially for a company known to make money off their users' data.
Conspiracy theories are excessively using the "what they don't tell you" reasoning, but that doesn't mean the reasoning itself is bad. It's just one requiring finesse and a somewhat more complete picture of an issue at hand. The best conspiracy theories are ones that balance this reasoning - any more, and they'd be trivially debunked; any less, and nobody would find them interesting.
--
[0] - And some even achieve success by crossing the line into squarely illegal behavior - see e.g. Uber.
That would explain the unescapable modal "These are our privacy terms" popup that every Google service displays the first time you hit it with a fresh browser...
Obviously nobody reads anything, they just click "Agree".
It's all so broken.
Worthless, commercially - yes.
But priceless and utterly invaluable for historical interest and academic research! I hope Google keeps their Ngram, search-stats and even geolocation history for everyone (anonymized and aggregated, of course) - it'll be an invaluable tool for researchers of all kinds to be able to see, for example, how our movement patterns varied from ~2015 when we started collecting that en-mass to some point in the future.
Ditto Google Street View archives - Street View is now almost 13 years old and I'm looking forward to being able to see how things were in everyday life back in 2008 compared to however things are when I'm old and grey).
Aggregated data is useful for building models, but individual user data is not useful for targeting that particular user.
For example, suppose you are ranking search results for Alice. It may be beneficial to include old data in your language model or indexing models. However, it would likely not be beneficial to include Alice-specific features that go back more than 30-90 days. Such features include counters for her actions (clicks, queries, scroll), embeddings of her activity, etc.
On the other hand, site and query specific features would likely be useful for years.
For an average user, sure. But suppose the user suddenly becomes e.g. an important politician. Now having their history older than 90 days is interesting.
Hard disks are cheap, you can keep everything and wait for the ticket that wins the lottery. If you are a company like Google of Facebook, it is almost guaranteed that you will have juicy material about most people who become important in a few decades. In a few decades, any politician who says "anti-trust" will immediately have their porn history or edgy things they wrote as a teenager leaked to public.
I'd think thats more valuable - raw data would not really be necessary unless they need to re-learn things.
I've used it for "When did we go there?" Type questions, or "What was the name of that restaurant?" And also, just sitting down to look at the history and see "What was I doing X years ago" brings back memories. I think it's neat to see something like, I went to the theater three years ago? What did I see? Ohhh, I bet it was X! I feel like it's something like an automated diary of my life.
On the Photo map you can see your photos grouped by location, on the normal photo timeline grouped by year/month/day and the ML powered search is good enough to recognize stuff like “book”, “restaurant”, “beach” etc.
It seems doable, but also like it would take some effort to get to be nearly as good as Google is now. I get why some people want to be private from Google, but it doesn't really bother me.
On the other hand, Chrome remembers some thing for far too long, sites that I haven't visited for a long long time still exist in chrome://settings/siteData?search=cookies+and+site which is rather annoying.
Otherwise they would have custom cut-off capabilities, and wouldn’t basically force everyone to turn collection back on at every turn after you’ve disabled it. I explicitly did it three times before giving up, as Assistant would turn it back on every now and then. That’s when I decided I’ll make an effort not to use any G service if I can avoid it.
Bob: sends the chart
Alice: Carol, could you please chart the amount of data by creation date and send me alongside with our storage costs?
Carol: sends the information
Alice: does quick cost benefit calculation, fires up an email to Diane in PR
Alice: here is how we spin it...
It'd be like you deleting a single doc file.
The original post from Google had this line https://www.blog.google/technology/safety-security/keeping-p...
"As always, we don’t sell your information to anyone, and we don’t use information in apps where you primarily store personal content—such as Gmail, Drive, Calendar and Photos—for advertising purposes, period."
Wait, what? Is this some weasel word soup going on?
I do not believe this for one second.
What about that is believable to you? You just take them at their word and call it a day?
They do, however, use your data to make money by slotting you into segments and selling the anonymized _segment_ data to advertisers. It's still extremely invasive, just in a sorta subtle way.
Selling your data directly is the sort of behavior I'd expect out of companies that are evil and don't care about being seen as evil (see: wireless carriers, "location intelligence" companies, etc)
I hadn't used my account in a while and was going to just "delete" what I could and keep the account. They have a "download your data" button but refused to let me use it claiming they couldn't be sure it was really my account (but not providing any way to convince them), so at that point I decided to delete. Hopefully they did actually delete it. For me, it still showed the older history after I enabled the "delete" older history option (possibly it would have been removed after some delay), so I manually deleted it (it wasn't all that much) and that did seem to remove it.
- Write your own mobile app that has access to fine grained location, logs it, and periodically uploads the location trace to your own server.
- Download the traces from your server and convert them to an appropriate format.
- Upload the traces to Google MyMaps - or use a map rendering engine you control if avoiding 3rd party online services is a goal - for visualization
By doing this you give up additional features provided by Google Location history - things like semantic place detections (i.e you were at airport X or mall Y) trace error correction, auto-segmentation of traces, etc.
But if all you care about is a seeing a trace on a map, then that should work.
This general approach is already widely used by software that has an external location logging system (i.e. wildlife tracking transmitters, commercial vehicle fleet tracking) but need to visualize/track those on a map.
Disclosure: Google employee but not currently working on Maps.
Use GPSLogger[1] to record the position. It can auto-upload to pretty much anything from DropBox to SFTP (FTP over SSH). It can auto-start on boot, run in background, and rotate files daily. (If you do SFTP, use ForceCommand internal-sftp in sshd_config to restrict the phone to file transfers.) No excessive energy usage.
Use GPSMapper[2] to visualize the GPX files on phone, or https://www.gpxsee.org/ on Linux (this tool works well even with hundreds of GPX files loaded, unlike all others I've tried).
None of the viewers are really "nice" unfortunately.
If someone needs a side project idea, a personal location history visualizer would be nice. All existing tools focus on individual tracks (eg day hikes), there's nothing that does much useful when you dump a few years worth of history into it. Actually most just freeze or crash. But there's so much that could be done (performance, filtering by time, exporting/sharing specific parts, postprocessing like removing outliers, correcting cut corners by matching against a map, detecting places/businesses and hyperlinking to services like AllTrails, visualizations like speed over time (this part is common) but auto-detecting and showing this for all traces along the same route (no software does this yet) (think "how fast was I on which part of my commute each day over the past year"), multi user eg. family capability, etc.)
Of course even if viewers are limited today, one can still collect the data in case someone writes a better viewer tomorrow ;-)
Google Takeout can export raw location history, and there are tools to convert that standard GPX.
Only issue I have with GPSLogger is that the accuracy filter is not adaptive. At a high value, you'll get outliers, at a low value, it won't record during flights.
[1] https://play.google.com/store/apps/details?id=com.mendhak.gp...
[2] https://play.google.com/store/apps/details?id=com.mendhak.gp...
Anecdote: I was really happy about the data Google maps saved about me a couple of weeks ago when I got a letter from the migration agency which asked me to send in a list of my journeys abroad since 2017. The list was supposed to have: date traveling out, date coming back, country and purpose. Last year I was abroad 19 times and 126 of the 355 days, without Google maps collecting the data in the background for me I would have no chance in hell to get it right, and I guess the migration agency has access to this kind of data too so they can double check if I was telling the truth or not.
But in 2018 when Google introduced the possibility to mark my data as deleted, I clicked the button and lost the ability to see the data for 2017 and early 2018 which I really regretted now, because I needed to try to come up with dates and countries according to photos on my phone which had GPS coordinates and my DSLR which had just the dates.
OTOH, if you claim you've gotten rid of data to one party, and then another party is able to dig that same data up in a legal proceeding, I expect the first party might be upset, and seek compensation.
Quite apart from that, if you have metrics that show long tail data are infrequently accessed, but do take up an appreciable (and ever growing) amount of resources, it's a OpEx/engineering win to actually get rid of that on a proactive basis.
So, you'd need to vet an ultra secure team, who is not allowed to talk about this data both to use and store it. Or you have to publish a representation to wider teams telling them 'these are secret feautures'. So, you'd need a secret code base nobody is allowed to talk about.
Corporations can do shady things, but it's hard to imagine this old data being useful anyway, to be worth all this hassle and possible repercussions.
Does this mean any data older than 3/18 months is deleted or that every 3/18 months they delete everything?
P.S. I assume that they may have already put a price on 3/18 month old data and decided that by then it has lost enough of its value to make such a loss worth it in order to buy a bit of good will from users and regulators alike.
so if you set it up for 3 months today, then any data on and before 3/15/2020 will be deleted
It looks like I have to do nothing but I'm not sure.
analytics, recaptcha, fonts, JS libraries.
Except they've inserted themselves at every layer, to the point where you can't avoid your data going via Amazon/MS/Google.
Just the way the US wanted it :)
I'm a machine learning noob but it feels like they could train a vague representation of you with your data to throw products or sites against and get a thumbs up or down.
Note: there's multiple auto-delete's to select, one for Web & App Activity, Location History, Youtube history.
People freak out about what they see in the Google profile, but you have to give them credit for being transparent about it, even before GDPR came into effect.
Presumably, the overwhelming majority of HN users already have an account.
https://www.propublica.org/article/no-warrant-no-problem-how...
On topic: this just in, person (corporation) who has lied and abused your privacy many times says they promise not to do it anymore, more at 11.
This seems to contradict "the right of the people to be secure in their persons, houses, papers, and effects, against unreasonable searches and seizures, shall not be violated, and no Warrants shall issue, but upon probable cause, supported by Oath or affirmation, and particularly describing the place to be searched, and the persons or things to be seized"
"The government can also get older unopened emails without notifying the customer if they get a court order that requires them to offer "specific and articulable facts showing that there are reasonable grounds to believe" the emails are "relevant and material to an ongoing criminal investigation" — a higher bar than a subpoena."
Looking back, I always felt something missing when I think of him, in comparison to Larry Page.
Right at this moment, it felt to me that Sundar is meant to be another typical professional manager for a "conventional company". The primary issue is that he, like other manager, lacks the inert urgency for their task.
Apple has been playing the privacy card at least 2 years ago. That by itself threatens Google at its core. Yet I never see the decisiveness in rethinking Google's identity; it was always reactionary, 2nd timed, backward looking...
Google is about to decline. Hopefully that will be something like the 90s' microsoft, and waiting for next revival...