Google Takeout: Export a copy of your data in Google services
takeout.google.com
takeout.google.com
Any organized efforts to export and donate data? Data is not data in some important senses unless it's aggregated. A "watch-them-watch-us" dynamic can't exist if one side is aggregated and the other isn't.
"You have no available activity to show at this time."
Maybe because I've never had a habit of linking phone apps, web apps, etc to Facebook. I tend to keep different accounts completely isolated. Which is why Google really bugs me, I don't like that my YouTube, Gmail, etc are all the same account.
You could probably try Firefox containers for that.
> Many people have more than one Google Account, like a personal account and a work account. Uses like that are fine.
Huh. What happens if you already have multiple accounts, and one of them gets banned?
Unfortunate that it pops up with a blank alert box, and then on clicking the button to proceed it informs me that it "Failed to clear history".
You can browse all the various data here: https://www.facebook.com/your_information/
This is the download page: https://www.facebook.com/dyi/
i've tried the takeout option about a year ago when i wanted to switch from google photos to apple photos and it was an absolute mess, it exported random zipped folders, all broken up (i think there were 14 of thema ll together), multiple images were duplicates, somehow it managed to corrupt apple format of pictures and i would have one picture with the same name that was something like 24KB and another one with 1.2MB, etc.
if it wasn't for the option (which they took away) of syncing photos to your google drive and copying them out that way, i have no idea what i'd have done, i'd have lost about 120GB of photos and videos going back to 2008.
so seriously, do all of you just use the takeout and store it away without actually opening up and trying to use it?
I've uploaded and expanded about half the archives on my Synology, and it's currently indexing everything, so I can't comment on the photo issues you've mentioned quite yet.
Overall, I'm happy there exists a mechanism to get my photos, but the quality of the experience is truly awful.
All that engineering talent and that's the best solution? Lol.
Source: working on a product engineering team for a larger company. Why would we upheave our roadmap for something that doesn't help provide explicit value to our product offerings? We've got a list a mile long of improvements and new capabilities to deliver on, and no engineer likes working through legal hoops anyways.
As for no engineer liking working through legal hoops, my observations disagree. Many absolutely don't care. Give them a well-described Jira ticket and they'll happily chip away at it for a year if it's necessary.
I downloaded the 7z without any problems.
I'm exporting it regularly (although I certainly don't have 120Gb of photos on it). You can choose an option for regular exports (every two months), and delivery method (e.g. Google Drive). Then I have a script that runs daily, mounts google drive, moves the takeout locally if it's present and removes from google drive (so it doesn't take space).
Then indeed, inside you have mess with some data in HTML, some in JSON, etc. But well, at least you can parse it... I have a library which I'm using as an API to various data exports, in particular, archived takeouts too (so I don't even have to unpack them to access)
- https://github.com/karlicoss/HPI/blob/master/my/google/takeo...
- https://github.com/karlicoss/HPI/blob/master/my/location/goo...
- https://github.com/karlicoss/HPI/blob/master/my/media/youtub...
Described this in more detail here: https://beepb00p.xyz/my-data.html#takeout
how is some regular schmuck that wants to move his data out of google to another service supposed to determine what they actually have to parse? the user simply uploads pictures into the system but gets garbage out?
the scarry part here is that google makes it extermely easy to suck in the data but for an average user it's extremly difficult to get back out and takeout is absolutely not a good solution.
If said person wants to view the data on their own time, HTML seems adequate. And JSON seems ideal if they plan on sending this data to a new service that ostensibly supports parsing Google's takeout.
Like mbox is fine for opening in a desktop client, but if you move from Gmail to Fastmail or Outlook or whatever, mbox might as well be a ClarisWorks spreadsheet file.
$ ls -lh takeout*
-rw-r--r-- 1 ben ben 39G Sep 26 18:29 takeout-20200925T172738Z-001.tgz
-rw-r--r-- 1 ben ben 38G Sep 26 19:20 takeout-20200925T172738Z-002.tgz
-rw-r--r-- 1 ben ben 35G Sep 26 20:06 takeout-20200925T172738Z-003.tgzThe data is not uploaded and instead parsed in the browser.
For the Google takeout make sure not to include data from Photos, Gmail, Youtube and Drive as they make the export too big. Also select "JSON" for "My Activity".
Feel free to try it out yourself:
Every photo has an additional json with some kind of metadata.
For me it was fine.
It does allow you to send the takeouts regularly to google drive, but the amount of data I have is too much for that / fills my drive.
Has this improved?
One of the positives of the Google+ shutdown was that Google Takeout saw a major overhaul around February of 2019. Third-party tools (Alois Bělaška's Friends+Me was invaluable https://blog.friendsplus.me/) still proved very useful, and gave capabilities missing from Google's offerings.
Dealing with data exports remains challenging.
I used a simple one-liner find/mv command to move all JPGs into the same folder... which made it more decent.
I rsync my Google Drive folder to network storage, without deletion, and I have OneDrive syncing that. So I have a local and cloud backup of everything that's ever been in Google Drive.
- the export size is around 176 gb. Mostly photos.
- it has the option to move to one drive or box. But 100 gb on Google will be 200 gb on one drive. Images are copied into multiple folders to recreate the albums. Note that google photos automatically create albums for family and trips.
- tried to use 2gb zip to split the files. We have to click and download 100+ pieces of zip files. Even if one file is corrupt we are done. All this shows in a modal window. We can't download more than 5-8 files at a time.
- split it into 5gb zip files. Now download numbers are manageable, but the network keeps dropping and we have to download again. We can only retry 3 times making the entire set useless.
- no options to separate videos and photos.
- we only have a week to takeout and test the whole thing.
TLDR; it's designed to make sure that we don't actually take out the files...
It's pretty clear they were not super concerned about making it a user-friendly process.
It’s not worth worrying over except as a reminder that rating systems need to handle bad-faith voting.
EDIT: It does seem that there are more people than I expected for whom the experience isn't as good as the one I had, so maybe Google should test this with (and make it work better on) internet connections that aren't as good as their office lines.
And I don't particularly want to hear "get a better network connection". My connection works just fine for everything else.
That said, if you're able to download everything, it's reasonably well-organized, although as stated elsewhere, there's a lot of duplication of data.
Plenty of solutions.
1. The interface requiring multiple downloads prevents automation or simply waiting out a large transfer, and not having a robust automated retry mechanism ensures wasted time and increases the odds of data loss.
2. Few people have a high-speed free WiFi network nearby. You’re not getting better results at Starbucks or the local library, and Google’s campus networks require logins even if you are one of the few people who lives near one.
3. Setting up a VPS and running downloads from a web app requires money and skills most people don’t have, especially if you care about not accidentally leaking your personal data. If you have enough data to matter, you’ll also hit many providers quota limits or bandwidth charges. If you navigate all of those challenges, you still haven’t solved the problem of getting it home - at best you can now use rsync to remove the manual component of the second transfer.
LoL
I have tried to spin off an Amazon EC2 instance to download and copy to an S3 bucket. But it logs out every few minutes disrupting the downloads. It will not allow to down the same file multiple times. If one zip fails, the whole set is useless.
I understand that big file transfers are hard, but that's the entire point of the service.
There are plenty of file transfer tools that can take a list of URLs, but this forces you to do everything in the browser.
I am somewhat worried that their great algorithms will one day decide it's there is some violation of their ToS and close the account. I had that issue even with a paid (well, voucher that came with a PC, but still) Microsoft Drive account, I used in an atypical way (no sharing, all contents encrypted).
In the past I used IMAP for backing up my messages, but over time my scripts to do so have fallen into disrepair... Whether takeout would be a way to do somewhat regular backups? Or might that trigger their algorithms, that you are not a good customer? Has anybody read the ToS whether anything is mentioned about takeout?
There's a product called "null" without an icon.
It is meant as a service for backup, and migration to other services.
Yet, if you're audited, it's just a series of IDs and numbers, nothing identifying there...Right?
If you ask Spotify for your data dump, you'll notice in a lot of the .JSON files the information is encrypted such that you can't understand it (it's just numbers). It's impossible to say whether it's actually stored like this or if they encrypt it before they provide the archive to you.
Meta data is almost impossible to legislate against, and as far as I can see, is entirely legal to collect and use as you see fit.
How many people in the world are on hackernews, named "lopis", use Firefox 65, have an IP address in $country, use this screen resolution etc etc
Easy enough to identify who you are.
See first definition of https://eur-lex.europa.eu/legal-content/EN/TXT/HTML/?uri=CEL...
You are identifiable by combining multiple factors, even if every individual factor is not enough to be identifiable.
If I understand it correctly, the EU 'personal data' (PD) concept is much wider than the US 'Personally Identifiable Information' (PII) concept. You are touching one of the differences here.
For GDPR purposes, data is PII if it can be used in combination with any other data to identify an individual. Doesn’t matter if the individual data points are not themselves identifying.
One thing I’ve been curious about is whether AI and algorithms that can potentially take a huge amount of anonymous data and “identify” a user (but not explicitly), only identify in the sense that the output of the AI was only possible by correlating individuals granularity enough. I’m almost certain the answer is yes. I’m not clear on whether GDPR addresses that issue or not.
By that definition, all data is PII. There is no information available on this planet that has not been influenced by people.
I'm not trying to be obtuse. I worry about this problem a lot. Obviously we need to keep companies from doing stupid stuff like storing the first digit of a Social Security number (can't identify someone by that!) and then the second digit (also not uniquely identifying!), etc.
On the other hand, what if I have web log files that only store URL, timestamp, and status code? Is that OK? If I get hits for two specific pages within a couple of minutes of each other, and there's only one person on the planet who would know about both those pages, I know they were visiting my site at that time.
People influence the world around them and it feels like privacy laws are trying to prevent companies from understanding that influence. At the same time every other incentive is pushing those companies to understand more.
I think that is a step too far. For example, it seems quite clear that a dataset of daily average temperatures from the top of Everest is not personally identifying information.
Hence, installing a camera that counts black-haired people, another that counts people entering some location, a third counting people having a BMW is perfectly fine. Merging the 3 recorded tapes to identify a person is not. Giving the 3 tapes to someone else is only OK if you guarantee somehow they wont do the merge.
None of this is the same as a blanket prevention.
Right answer: You can go to [1] and [2] and [3] and ask them to delete your information. It's important to retain a copy in writing that they have removed your information. If a copy is ever found online (in a data breach or otherwise) you would be able to enact legal rights as a result of their GDPR breach. I would encourage people who upload data to leave "fingerprints" in their accounts, such as certain photos, emails, and other data that you have ONLY created on this service (for example, email your own gmail account a unique email, if it's ever leaked, you know where it came from).
It's the same way Spotify's GDPR tool does NOT give you all the information they store, yet if you ask via their DPO (usually privacy@) you get a lot more data, rather sneaky way of hiding their true data collection.
ALWAYS use email or a physical letter, ALWAYS get a reply by the organization when enacting your GDPR rights, your lawyer/legal authority will be very thankful ;)
AND NEVER EVER USE AUTOMATED TOOLS! The chances are, there is data that isn't included within them. For example, go ahead right this second and submit a SAR for "technical log information" to Google, this data is NOT included in their official tools and you will be amazed how much they're storing!!
[1]: https://support.google.com/policies/answer/9581826?hl=en
[2]: https://support.google.com/policies/contact/sar
[3]: https://support.google.com/legal/troubleshooter/1114905?p=pr...
This service will generate an email with a general deletion request. Modify it to ask for the specific information you want deleted.
Disclaimer: I'm the creator of this service.
Google: "Listen here you little sh*t"
You're complaining about the one way google users have to reclaim data that is rightfully theirs?
You know that if "big brother" wants data, big brother gets backdoors, not polite export requests?
I'm sure it may get used but I highly doubt it's a primary tool in any way. PRISM revealed how much custom tooling is made specifically for governments and for giving back data to them, automatically addressing their requests etc. Takeout is slow, bulky, and its audience is the end user.
Just look at Google's [transparency report](https://transparencyreport.google.com/user-data/overview?hl=...) which shows requests back to mid/late 2000.
Additionally, the scope of what's provided in response to legal process depends heavily on who and what was requested and so the software is far more complex than "download all the user's data". Investigators often aren't aware of what data these companies store and if it's not specifically requested then it's not provided. Lawyers basically copy and paste their last successful warrant/wiretap/whatever and send it to the judge because that's how the legal system works.