Why Google Takeout is sooo bad
marcin.cylke.com.pl
marcin.cylke.com.pl
I will caveat that though, I am bullish on Google and AI in general given the incredible talent and vast amount of data Google has access to. I think eventually they will make a technological breakthrough that puts them back in the leader position - they had it with transformers and just didn't know how to turn it into a product.
As for a breakthrough to leadership, Demis' approach of using successively more complex video game environments to develop AGI is absolutely the right path, so I wouldn't be surprised if DeepMind generates a prototype "AGI" first, but I would be VERY surprised if goog successfully capitalized on that.
Huh!
That's probably a good reason to avoid organisations of these sizes unless you have a good reason not to, but ... that applies to any organisation of comparable size.
However, they way they approach it varies considerably, and I'd expect Google engineers to address it in a way that minimizes problems resulting from, say, changes in product A causing problems in product B. I'd bet some work on that has already been done because these folks aren't stupid but apparently that's not enough.
If you go this route, then you'll need to organize your media in whatever photo manager you use, and then again on $BIG_CLOUD. Yes, your photo manager will sync some things like titles, comments, and tags as you upload new media, however not all things are synced, such as the event(s) that you want your media to show up in. Also if you make a change in your local library to media that's already been published to $BIG_CLOUD, then those changes will not be reflected there.
Personally I use Shotwell under Linux: https://wiki.gnome.org/Apps/Shotwell and I wrote a program that generates a static HTML site based on my library: https://github.com/masneyb/shotwell-site-generator. When I make a change to my media library in shotwell, then the static site is regenerated to reflect the most recent version of my site. This also makes it super easy to backup my photos to $BIG_CLOUD (like Amazon S3) for redundancy, while retaining full control of my media.
I have my generated site on a password protected website that my family has access to. When I need to share photos with friends, I'll upload them to a photo hosting service like Google Photos or Flickr.
Combined with the tech (RAID array, backups, sharing script), it also helped to have a manual practice of culling photos.
I didn't cull as selectively as I might pick photos to cold-submit to a publication. But if I had several almost identical images from the same event in my archive, I'd try to delete all but one of them.
Reducing space requirements to 1/4 has home IT benefits: maybe don't need that NAS or bigger drives yet, backups run 4x faster, backups might fit on a single backup medium or much less expensive one, can afford that second big local drive for a little extra RAID-mirroring protection, etc.
It's also good encouragement to be a little more judicious about pressing the button on the camera that makes more culling work. :)
(Personally I use rsync (FolderSync) from the phone to a home server and pull things into my (kphotoalbum-oriented) workflow that way. But it's more about personal control and paranoia than being something that would be useful to random people for whom vendor lock-in is actually a threat...)
Curation is key to avoid having a mess down the road since I may take 10-15 photos of the same scene, and only save the 1 or 2 that turned out the best.
Even for us, Google Takeout is a complete mess that fails all the time, or straight up corrupts files that need to be exported. It doesn't surprise me at all that the service sucks for general users, but the fact that it's terrible for Enterprise customers really tells you all you need to know about Google.
Similarly to the OP, we also run into issues with logs wherein the Google Admin Console also just straight up doesn't provide detailed information about what went wrong, and questions via our relationship manager often get passed around for what feels like years.
Thanks, Google.
Also a trick to download which avoids using wifi is to rent a tiny VPS. Then do takeout -> email link (not google drive). When you go to download it, open chrome devtools -> networks tab. Click "download" and find the networks row which is for that download (is preceded by a 301 redirect). Right click -> "Copy as cURL" and then can copy into that VPS
It's a way of spoofing the connection with cookies and everything so you download over the fast VPS connection. Then can use `rclone sync` or whatever to copy to a backup place
I’m no Google proponent but I hate people who spread outright wrong information that they could just double check.
If you want to download a snapshot every so often, you should be able to give it a token representing the last timestamp, and have it give you all the things that have been added since then - most datastores can do this. Then, if the service supports it, it should also indicate changes to old data (including deletes) - but this requires some additional state so it probably wouldn't work for every case.
I use dedicated tools for backing up my most important Google data -- rclone for my Drive, and gmvault for my Gmail. And I'd use gphotos-sync if I used Google Photos (but I don't). And around once a year I use Takeout for the rest -- stuff like Calendar, Contacts, etc.
Takeout doesn't really fit the author's use case. It's intended for migrating your data to another service, not for regular backup.
Once you've downloaded the entire new Takeout, there's no reason to deduplicate -- just delete the old Takeout.
My use case is a have a local NAS that i use for backup but i also want things backed up offsite, so i mirror the backups to b2 (and soon to glacier).
I would download and extract the takeout archive locally, then run borg with the NAS as the borg repo. It tries to dedup and only store incremental data in the Borg repo.
If the takeout data consistently has enough of the same “shape”, the b2/s3 storage would only grow by roughly my incremental takeout archive size, rather than storing 200 more GB every time I export a takeout.
So yah, it would use a lot of space locally and temporarily, but the idea for me is to minimize cloud storage but also being able to extract files from older takeout archives.
You don't need to keep the previous backup on the disk, it's enough to have it on the backup destination (at least in the case of borg).
But, as stated in the blogpost, I'll check gphots-sync in near future.
The problem with that kind of approach is that I need to setup my own infrastructure of some kind to run that sync app on it. With Google Takeout I can just do that to external vendor.
https://gunargessner.com/takeout
One of the tricky things is making sure that you're not duplicating things you've already backed up. The way I handle it is through a combination of date ranges and tags.
One good thing I noted in Takeout though, is it converts your google docs stuff to locally-viewable formats easier than other options (and of course a lot better than you'd get by just syncing in Drive as they are just links to the gdocs in that case!).
Another issue where google makes transferring stuff harder is just transferring stuff to another Google Drive. I had to do this from company account to my own (with a valid reason, I was a cofounder and we were shutting down). There was no built-in way so I programmed my own recursive copier that retains google docs as google docs and files as files, an also manually manages to retain comments and where they apply to in the docs. It still wasn't possible to retain edit history which sucks, considering if these were just normal files doing be most basic folder copy would do all of this.
The backup process puts both the underlying photo assets plus the DAM's DB of metadata into backup automatically and continuously. I could (always) be more paranoid, but so far this has worked. And I have successfully restored from the cloud backup after a motherboard death on my primary local machine.
I think using a 3rd party cloud system as the primary system and source of truth is making life harder unnecessarily.
This, and the fact that Google Photo's web UI is a poor alternative to a good native photo manager is probably a large part of why I pretty much don't really bother with Google's ecosystem of apps and devices at all anymore.
I don't think this can be correct. Takeout was released June 28, 2011. GDPR was passed April 14, 2016 and went into effect two years later.
I seem to be luckier than sibling comments; my gmail+photos tasks have been reliable. Three big differences from the author: I have closer to 40GB than 200GB of photos, I download the zips from Google rather than having it post 200GB+ to OneDrive, and I unzip my backups when they happen and move them to network storage, rather than rely on them to be 100% automated. I do recall one manual run failing and I was notified.
I do wish some pieces were a little more intuitive, for example how the Drive vs Docs/Sheets/Slides overlap is handled.
I know anti-Google sentiment is very high here, and disclosure that I worked there in the past, but IMO Takeout is a user-friendly effort and they're well ahead of other companies, big tech and especially smaller and non-tech, there. I might be selling Facebook short; they do have an exporter but I haven't used it. [Edit: of course my opinion would be different if I ran into broken exports like a couple sibling comments report.]
> they're well ahead of other companies
Agreed; it has never had a great UX but I'm glad it exists. I would like very much for it to be automate-able, as it is, it's pretty ADHD-hostile. I have to remember to start the process, remember to go get the files, and remember to deal with the files after they're done downloading (even on "gigabit fiber", it still takes long enough to force a context switch to something else and thus another opportunity to forget to deal with it).
Exactly. It really bothers me that people make these assumptions that all companies are bad and must be forced to do things that help users. And in this case, the author straight-up lies.
Google did NOT do this in response to legal regulations. They launched Takeout before the GDPR was even first proposed, let alone passed. They did it on their own, precisely in response to user desires not to be locked in. You're more likely to start using Google if you always know you can get your data out.
Also, to be clear - I don't see Google as inherently bad. My opinion here is just about Google Takeout.
I filed like five support requests and never got any actual replies.
My next step is going to be moving off google for most things, but amespecially for photo storage (80% of my usage) which I was relying on google takeout for a backup solution.
As a former Googler i have pretty much given up on google being useful these days. :-/
Well, that's just simply false. Google Takeout was in development and released much earlier than GDPR was published. It was an internal project of a weird Google team known publicly as the "Google Data Liberation Front". I wasn't actually at Google during any of this, but I recall Google Takeout from ages ago, and it's hard to forget, in part because it seemed so novel for the time, and also the name "Data Liberation Front" stuck out as being surprisingly bold.
(Of course, Takeout has probably expanded greatly due to fears of regulation, but according to Wikipedia it was in development for four years upon being released in 2011, and I don't think Google was as worried about data portability regulation by that point in their history.)
--
That aside, I was actually genuinely curious to see how bad the process would be for me, as I've been gradually de-googling myself for years now. I chose to export my Google Photos library, which is apparently around 60 GiB. It isn't growing, because I no longer use Google Photos. I opted to just export them to 10 GiB tgz chunks to download directly. About an hour or so later, I got an e-mail with the download links. I'm currently downloading them. The download speed is tolerable: it's between 10 MiB/s and 30 MiB/s, jumping around a bit depending on the part file (probably random chance to some degree.)
(Edit: I did indeed finish downloading the batch with no problems. Seems fairly complete. The formats are documented, not sure how much stuff supports them.)
So at the very least this seems like a "YMMV" situation. For me, this works perfectly well, which is great news considering I didn't really have another contingency plan for this.
Granted, I'm not sure what I'll do with the exported data yet.
As for the way it works for me - I've tried multiple times and it always fails - but I backup to other cloud.
Probably won’t get rid of Google entirely but definitely need to do something to derisk this.
https://github.com/GAM-team/got-your-back
It is, I believe, written by a Google employee. It's a really great CLI tool that allows you to dump mailboxes to disk as a bunch of text files. You can also restore them to another account. Lots of flexibility as well.
I gave up on Takeout and Google's own Data Migration tool, both of which I found just too flaky and inconsistent - my experience was not dissimilar to that in the OP.
I do regular copies from google takeout to AWS Glacier Deep Archive (via S3) , it's cost effective and works well. I use rclone for the reading and writing.
Just spin up a VM for the duration of the processing. You can use the intermediate step to filter, transform, prune , chunk or edit files as well.
Last I used it, probably a year ago, those options were not there. I'm 99% sure. And I can find lots of forum posts over the past decade complaining that Takeout doesn't support uploading to other cloud providers, and that it doesn't support any kind of automation.
When did Google add this?
Bear in mind that the scheduling support is extremely basic -- the only option available is to schedule six exports, one every two months. You can't change the frequency or the number.
Also, you can't pick when the schedule starts, so if you want backups every two months indefinitely, you have to remember to schedule the next set of backups two months after the final backup of the previous scheduled job finished.
It's better than nothing, but only just.
My main use case is backing up Gmail accounts of users in our GSuite org, before deleting an account.
1. Reset password and login as user 2. Initiate Gmail Takeout, export to Google Drive 3. Return to admin, and delete user account, using the option to xfer their entire Drive to admin user
If you're doing recurring backups, there's likely a better solution (e.g., file syncs from Drive/Photos).
A few weeks ago I used Google Takeout to download more than 40TB of Google Drive data, compressed to 50GB zip files. It worked perfectly.
It won’t get them promotions, raises, or even kudos if they make takeout really awesome. It doesn’t generate revenue. It only costs Google to have it. As long as it meets the minimum to satisfy regulators, it won’t be touched.
Software tends to rot, though, and software at Google double so.
These guys who initiated the effort could mean well and put all their efforts into it. That doesn't invalidate the parent's point a tiny bit -- nobody is incentivized to make it good today.
Yes software rots, but some much more than others. Look at Gmail, Search and Android. See the difference? Some generate revenue, others don't.
The problem is more likely a lack of prioritization at the leadership level.
Great. If you're in leadership why would you prioritize a feature making it easier for people to leave your platform when you could instead prioritize a new feature that might generate value for the company?
There's no incentive to make Takeout any better than it is today.
Yeah, but there better be a high-up patron for these products because Google is notoriously stingy with promo.
Source: quit Google right after L3->L4 because another company was willing to offer me an extra $200k/yr and L5. I've since been promoted at THAT job, and am now looking again because they gave me a raise to the bottom of the next band and that's dumb.
Why is that dumb? By your own account they’re already paying you at least 200k/year.
There’s a limit to how many large raises you can give if you intend to give x% for the rest of time.
The tl;dr for this is that if the company makes getting paid a market rate and promoted internally more difficult than just interviewing, they should expect people to just leave.
I'm really not sure how the economics of this work out. Obviously Google has a much easier time swapping engineers in and out (it's responsible for basically everything that people hate about the company both internally and externally) but there are still specific teams where engineers leaving represents significant knowledge loss.
Hell, companies that DON'T follow the same engineering practices that let Google hotswap engineers still do this and there's no way it doesn't have significant hidden costs for them.
If you still expect 200k raises in such a position you are likely in for a bad time (nothing is guaranteed though, you might get lucky).
This is not even remotely close to true for high-paying SWE jobs in the Bay, and the part of me that grew up in the middle of nowhere still has a hard time believing it.
Staff engineers make between $500k to $1M a year. My first year at Google I was sat in a row of absurdly high-level engineers, many of whom made even more.
At Google is already not the norm is it?
Google Takeout existed long before Google Cloud existed, and before GDPR existed. At any rate, it was novel at the time, meaning other vendors were not so compelled. So this is an obviously incorrect statement.
I suppose the article author no longer has the original photos? Because the obvious solution is to start from the originals, and store them in the multiple places.
Lastly, this article is not why. It's how. That's too bad because I am more interested in the why.
The manager in charge at the time (a) hated the project and didn't want to take responsibility for it because it was a steaming pile of poop and (b) Google Takeout being a steaming pile of poop served Google's interests of keeping users locked in
its almost as good as photo's, better in some ways.
Next will be emails. though thats going to be a while away.. hosting emails sucks.