Google Drive is a backup solution. I can think of a lot of reasons to have a lot of small files, whether they be for logging purposes, analysis, training data, etc.
In fact, I can think of more legit uses than nefarious ones! Especially since an illegal use would be easy to trace back to you personally!
What normal user has millions of typical data laying around? Not many, if any.
Enough that this policy is impacting some of them.
Again, realize that this isn't just talking about free accounts. It's paid accounts, too, regardless of the plan size. 2TB... 5 million file limit. 20TB... 5 million file limit. 1 plan with 10 users... still a 5 million file limit.
If you can't see that this could be a problem for normal customers (or normal accounts that may comprise multiple users), then you're not thinking creatively.
I wouldn't doubt that I have millions of files backed up on various drives, simply because I back up every system, and I've owned computers since the 90's. And I'm not intentionally trying to generate data!
Edit: why the down votes? It was an honest question.
Backup is a built-in function of the Drive desktop app (previously, “Backup and Sync” was one of two separate Drive desktop apps.)
Backup to drive is a built-in core feature of Android.
I have about 15M files in my Dropbox. The limit of my box is the size of my files combined AND their own size not their number. "Files uploaded to dropbox.com must be 50 GB or smaller. All files uploaded to your Dropbox must be smaller than your storage space. For example, if your account has a storage quota of 2 GB, you can upload one 2 GB file or many files that add up to 2 GB. If you are over your storage quota, Dropbox will stop syncing."
Only things I can think of are mapping tiles. Because while logs can take up space they don't generally create that many separate files...
Locally you could process those pretty quick. Or in cloud block storage. But the amount of latency involved in each cloud API operation to create a new file... yikes.
I can understand why a company would care about the access policies; if you are storing some kind of database that is accessed constantly and routinely then it's nice to know -- but when did we become okay with categorizing bits beyond that?
I don't really care for the future of needing a subscription to backup cat photos, specifically. I'd rather bits be bits except in limited instances where the data is priced based upon access frequency/retention/guarantees. (yeah, it stands to reason that cold storage that has limited access should be a fair bit cheaper than hot storage where the data is ready to go ASAP.)
This kind of thing just accelerates the stenography arms-race wherein people try to store their terabytes of data as cat pictures.
Just for performance, if you insist on backing up the content of 5M files to Google Drive, you really need to write a custom script to tarball and compress whole directories on-the-fly or something.
I mean, 5M files is like multiple sets of map tiles for the entire world. Sometimes people definitely need to work with those, but using cloud backup through the Google Drive API interface is just not the way to go there...
But again, if you ever want to retrieve all 5M files, I hope you're willing to wait a few weeks-to-months, because that's what it's going to take...
Which is why it's a terrible idea not to tarball them first to something more manageable in terms of latency per file operation.
FAT16 has a limit of 65460 total files. FAT32 limits 65535 files in a directory.
The RISC OS ADFS filesystem had a limit of 77 files in a directory; I used this until about 1994.
There are also limits to number of files in a directory.
In the Python word, I just created a new virtual environment on Linux, and that had 1339 files. I then loaded it up with a few standard players (numpy, pandas, scipy, scikit-learn, matplotlib, dask, and django) - this ballooned the file count to 16,945.
If you can't imagine how that might inflate, I don't know what to tell you.
What on earth are the main contributors to that file count? I can't even begin to imagine, but I'm incredibly curious now.
Also, that seems like an enormous waste of disk space. Assuming a 4 KB block size means roughly an average of 2 KB of space is wasted per file, which for 700K files would total well over a gigabyte of wasted space. Yikes.
If you have 5 million users, you should be using a real system
https://issuetracker.google.com/issues/268606830?pli=1
> What it should say is that no matter how much you pay you have to adhere to the 5m file limit.
This is about a single user. No user should be storing separate files for 5M customers in their individual Google Drive!
That's what databases and storage solutions like S3 are for. Use the right tool for the job.