Cleaning up my 200GB iCloud with some JavaScript
andykong.org
andykong.org
Turns out that the API doesn't expose "file size", at least I didn't find a straight-forward way.
I think that all "photos" or "videos" are just a view of the underlying "photo or video object". If you crop a video, the full-size video will remain. Only if you export the video, it will be cropped and the smaller file size will manifest.
I guess that's why the file sizes differ.
[Edit: someone created an AppleScript to query file sizes - I didn't test it, yet: https://discussions.apple.com/docs/DOC-250000422 ]
This also helps explain why my iPhone storage always seems to be at its limits, despite my obsessive management.
Anyway, we all know foss unix ftw :P
It's been possible to create a clip from a video file that merely changes what parts of the video are displayed without effecting the data in the original since the Classic Mac OS days.
If you want to completely remove unwanted portions of a video to reduce the size without a loss of quality, there are many options. LosslessCut is a cross platform option that is both free and open source.
Yup, the Photos app keeps the unmodified original file, and then any edits/crops are stored separately. You can always revert to the original file and redo your edits. So they might be storing multiple copies of the same image, with and without edits.
Which API were you looking at for "file size"?
I was able to get the size data from Photos.app with the PhotoKit API [1]. I've only tested it with my library of ~26k items, but it was useful for getting an indicator of the biggest items. (Although I didn't think to check whether exporting a 1GB video caused my iCloud usage to drop by 1GB.)
Do you know more? The introduction says "Retrieve asset metadata or request full asset content.", but I can't find clarification when it actually accesses full content.
AFAICT, PHAsset is only metadata. When I'm downloading the full-sized images, I use PHImageManager.requestImage() and pass in the PHAsset I'm looking at [1][2]. I know there's something similar for video, but I've never used it.
You can control the behaviour by passing a PHImageRequestOptions instance. This includes an isNetworkAccessAllowed bool which controls where Photos.app will download the file from iCloud if not present locally, and it defaults to false.
[1]: https://developer.apple.com/documentation/photokit/loading_a...
[2]: https://developer.apple.com/documentation/photokit/phimagema...
[3]: https://developer.apple.com/documentation/photokit/phimagere...
After all, the file might not be present locally, so opening it should go through the Photos.app. But once you call AppleScript to open the file, you might as well use the AppleScript to comb through the database like the script I linked to earlier.
"since" -> "some". ugh. I had a bad swiping day apparently.
osxphotos query --min-size 100MB --add-to-album "Big Files"
Finds all photos, videos > 100MB and adds them to the album "Big files"
`osxphotos query --help` for more info or `osxphotos docs` to open docs in the browser.
Disclaimer: I'm the author.
The example given at the end is interesting:
> So iCloud says the video is 128MB, I download it and the video is actually 48MB, and my free storage increases by ~170MB when I deleted it. Interesting!
This suggests that iCloud isn't simply misrepresenting the size of the example file, as then you'd expect that deleting the 128MB file would clear ~128MB of iCloud space. Instead, the deletion clears roughly the space it reports (128MB) plus the space of the downloaded version (48MB): 128MB + 48MB = 176 MB - which might be close enough, allowing for rounding errors, as iCloud reports the free space (from the article's example) to the nearest 10 MB.
When you buy a hard drive or USB stick, you get a certain amount of GBs to use as you please. If you put a 1GB file on it your free space decreases by 1GB (yes this is filesystem dependent and you might lose a few KBs for metadata, but the choice of filesystem is up to you and not mandated by the storage decide). It doesn't matter that the NAND controller probably used a few megabytes of the overprovisioned area to store its block mapping tables, or maybe even duplicated your data for its convenience - you were never charged for that overprovisioned area.
Here, you are sold a storage device (that you access over HTTP instead of SATA/PCIe), but when you write a 1GB file, they duplicate/convert/etc it for their convenience yet still charge you to store those duplicates you haven't asked for. That's new and unexpected.
For you (and me) perhaps, but not for most people[1] who appreciate having the file system abstracted away.
[1]: https://news.slashdot.org/story/21/09/27/2032200/students-do...
So when you’re syncing photos to iCloud, it’s not just the individual files that get synced but it’s the “Photos Library” managed container of the Photos app.
If you add individual files directly in Finder or the Files app then their size matches exactly both in iCloud and on the local file system.
Ultimately you’re increasingly tethered to some service for your storage that you pay for periodically based on total storage yet you have little-to-no information how to best optimize that storage if you want to operate in a fixed cost bracket or lower storage/cost ratio. So as a consumer, do I just wave my hands and keep throwing more and more money at the problem, especially now that devices are increasingly pushing everything, including storage, as a subscription service to meet my actual functional needs (that realistically could be met by local storage options if manufacturers didn’t have a vested interest in pushing me towards service based storage solutions)?
The modern business strategy in technology is simply hiding behind complexity. The cost is too complex for you to understand, it gives too much information away about our internals to competitors, and so on. Yet somehow these metrics are derived to assure the business is operating above cost because when the rubber meets the road it must be done, yet when the consumer wants to understand it’s suddenly too complex. The problem is that tech in many cases is growing to scales that really is too complex and business managers know this, so it’s often a valid excuse to hide behind. Conveniently that’s where they focus on investment and padding margins though.
Yes.
I can go and buy 1TB of Microsoft OneDrive or 2TB of Google Drive for less than a Franklin a year, and most people won't even need 1TB let alone 2TB. Both Microsoft and Google also offer 100GB plans for a Jackson a year, which is what I purchase myself. The average person can get by paying a Washington per month to Apple for 50GB of iCloud.
The amount of money I would save from managing photos myself locally isn't worth the time spent nor the money spent on the hardware.
EDIT:
For the downvoters, consider this: If I were to manage all this myself, I would need at least three storage mediums with one being a different form factor to satisfy the 3-2-1 backup scheme. I would also need to procure arrangements for that third backup copy in the 3-2-1 scheme. And I would need to spend time managing it all.
That is going to cost me more than a Franklin per year. Life is short, my time is precious, and my money is ultimately expendable.
Particularly the older I get the more I value my finite free time. Throwing $20 at something to remove a problem that would take me hours (not to mention a large startup cost) to do myself is just an obvious choice.
Remember it's not just buying the equipment, it's maintaining and understand it as well (e.g. I have to be familar with how ZFS works, how to restore a failed node, write some scripts, etc.). And with every backup solution you also need to be familar with the restoration process and test it occassionally to make sure it actually works as expected.
Wouldn't recommend it to Joe Normal. I've spent a good 5-6 days just rebuilding parity while switching disks. During which the server makes a godawful noise and the performance is degraded because of constant disk load.
It's great to have 40TB+ of storage right next to me though, the performance is great and I can run every self-hostable service imaginable with it.
But I still subscribe to Apple One for the whole family, it just works for every device (3-4 phones, 3 tablets, 2 laptops). I do have some backups running from those to the local system, but mostly it's for quick restores.
The only thing I'm adamant about is to keep your generated content for yourself. Don't trust FB, Instagram, Youtube or whatever to be the only storage of anything you create. Keep the master copy where you control it, publish it to other services.
It's a little absurd to think people don't need more than 2TB - especially on HN. Gamers will likely have 2TB in games alone, videographers often have many TBs of videos and photos from weddings and events in their life, many that care about health may have a few TB in genomic data mirrored on their computers to analyze, etc.
I would imagine it's hard to find people that wouldn't have TBs of data, if they were allowed to do so. The reason many people don't have TBs of data is they're limited by these exact companies you're claiming 'solve the problem' by offering limited storage.
It is notable however, that having better tools to organize, deduplicate, and compress data would be helpful to reduce some of the size of data that many people have. Over the years I've noticed my family will have multiple tar.gz archives, zip archives, etc, which (after extraction/unencryption) will share 20% files here, 10% files there, a 4kb jpg that's the same as a 100MB PNG here and there, etc. So yes, those 10TB archives may end up being 5TB if someone spent the time to really comb over, understand, make good decisions, and organize that data. But I have not yet seen anything that can scratch that surface yet, other than perhaps https://github.com/jjuliano/aifiles - but I won't use it until it's local only and has guarantees not to destroy data without explicit permission. An overlay filesystem that shows compression/deduplication with LLM capability like aifiles is probably the best option here.
However, I wouldn't imagine that most people's life data is less than 2TB even with all of this - it's mostly imposed as an artificial constraint by these companies.
If I were to install all the games on my steam account it would be many hundreds of terabytes of total storage. In the end I have about a terabyte of games on my computer. And of that 0 bytes are in my cloud storage.
I've been an amateur photographer for over 15 years. I tend to curate the photos I keep, largely because I don't need 20+ pictures of the same scene. Its more of a burden to casually flip through my photos if the majority of them are near duplicates. In the end my total collection is only several hundred gigs.
Most people aren't videographers.
Most people in my family have far less than even 50 gigs of actual data they care about. They maybe take a dozen compressed photos a week, maybe 30 minutes of videos a month. A lot of my friends take even fewer photos and pictures.
But in the end, that's not the same as my cloud storage where I'm being metered by my bytes.
> That is going to cost me more than a Franklin per year. Life is short, my time is precious, and my money is ultimately expendable.
When you're looking at cloud services, you need to perform your own off-site backup. Apple, Google, Microsoft, etc. will maintain copies that they'll restore in the event of a hardware failure. But, if your account gets compromised or a buggy sync or bad API event happens, your data is gone. They're not going to go restore it from tape for you. This is a big part of why I do have an in-home NAS. Maybe you have everything sync'd with a laptop and that has you covered, but Apple's expanded storage options are outlandishly expensive so I doubt many with the 2TB+ plans are able to do that. (Yes, you could use external storage, but that's also rather inconvenient for a Photos.app library.)
We could both get what we want if these storage operations weren't wrapped up in proprietary APIs. If I use iCloud I get a seamless experience on macOS, but no access at all on Linux. If I use Dropbox I get access on Linux, but little more than photo sync on an iPhone. Given the decades of precedent with filesystems and I/O APIs, I suspect we could have an abstraction layer and an implementation layer that would allow for interoperability. Anyone that wants to pay for iCloud are free to do so, others could use their preferred storage engine. But, allowing access into the walled garden is far less profitable.
For most people, storage needs are going to increase over time (more + higher resolution photos & videos, larger apps, document storage, etc.). 6TB for a family is not unreasonable and that's what? Three Franklins and three Jacksons per year + whatever for an external drive for your offsite backups. What comes after the 6TB option? Storage costs have decreased drastically over time, unless you're using a proprietary service; consumers are not benefiting at all from those gains in efficiency.
Well, 12TB.
And if you are head of household and share storage, you can combine storage plans. Mine currently shows “2.3TB of 14TB used”.
I know you're just following the theme set by the parent commenter, but there are a bunch of us folk on HN who aren't US residents, and have no idea how much those presidents mean in terms of currency.
Microsoft in particular has a 6TB for $100/year family plan, sharable with up to 5 other family members for a total of 6 persons each with 1TB. Google's plans can all also be shared with up to 5 other family members, though their bytes-per-dollar can't compete with that particular Microsoft family plan.
Basically: Local storage with personal management needs to be very easy, cheap, and carefree (which it isn't) to compete practically with cloud storage.
The only exception is if one's needs are niche and specific. I actually have a Synology NAS at home that I keep most of my data on, but that's because my data is mostly "bottle of rum" and "Linux ISO" in nature and thus not something I can throw on cloud storage in the first place.
- $100,000: Wilson
- $1,000: Cleveland
- $500: McKinley
- $100: Franklin*
- $50: Grant
- $20: Jackson
- $10: Hamilton*
- $5: Lincoln
- $2: Jefferson
- $1: Washington
* not a president
> I can go and buy 1TB of Microsoft OneDrive or 2TB of Google Drive for less than $100 a year, and most people won't even need 1TB let alone 2TB. Both Microsoft and Google also offer 100GB plans for a $20 a year, which is what I purchase myself. The average person can get by paying $1 per month to Apple for 50GB of iCloud.While we're at it, iCloud+ offers these monthly storage plans now:
United States:
50GB: $1
200GB: $3
2TB: $10
6TB: $30
12TB: $60
See everywhere in the world here:I imagine the other reason is because they’re not mutually exclusive: For instance, Synology makes it easy to have both an in-home NAS and cloud sync.
,End Balance after x years, 4% per year
10, 1,578.36, 63.13
15, 2,788.81, 111.55
30, 10,207.30, 408.29
40, 21,460.96, 858.44
I converted this into a TamperMonkey/Greasemonkey script. Also added a feature to "hide" all elements that do not match the threshold.
I personally use immich[1], a very complete solution with iOS / Android App, Server-Component and Sync / Backup option.
[1]: https://immich.app/
To me it’s better than any other interface I’ve tried. Commercial or home lab.
The features I appreciate the most are:
- Auto-Synchronisation on Android AND iOS (this is hard to find in any other app)
- Photo sharing (Accounts can have partners to share all their photos with - ideal for me and my wife)
- Deduplication
- The Web Interface
The feature I miss the most is tags[1].Once the app stopped working but there was a clear message on the repository / homepage that the server has to be upgraded. Since it is docker based, it was very easy to upgrade without losing any data. Same applies for backup...
So if you ask me, there is no need to worry - but I would not use it as only option to store my photos and it does not replace an (off-site)-backup.
[1]: https://github.com/immich-app/immich/discussions/1651#discus...
If you open iPhone settings and browse to:
> Apple ID > iCloud > Photos
There is an option to 'Optimise iPhone Storage' which is enabled by default. This states: > If your phone is low on space, full resolution photos and videos are automatically replaced with smaller device sized versions. Full-Resolution versions can be downloaded from iCloud at any time.
This seems perfectly reasonable to me.Encoding additional videos takes a lot of processing power. When you upload a video to youtube, the higher-quality versions aren't available for sometime. There is likely a queuing system, maybe even a minimum weight time before a video is processed.
This deals with the storage on the phone, which you can find in Settings > Manage storage. It has nothing to do with the storage in the cloud.
If you shoot RAW+JPEG (not a super-rare thing to do, for photo enthusiasts) then Apple Photos links the two images. Which is useful, rather than having a bunch of kinda-duplicates littering your library you can easily toggle between RAW and JPEG.
But this combining, along with the file system design described in this article, makes it impossible (as far as I can tell, anyway) to easily separate them and delete the RAWs. So years later, I have HUGE RAW files that I'll never touch that I can't delete, because I want to keep the much smaller JPEGS.
Any method that I've found to clean them up (exporting the originals, deleting them from the library, and then re-importing the JPEGs only seems easiest) will lose all of the years of metadata that I've built up in the library.
So I have to upgrade.
But in RAW/JPG world, the metadata and edits is already a solved problem with Sidecar files + EXIF data. Sure, EXIF fields are kinda messy but I'm sure it's better than Apple has rolled by their own.
> Any method that I've found to clean them up (exporting the originals, deleting them from the library, and then re-importing the JPEGs only seems easiest) will lose all of the years of metadata that I've built up in the library.
Apparently when you File/Export Unmodified Originals it will export the RAW+HEIC and a separate sidecar file containing the metadata. You can then move the RAW file away and import the HEIC file, which will autoimport the sidecar metadata file too.
You lose edits though, although it seems you can "copy edits" somehow. Surely a technically inclined person can AppleScript their way through this...
Yet it seems needlessly cumbersome and should be a built-in function in Photos.app, it's clearly not prioritized because it helps funnel people into higher iCloud tiers.
Which would be annoying by itself. But then iCloud has an arbitrary limit of 2TB as well, which does not seem to be able to be increased.
The open source tool osxphotos (https://github.com/RhetTbull/osxphotos) can help with this. You can export the JPEG images while preserving metadata using the thrid-party exiftool utility:
`osxphotos export /path/to/export --has-raw --skip-raw --exiftool`
This exports all images that have a raw pair but skips the raw component then uses exiftool (https://exiftool.org/) to write the metadata (keywords, etc.) to the exported JPEG files. You can then re-import these into photos either by dragging them or by running `osxphotos import /path/to/export/*`
Both the export and import commands have many other options for controlling export directory, etc. `osxphotos help export` or `osxphotos docs` to open docs in browser. (Disclaimer: I'm the author of osxphotos)
The issue with the recent changes with Apple is that they increase the prices for no good reason. We're always going to take photos/videos and their sizes keep increasing with modern tech and capabilities.
Often the photos app doesn't even detect the iphone even when plugged in, and it's a serious bug that Apple has neglected for years.
If you ever try a large-scale import/export into macOS Photos be prepared for 100% CPU, endless spinner cursors and the process ending up killed because it ran out of memory.
It's works the same way as Google Takeout. To get to it do this:
Sign in to your Apple ID account page at appleid.apple.com on a Mac, iPhone, iPad or PC. Go to “Data & Privacy” and select “Manage Your Data and Privacy.” On the following page, go to “Get a copy of your data” and select “Get started.”
Let me revolutionize your life:
https://github.com/icloud-photos-downloader/icloud_photos_do...
Edit: Seems like it has gotten support since I’ve last looked: https://github.com/icloud-photos-downloader/icloud_photos_do...
I use PhotoSync[0] to copy photos from my iPhone to a NAS. It's an excellent program.
It will even download photos from iCloud as needed and can do format conversions. I run it every few days to push new photos to my NAS so that I've always got a local copy (these also get backed up to Backblaze B2 nightly). The format conversion lets me keep HEIC+JPG pairs of my photos, so I have the original and something that's more readily usable.
What I really want is something that will do the same thing, but with iCloud Drive. I keep a bunch of stuff in there and it bothers me that I don't have a reasonable way to back it up. Apple's recommended methods[1] leave a lot to be desired.
In today’s age where storage is a commodity, this ought to be priced per GB used.
I can only swap that out for a 200GB increment. This is weird because I'm sure you used to be able to do that.
So I jumped 5, 50, 200, 2000.
>After you subscribe to Apple One, you can buy more iCloud storage if you need more. With both Apple One and an iCloud+ plan, you can have up to 14TB of total iCloud storage.
If this is widespread, this could be seen as apple bloating figures to push people to upgrade, which could lead to a lawsuit, no?
IANAL
It's similar levels of reliability to an FTP account mounted with curlftpfs[1] - except the latter at least fails in understandable ways and can be debugged.
No, the OP doesn’t take into account that adding media to Photos is not a simple “copy and store a file”. The Photos app (like all apps) has its custom Photos Library file format. So when one ads a picture or a video, the Photos app analyses it and stores all kinds of metadata that’s needed for the Photos app to work, including edit history and other bits. This is what eventually gets synced to iCloud.
The smart thing about having both an iPhone and a Macbook should potentially be; snap a bunch of photos, instantly have them available on the computer - but no. Apple apparently chooses a random time depending on 100 factors to upload the photos in the next 30 minutes to 7 days.
So you often have to airdrop a bunch of photos files completely invalidating the purpose of the sync function.
iClouds bad pricing and storage issues led me to switch to Google Photos, It's way better. The lack of native support on iOS can be a little cumbersome though.
The only thing I wish these cloud providers offered was a way to deduplicate photos / videos. That would make my life so much easier.
Same here ... but of course that would make it easier for customers to not have to pay more, so sadly don't think it's ever happening.
https://apps.apple.com/us/app/photosweeper/id463362050?mt=12
It was especially painful on a Windows PC.
The biggest selling point of other mp3 players for me was the ability to transparently copy and paste files into the filesystem.
I’m an Apple fan, but still don’t like the idea of having all my digital life on iCloud only.
You need to re-authenticate every 90 days due to 2fa restrictions but it works great.
Free 5GB is not enough for backing up iPhone system data any more.
Most insane is that if you remove some apps from backup, then it iCloud usage goes from 4.5GB to 4.6GB.
Also, (automated) version history might have caused this! Maybe (automatic) image optimiztion saved old and new versions of everything.
Unfortunately I haven’t found a way to automatically delete all the videos/photos that I send which still have the original in the photo stream. It would be awesome to be able to automate that.
https://www.igeeksblog.com/how-to-use-smart-search-filters-i... https://www.igeeksblog.com/delete-multiple-imessage-photos-a...
No idea what this means so can't tell you whether it works for me
I chuckled!
I believe the size difference has to do with the encoding, can’t tell for a fact since you didn’t show that part, but maybe it gets re-encoded when you download it hence the difference.
You don’t say. I’m sure this is just a harmless bug and not the source of $20M extra revenue per year across their hundreds of millions of active iCloud users.
Finds duplicates, screenshots and big vids. Had a random vid that was 1.8gb for reasons unknown
Doesn't always seem to be a accessible though
Pretty sure that can enumerate and query things like the filesize of the photos and videos in your iCloud.
Other companies would build the redundancy costs into the final pricing, but Apple is known to "think different".
1. older videos have a larger file size because they are encoded less efficiently (h264 vs h265)
2. gigabytes vs gibibytes (like how a 500GB drive formats to ~465 GiB)
3. rounding errors
4. Apple fucking its customers
5. combo of the above
I have a very similar issue with iMessage in iCloud, and unlike photos, there’s no web interface. You have to interact with it on your device. Thankfully I have a Mac so I can load up the chat database from there and see what’s using the space, but cleaning it up has been a nightmare: I have a script to find the big attachments, but if I delete them then some agent that manages the database goes into a loop for like five minutes per deletion. Then the change gets synced to iCloud (at least, I assume). So unless I fully reverse what the deletion process does and whether it is possible to do a batch operation I’m basically at the mercy of the front end they provide to clean things up, which as I mentioned earlier is absolutely not designed to make it easy to do this.
It's unlikely to be a "static" bug affecting everyone because it would've been caught in QA and/or user reports and fixed already (this is not a new issue, been happening for a year at least).
Therefore my hypothesis is that it's dependent on corrupt persistent data that a relative minority of users have which is not easy for Apple to detect/replicate.
Since the macOS Messages app is some Catalyst abomination, it won’t even support expected keyboard shortcuts for navigation, selection and deletion. Having to do everything using the trackpad is a slow process.
It’s frustrating that messages is so bad.
Even in open source apps where in theory we can contribute functionality like Signal it’s difficult. E.g. it’s impractical to delete lots of media in groups in the iOS app. I added a PR last year to help this but still waiting for a merge even though it has been approved :(
Edit: PR for the curious
And this is why they (and many companies) behave this way, there's no incentive to fix a problem that you will just pay them more money to fix. Fixing the bug means reduced revenue. Complaining about the problem sadly won't help, you can only stop using the app and stop buying their products.
The user’s options are quit the ecosystem or pay up. Apple’s lock-in monopoly means they get to shove their customers into another monthly fee.
This is a classic example of the harm done to consumers by allowing anticompetitive practices. (The consumer has no choice but to pay)
When the US government gets off their butts and investigates Apple for antitrust, this will be one of the key findings. The fine should be some multiple of the amount of money they have illegally extracted from their customers.
Same with HomeKit Secure Video (HKSV). The only way to access recordings is through the Home app. I'm surprised nobody has figured out a way to browse them outside of the app.
And such devices will be widely available, also in non-first world countries, with proper OEM warranties and support. Hell, a local manufacturer can just build for that spec.
I know it’s like a wishful dream. But if this happens, and when this happens, I hope a lot of us will be able to breathe free and hopefully would be candidly able to shit upon legacies of those so called fucking visionaries, who were barely not subhuman and were just rather pathetic jerks, who ensured such shit-show of walled gardens and opaquely implemented utterly inferior systems where users go and get stuck in one of the two houses of the duopoly to experience a lot of shitty things including, but not limited to, some variation of Stockholm syndrome and clear and conscious apologism.
Why? Because neither of the two is acceptably good and they have become so big and they have closed it down so much that nobody else can even make a dent even if they try. And they try!
So yeah, until then I will rant and feel shitty about both my iPhone 14 (as they call it - my “daily driver”) and Pixel 5a (my bread and butter phone; form factor wise less shitty one though).