A disk so full, it couldn't be restored
sixcolors.com
sixcolors.com
Use an external storage device as a Mac startup disk https://support.apple.com/en-us/111336
Was surprised to learn that with Apple silicon-based Macs, not all ports are equal when it comes to external booting:
If you're using a Mac computer with Apple silicon, your Mac has one or more USB or Thunderbolt ports that have a type USB-C connector. While you're installing macOS on your storage device, it matters which of these ports you use. After installation is complete, you can connect your storage device to any of them.
* Mac laptop computer: Use any USB-C port except the leftmost USB-C port when facing the ports on the left side of the Mac.
* iMac: Use any USB-C port except the rightmost USB-C port when facing the back of the Mac.
* Mac mini: Use any USB-C port except the leftmost USB-C port when facing the back of the Mac.
* Mac Studio: Use any USB-C port except the rightmost USB-C port when facing the back of the Mac.
* Mac Pro with desktop enclosure: Use any USB-C port except the one on the top of the Mac that is farthest from the power button.
* Mac Pro with rack enclosure: Use any USB-C port except the one on the front of the Mac that's closest to the power button.
Hilariously this failure case doesn't seem to be listed in the docs. https://developer.apple.com/library/archive/documentation/Sy...
i find it really frustrating though. why not just reserve some space?
It reserves a percent of your pool's total space precisely to avoid having 0 actual free space and only allows using space from that amount if the operation is a net gain on free space.
https://github.com/openzfs/zfs/blob/99741bde59d1d1df0963009b...
This is a brokwn implementation.
They also said it was mainly used for other issues, such as fragmentation. In other words, this was stated as a fix for the file delete issue.
How does this invalidate my comment, that this was a broken implementation?
It doesn't matter if it will be fixed in the future, or was just fixed.
https://btrfs.readthedocs.io/en/latest/btrfs-filesystem.html
> GlobalReserve is an artificial and internal emergency space. It is used e.g. when the filesystem is full. Its total size is dynamic based on the filesystem size, usually not larger than 512MiB, used may fluctuate.
With that in mind, you can see how we get in a scenario where deleting a file will require a minor bit of storage for recordkeeping the old and new states, before it can actually free up the storage by releasing the old state. There is supposed to be an escape hatch for getting yourself out of a situation where there isn't even enough storage for this little bit of record keeping, but either the author didn't know whatever trick is needed or the filesystem code wasn't well-behaved in this area (it's a corner-case that isn't often tested).
1. Just remove some files - ZFS will attempt to do the right thing
2. Remove old snapshots
3. Mount the drive from another system (so nothing tries writing to it), then remove some files, reboot back to normal
4. Use `zfs send` to copy the data you want to keep to another bigger drive temporarily, then either prune the data or if you already filtered out any old snapshots, zero the original pool and reload it by `zfs send` from before.
You can have cheap defrag but comparatively brittle filesystems by making things modifiable in place.
You can have filesystem that has as its primary value "never lose your data", but in exchange defragmentation is expensive.
With snapshotting, especially with filesystems that can only write data through snapshots (like ZFS), blocks can be referred to by many pointers.
It's similar to evaluating liveness of object in a GC, except you're now operating on possibly gigantic heap with very... pointer-ful objects, that you have to rewrite - which goes against core principle of ZFS which is data safety. You're doing essentially a huge history rewrite on something like git repo, with billions of small objects, and doing it safely means you have to rewrite every metadata block that in any way refers to given data block - and rewrite every metadata block pointing to those metadata blocks.
In fact, the main difficulty with garbage collectors is maintaining real-time performance. Throw that constraint out, and the game changes entirely.
You can attempt to add an extra indirection layer, but it does not really reduce fragmentation, it just lets you remap existing blocks to another location at a cost of extra lookup. This is in fact implemented in ZFS as solution for erroneous addition of a vdev, allowing device removal though due to performance cost its oriented mostly at "oops, I added the device wrongly, let me quickly revert".
You're missing the part where (c) is forbidden by design of the filesystem, because ZFS is not just "Copy on Write" by default (like BTRFS, which has in-place rewrite option, IIRC) nor LVM/disk-mapper snapshot which similarly don't have strong invariants on CoW.
ZFS writes data to disk in two ways - a (logically) write-ahead log called ZFS Intent Log (which handles synchronous writes and is read only on pool import), and transaction group sync (txgsync), where all newly written data is linked into new metadata tree, sharing structure with previous TXG metadata tree (so unchanged branches are shared), and the pointer to the head of the tree is committed into on-disk circular buffer of at least 128 pointers.
Every snapshot in ZFS is essentially a pointer to such metadata tree - all writes in ZFS are done by creating a new snapshot. The named snapshots are just rooted in different places in filesystem. This means that sometimes even in case of catastrophic software bug (for example, master branch had for few commits a bug where they accidentally changed on-disk layout of some structures - one person ran master branch and hit that resulting in pool that could not be imported... but the design meant they could tell ZFS import to "rewind" to TXG sync number from before the bug)
Updating the blocks in place violates design invariants - once you violate them, the data safety guarantees are no longer guarantees. And this makes it into minimally offline operation, and at that point the type of client that needs in-place defragmentation can reasonably do the two-space trick (if you're big enough, to make that infeasible, you're probably big enough to easily throw in an extra JBOD at least and relieve fragmentation pressure).
To make latter paragraphs understandable (beware, ZFS internals as I remember them):
ZFS is constructed of multiple layers[1] - from the bottom (somewhat simplified):
1. SPA (Storage Pool Allocator) - what implements "vdevs" - the only layer that actually deals with blocks. It implements access to block devices, mirroring, RAIDz, draid, etc. and exposes single block-oriented interface upwards
2. DMU (Data Management Unit) - An object oriented storage system. Turns bunch of blocks into object-oriented PUT/GET/PATCH/DELETE like setup, with 128bit object IDs. Also handles base metadata - the immutable/write-once trees for turning "here's a 1GB blob of data" into 512b to 1MB portions on disk. For every given metadata tree/snapshot, there is no in-place changes - modifying an object "in place" means that new txgsync has, for given object ID, a new tree of blocks that shares as much structure with previous one as possible.
3. DSL / ZIL / ZAP - provide basic structures on top of the DMU - DSL is what gives you "naming" ability for datasets and snapshots, ZIL handles the write-ahead log for dsync/fsync, ZAP provides a key-value store in DMU objects.
4. ZPL / ZVOL / Lustre / etc - Those are the parts that implement user-visible filesystem. ZPL is ZFS Posix Layer, which is a POSIX-compatible filesystem implemented over object storage. ZVOL does similar but presents emulated block device. Lustre-on-ZFS similarly talks directly to ZFS object layer instead of implementing ODT/OST on top of POSIX files again.
You could, in theory, add an extra indirection layer just for defragmentation, but this in turn makes problematic layering violation (something found at Sun when they tried to implement BPR) - because suddenly SPA layer (the layer that actually handles block-level addressing) needs to understand DMU's internals (or a layer between the two needing bi-directional knowledge). This makes for possibly brittle code, so again - possible but against overarching goals of the project.
The "vdev removal indirection" works because it doesn't really care about location - it allocates space from other vdevs and just ensures that all SPA addresses that have ID of the removed vdev, point to data allocated on other vdevs. It doesn't need to know how the SPA addresses are used by DMU objects
> Updating the blocks in place violates design invariants - once you violate them, the data safety guarantees are no longer guarantees.
Again - you can copy blocks prior to deleting anything, and commit them atomically, without losing safety. The fact that you (or ZFS) don't wish to do that doesn't mean it's somehow impossible.
> the type of client that needs in-place defragmentation can reasonably do the two-space trick (if you're big enough, to make that infeasible, you're probably big enough to easily throw in an extra JBOD at least and relieve fragmentation pressure).
You're moving goalposts drastically here. It's quite a leap to go from "has a bit of free space on each drive" to "can throw in more disks at whim", and the discussion wasn't about "only for these types of clients".
And, in any case, this is all pretty irrelevant to whether ZFS could support defragmentation.
> this makes it into minimally offline operation
See, that's your underlying assumption that you never stated. You want defragmentation to happen fully online, while the volume is still in use. What you're really trying to argue is "fully online defragmentation is prohibitive for ZFS", but you instead made the wide-sweeping claim that "defragmentation is prohibitive for snapshotted filesystems in general".
I did say that there are trade offs and that some goals can make things like defragmentation expensive.
ZFS' main design was that it nothing short of (extensive) physical damage should allow destruction of users data. Everything else was secondary. As such, the project was not interested, ever, in supporting in-place updates.
You can design a system with other goals, or ones that are more flexible. But I'd argue that's why BTRFS got undying reputation for data loss - they were more flexible, and that unfortunately also opened way for more data loss bugs.
That's not true. That was only in the beginning -- "impossible" was only what I originally took (and would still take, but I digress) your initial comment of "ability to defragment is not free" to be saying. It's literally saying that if you don't pay a cost (presumably, performance or reliability), then you become unable to defragment. That sounded like impossibility, hence the initial discussion.
Later you said you actually meant it'd be "prohibitively expensive". Which is fine, but then I argued against that too. So now I'm arguing against 2 things: impossibility and prohibitive-expensiveness, neither of which I'm hung up on.
> ZFS' main design was that it nothing short of (extensive) physical damage should allow destruction of users data. Everything else was secondary.
Tongue only halfway in cheek, but why do you keep referring to ZFS like it's GodFS? The discussion was about "filesystems" but you keep moving the goalposts to "ZFS". Somehow it appears you feel that if ZFS couldn't achieve something then nothing else possibly could?
Analogy: imagine if you'd claimed "button interfaces are prohibitively expensive for electric cars", I had objected to that assertion, and then you kept presenting "but Tesla switched to touchscreens because they turned out cheaper!" as evidence. That's how this conversation feels. Just because Tesla/ZFS has issues with something that doesn't mean it's somehow inherently prohibitive.
> As such, the project was not interested, ever, in supporting in-place updates.
Again: are we talking online-only, or are you allowing offline defrag? You keep avoiding making your assumptions explicit.
If you mean offline: it's completely irrelevant what the project is interested in doing. By analogy, Microsoft was not interested, ever, in allowing NTFS partitions to be moved or split or merged either, yet third-party vendors have supported those operations just fine. And on the same filesystem too, not merely a similar one!
If you mean online: you'd probably be some intrinsic trade-off eventually, but I'm skeptical it's at this particular juncture. Just because ZFS may have made something infeasible with its current implementation, that doesn't mean another implementation couldn't have... done an even better job? e.g., even with the current on-disk structure of ZFS (let alone a better one), even if a defragmentation-supporting implementation might not achieve 100% throughput while a defragmentation is ongoing, surely it could at least get some throughput during a defrag so that it doesn't need to go entirely offline? That would be a strict improvement over the current situation.
> But I'd argue that's why BTRFS got undying reputation for data loss - they were more flexible, and that unfortunately also opened way for more data loss bugs.
Hang on... a bug in the implementation is a whole different beast. We were discussing design features. Implementation bugs are... not in that picture. I'm pretty sure most people reading your earlier comments would get the impression that by "brittleness" you were referring to accidents like I/O failures & user error, not bugs in the implementation!
Finally... you might enjoy [1]. ;)
[1] https://www.reddit.com/r/zfs/comments/1826lgs/psa_its_not_bl...
Rinse and repeat.
the trick is to truncate a large enough files, or enough small files, to zero.
not sure if this is a universal shell trick, but worked on those i tried: "> filename"
Since then I memorized this: `cat /dev/null >! filename`, and it has worked on systems with zsh and bash.
I believe "> filename" only works correctly if you're root (at least in my experience, if I remember correctly).
EDIT: To remove <> from filename placeholder which might be confusing, and to put commands in quotes.
It saved me just yesterday when I needed to truncate hundreds of gigabytes of Docker logs on a system that had been having some issues for a while but I didn't want to recreate containers.
"truncate -s 0 /var/lib/docker/containers/**/*-json.log"
Will truncate all of the json logs for all of the containers on the host to 0 bytes.
Of course the system should have had logging configured better (rotation, limits, remote log) in the first place, but it isn't my system.
EDIT: Missing double-star.*
openat(AT_FDCWD, "file", O_WRONLY|O_CREAT|O_TRUNC, 0666) = 3
man 2 openat: O_TRUNC
If the file already exists and is a regular file and the
access mode allows writing (i.e., is O_RDWR or O_WRONLY) it
will be truncated to length 0.
... MULTIOS=1 > file
- zsh isn't POSIX compatible by defaultHowever, it won't work in bash. It will create file named "!" with the same contents as "filename". It is equivalent to "cat /dev/null filename > !". (Bash lets you put the redirection almost anywhere, including between one argument and another.)
---
[1] See https://zsh.sourceforge.io/Doc/Release/Redirection.html
In that case I'll just always use `truncate -s0` then. Safest option to remember without having to carry around context about which shell is running the script, it seems.
: is a shell built-in for most shells that does nothing.
https://support.apple.com/en-us/108900
That said, no idea why they can’t be used in this case
My intuitive guess here is how the ports are connected to the T2 security chip. One port is as you said a console port that allows access to perform commands to flash/recover/re-provision the T2 chip. Same as an OOB serial port on networking equip.
The rest of the ports the T2 chip has read/write access to devices connected to it. Since this is an OS drive, I'm guessing it needs to be encrypted and the T2 chip handles this function.
The rest of the code necessary to boot from external sources is located on main flash
This is like saying my software did not work because it was based on an incompatible version of some library. Maybe so, but that is a bad excuse. Implementing systems is hard, and like the rest of us, Apple should not get away with bad excuses. And this is even more true because they control more of the stack.
The 3 users of this feature on this planet are already happy that it’s even possible at all. The only thing Apple could do is to document this clearly like adding a text in the boot drive selector.
Also on my mbpro at least the mentioned port is the one closest to the magsafe connector and may have funny electrical connections to it, perhaps.
My bet is that you can get nearly the same functionality with single user mode vs booting from external media, but I only have a vague understanding of the limitations of all three modes from 3-5 uses via tutorials.
iirc, not all ports were equal when it came to charging with the m1 macs, so this is actually not so surprising.
macOS continued to write files until there was just 41K free on the drive.
I've (accidentally) ran both NTFS and FAT32 to 0 bytes free, and it was always possible to delete something even in that situation.
Digging around in forums, I found that Sonoma has broken the SMB/Samba-based networking mount procedure for Time Machine restores, and no one had found a solution. This appears to still be the case in 14.4.
In my experience SMB became unreliable and just unacceptably buggy many years ago, starting around the 10.12-10.13 timeframe; and now it looks like Apple doesn't care about whether it works at all anymore.
I hate to think what people without decades of Mac experience do when confronted with systemic, cascading failures like this when I felt helpless despite what I thought I knew and all the answers I searched for and found on forums.
I don't have "decades of Mac experience", but the first thing I'd try is a fsck --- odd not to see that mentioned here.
If I were asked to recover from this situation, and couldn't just copy the necessary contents of the disk to another one before formatting it and then copying back, I'd get the APFS documentation (https://developer.apple.com/support/downloads/Apple-File-Sys...) and figure out what to edit (with dd and a hex editor) to get some free space.
Apple dropped Samba in favor of their own implementation a long time ago after Samba adopted GPLv3:
https://lists.samba.org/archive/samba-announce/2007/000122.h...
https://www.engadget.com/2011-03-24-apple-to-drop-samba-netw...
I'd like to see zsh also adopt GPLv3 to call Apple's bluff.
https://github.com/fish-shell/fish-shell/blob/master/COPYING
https://git.kernel.org/pub/scm/utils/dash/dash.git/tree/COPY... (Default Ubuntu shell since 6.10)
I'm less familiar with fish, but based on a very fuzzy awareness, it's at least fully-featured.
I do encounter dash on a few systems. It's the default shell on my OpenWRT networking kit, for example. I've installed bash where those systems have enough storage to accommodate it.
Changes are usually batched to reduce the amount of tree changes to a manageable amount. A bonus of this design is that a filesystem snapshot is just another reference to a particular tree.
This requires space, but CoW filesystems also usually reserve an amount of emergency storage for this reason.
s/unusual/usual/ surely.
Tested a few years ago throughput to a big NAS connected in 10gigEo from a Hackintosh with BlackMagic Disk Speed Test :
* running Windows, SMB achieves 900MB/s
* running MacOS, SMB achieves 200MB/s
* running MacOS, NFS and AFP both achieve 1000MB/s
Anything related to professional work is a sad joke in MacOS, alas.
(People keep repeating that AFP is dead, however it still works fine as a client on my Mac Pro -- and performs so much better than SMB than it's almost comical).
> ran both NTFS and FAT32 to 0b and was able to delete something.
AFAIK those aren’t journaled, no?
However sometimes filesystems can't do that. For those cases, hopefully the filesystem supports: resize-grow, resize-shrink, and either additional temporary storage or is on top of an underlying system which can add/remove backing storage. You may also need to use custom commands to restore the filesystem's structure to one intended for a single block device (btrfs comes to mind here).
/dev/null is magical and worth reading into
I knew that ZFS was better about this, but even so I still got that "oh... hell" sinking feeling when you really bork something.
Because they dont sell Time Capsule anymore. And they want you to backup everything to iCloud to grow their Services Revenue.
But we shall see!
Hopefully they will provide the ability to backup and restore file versions!
But iCloud isn’t a backup.
Its sync.
And it will happily sync corrupt files and does not provide any versioning.
The best TimeMachine is an SSD connected locally to the Mac.
The second best is an SSD running on a Mac setup with TomeMachine Server.
Then you are lucky if backups continue to work. And even luckier if you can sensibly restore anything via the hellscape that is the interstellar wormhole travel interface! ;-)
BackBlaze is very reliable at least! Not cheap with a house full of computers to backup though :-/
It does not even provide a basic progress bar when used on the phone.
My current desktop Mac environment is a direct descendant of my original Mac from 2004 thanks, largely, to Time Machine.
But hey, we get new emojis and moving desktop wallpapers…
Both machines on wired gigabit ethernet, yet the restore took more than 24 hours. And that was for just a 1TB disk.
So by adjusting the Linux kernel tunable "spa_slop_shift" to shrink the slop space, you can regain up to 128GB of bonus space to successfully complete your file deletion operations:
https://openzfs.github.io/openzfs-docs/Performance%20and%20T...
As does ext4 (although they call the space "reserved blocks"). 'man tune2fs' for details. As well as most other modern (and not so modern[0]) filesystems.
[0] As I recall, the same was true for SunOS'[1] UFS back in the 1980s.
I believe this problem in is only relevant to CoW filesystems. With ext[234]fs you can set the reserved blocks to 0, fill the fs, and always remove files to fix the situation.
It's kinda like how almost any 1980's MS-DOS shareware terminal program was really good at downloading files over a limited-bandwidth connection, but current versions of MS Windows are utter crap at that should-be-trivial task.
In a different field: In the early decade of Wikipedia it often had to be explained to people that (at least from roughly 2004 onwards) deleting pages with the intention of saving space on the Wikipedia servers actually did the opposite, since deletion added records to the underlying database.
Related situations:
* In Rahul Dhesi's ZOO archive file format, deleting an archive entry just sets a flag on the entry's header record. ZOO also did VMS-like file versioning, where adding a new version of a file to an archive did not overwrite the old one.
* Back in the days of MS/DR/PC-DOS and FAT, with (sometimes) add-on undeletion utilities installed, deleting a file would need more space to store a new entry into the database that held the restore information for the undeletion utility.
* Back in the days of MS/DR/PC-DOS and FAT, some of the old disc compression utilities compressed metadata as well, leading to (rare but possible) situations where metadata changes could affect compressibility and actually increase the (from the outside point of view) volume size.
"I delete XYZ in order to free space." is a pervasive concept, but it isn't strictly a correct one.
I was lucky: I had an additional APFS partition that I could remove, thus freeing up disk space. Took me a while to figure out, during which time I was in a proper panic.
---
https://apple.stackexchange.com/questions/338721/disk-full-t...
I’m in a pickle here. macOS Mojave, just updated the other day. I managed to fill my disk up while creating a .dmg, and the system froze. I rebooted. Kernel panic.
Boot to Recovery mode. Mount the disk. Open Terminal.
–bash–3.2# rm /path/to/large/file
rm: /path/to/large/file: No space left on device
Essentially the same issue as this Unix thread from ‘08! https://www.unix.com/linux/69889-unable-remove-file-using-rm...
I’ve tried echo x > /path/to/large/file, no good.
It’s borked. Does anyone have any suggestions that aren’t “wipe the drive and restore from your backup”?
Kind of like the old Unix file systems that would reserve 5% for root.
[0] https://www.cockroachlabs.com/docs/v23.2/cluster-setup-troub...
I am certain this was a result of APFS being copy-on-write and supporting snapshotting. If no change is immediately permanent, but instead old versions of files stay around in a snapshot, then if you don't have enough space for more snapshot metadata you're in trouble. Maybe they skip the snapshot in low disk space situations, but they still have the copy-on-write metadata problem.
In contrast, after accidentally maxing out the space on my windows 11 office laptop which has a single data and boot volume, I was still able to boot it and sort the issue out.
From the tiny beginning I started being able to delete bigger and bigger spaces until finally it was clear and then of course I resized the partition so that wouldn't happen again. The End.
And I feel like that ought to be the lesson for power users: always leave a bit of slack space after your partition.
echo > file # delete content of file first
rm file # now it should workTook us a little while to figure out that the problem was the database file was so fragmented NTFS couldn't store more fragments for the file[1].
What had happened was they had been running the database in a VM with very low disk space for a long time, several times actually running out of space, before increasing the virtual disk and resizing the partition to match. Hence all the now-available disk space.
Just copying the main database file and deleting the old solved it.
[1]: https://superuser.com/questions/1315108/ntfs-limitations-max...
She somehow managed to fill up the entire 512GB. Updates were unsuccessful, she couldn't make calls and wasn't able to delete anything to make room.
She couldn't even back up her phone through iTunes, the only option was to purchase an iCloud subscription and back up to the cloud in order to access her photos.
The solution was to boot into recovery and mount the Data partition using Disk Utility.
I don't recall where the Data partition gets mounted but I think it is:
"/System/Volumes/Macintosh HD - Data"
Or just Data, since Sonoma. It will be clear from Disk Utility.
Then close Disk Utility and go into Terminal and run rm on a big unseeded file.
You can find one using:
find <data-mnt> -size +100m
Using rm will fail.
Unmount the Data partition and run fsck on it.
This completes the deletion.
From there enough more space can be freed in recovery to have a healthy buffer, then reboot normally and finish cleaning.
It seems that when the volume gets so full that rm doesn't work anymore the filesystem also gets corrupted.
HTH and that I didn't forget anything.
In general I’ve had good success with Time Machine. I, too, have lost TM volumes. I just erased them and started again. Annoying to be sure but 99.99% of the time don’t need a years worth of backups.
The author mentioned copying the Time Machine drive. I have never been able to successfully do that. Last time I tried I quit after 3 days. As I understand it, only Finder can copy a Time Machine drive. Terrible experience.
That said, I’d rather cope with TM. It’s saved me more than it’s hurt me, and even an idiot like me can get it to work.
I did have my machine just complain about one of my partitions being irreparable, but it mounted read only so I was able to copy it, and am currently copying it back.
I don’t know if this is random bit rot, or if something is going wrong with the drive. That would be Bad, it’s a 3TB spinning drive. Backed up with BackBlaze (knock on wood), but I’d rather not have to go through the recovery process if I could avoid it.
Problem is I don’t know how to prevent it. It’s been suggested that SSDs are potentially less susceptible to bit rot, so maybe switching to one of those is a wise plan. But I don’t know.
rsync -av $SOURCE $DEST has never let me down. Copy or delete on Time Machine files using Finder never worked for me.
> Problem is I don’t know how to prevent it. It’s been suggested that SSDs are potentially less susceptible to bit rot, so maybe switching to one of those is a wise plan. But I don’t know.
OpenZFS with two drives should protect you from bit rot. ZFS almost became the Mac file system in Snow Leopard.
AFAIK it's fixed now, because btrfs reserves some space and reports "disk full" before it's reached. macOS probably does the same (I'd hope), but it seems in this case the boundary wasn't enforced properly and the background snapshot caused it to write beyond.
https://zfs-discuss.opensolaris.narkive.com/BQ7RMcjo/cannot-...
The new problem with that reserved pool mechanism is that in 2024 it's probably way too big, because it's essentially a small but fixed percentage of the storage size. Don't let people use thresholds of total size without some kind of absolute cap!
If the POSIX API does have some limitation which would prevent this error from occurring with higher level APIs (which I sincerely doubt), macOS should simply start failing with errno = ENOSPC earlier for POSIX operations.
There is no other system that behaves like this, and we wouldn't be making excuses like this if Microsoft messed something basic up like this.
I understand the logic, but typically I've seen filesystem implementations block writes once metadata volumes become close enough to full. Also, and I don't know if this is a thing on modern filesystems, you used to be able to reserve free space for root user only, precisely for recovering from issues like this in the past.
I also heard of this happening to regular users downloading stuff with safari. It is simply terrible design on apple's part that you can kill a macOS install simply by filling it up so much that it becomes possible to not be able to delete files.
All too familiar. I have two Macs. I upgraded one of them to Sonoma and ever since then it has been nothing but headache and disappointment. Starting from the upgrade having failed (meaning I had to completely wipe the disk and install Sonoma from scratch, luckily I still had data), to problems with Handoff, the firewall seems to not work, Excel very slow etc etc.
I don't recommend using Sonoma.
Bluetooth audio was a joke even before the update, and now it's almost unusable.
As a side node, Time Machine has been pretty garbage lately. I back up to a local Synology NAS and letting it run automatically will just spin with `connecting to backup disk` (or some such message), but running manually works just fine.
Luckily, all important files were in the cloud, and you could write a blog post describing these monumental failures.
In the end I was able to mount and rescue the data using https://github.com/libyal/libfsapfs
I followed this guide: https://matt.sh/apfs-object-map-free-recovery
One day she had discharged the battery completely, shutting down the phone; after recharging, she tried to restart it, only to be sent into a boot-loop. There is no (official) way to resolve this except repeatedly reboot and hope that at least once, Springboard loads and you can immediately jump into Photos and start mass-deleting, or at least connect to a computer and transfer the media out of the phone.
There was a time when OSs could deal with the system disk being full quite well. And not so long ago.
I'm always amazed at how the file system survives just fine, but the machine doesn't even crash!
I'm not sure where you got your experience from.
My experience is that Windows and many of its programs will become very unstable with 0 b on the system drive. And about 3 times out of maybe 50, the system also became unbootable. I’ve learned to do whatever I can to free up space before restarting for stability.
The last time I’d regularly run out of space on Win was around Windows 98 times. I never had a problem then. Now in Windows 11 times, it’s a real headache.
Not sure how you’re so lucky.
Still, whatever this APFS bug is, the conditions to trigger it are more specific than just filling up the disk.
I do not know how to fix it without reinstalling
Use Backblaze if you don’t care about privacy, rsync+ssh to a selfhosted zfs box if you do.
EDIT: I checked your tool. It's a 1000 bucks to restore 4 TB in 48 hours. If the house burns down, insurance will cover that. I guess now I know I gotta check those drives a bit more.
What? This tool is exceptionally out of date. Retrieval cost is $30/TB at the high end, and for glacier deep archive and a 48 hour window it only costs $2.50/TB. (Plus a few cents per thousand requests, so maybe don't use tiny objects.)
Glacier's percentage-rate-based retrieval pricing was only active from 2012-2016.
The bandwidth charge of $90/TB is still accurate. Though there are ways to reduce it.
Why use closed source crypto for money when free software that can be reviewed is available gratis? There are much better options.
It’s worse than this. The private key for data decryption is sent to their server by the installer before you can even set a PEK. Then, setting the PEK sends the password to them too, since that’s where your private key is stored. So you have to take their word not just that they never store the key and promptly delete unencrypted files during restoration, but also that they destroy the unprotected private key and password when you set up PEK. It’s a terrible scheme that seems almost deliberately designed to lull people into a false sense of security.
The average user just needs to be able to ask the question in a decent place. E.g., Hacker News or a fitting Stack Exchange site. Some developers not afraid of touching the kernel will see the question (or one like it), and if no workaround (e.g. truncation) is found to be acceptable, someone may decide to look into the kernel source to see if it's feasible at all. They may find a lower level function that deletes without writing metadata. Or they may find the function in the filesystem driver's source code where the metadata is written first, and if that was successful, the actual data is written. In the easiest case, you could create a copy of the function with the calls swapped, and a live CD with the modified driver could be created. (Course, this solution is quite unsafe, as writing the metadata could still fail for some other or related reason, so it's a bit of an emergency solution.)
There are two other filesystems that were mentioned in the discussion here, btrfs and ZFS.
ZFS solved the problem by reserving space, so creating such a tool isn't needed. (However, ZFS is not part of Linux, so I'm not too interested in digging into the details.)
btrfs users apparently accept this as a fact-of-life, but have what they consider decent-enough workarounds, see e.g. https://www.reddit.com/r/btrfs/comments/ibjrpm/can_i_somehow....
(I use neither ZFS nor btrfs; I prefer boring filesystems, thank you very much.)
For a filesystem where this happens, it would not be simple and it would require a lot of experience to get right.
> Or more likely, somebody else would have been there before you and you could just use their tool.
I don't think open-source makes such a tool much more likely to exist.
Pet peeve (or alternatively: correct me if I'm wrong), Samba is not a protocol. It's a software suite that implements the SMB (Server Message Block) protocol
:> somefile
Note the colon, which stands in for the command as a null operator.
An additional hint is that if the system is so starved you can’t run ls, again call On the shell for help and use “echo *”, which will show a list of file names (without sizes) from which hopefully you can select a large file (or few) based on name alone. Colon-greater them into zero byte status and watch your system become useable enough to begin normal recovery efforts.
From an old Unix admin; I hope that helps someone.
This looks like a huge bug, and the elephant in the room.
What's the reason for `rm` requiring space left on the device?
But I found myself in the exceedingly frustrating and confusing situation of not being able to free space by deleting files:
- 'rm <filename>' from bash
- Emptying Trash.
- Deleting ".Trashes" directories.
- Booting single-user and attempting to delete files.
- Booting to the Recovery system, invoking terminal, and attempting to delete files.
- Attempting to remove snapshots using the 'tmutil' utility, either booted normal, safe-mode, single-user, or rescue. Best I can tell, tmutil simply would not run under single-user or rescue modes.
Final solution was to repartition the hard drive, re-create filesystem, reinstall the OS, and recover user files from Time Machine. This ... took a while.
My take-away is that MacOS behaves exceedingly poorly under any number of high-resource-utilisation modes (memory, CPU, or disk usage).
Some discussion from the time on the Fediverse: <https://toot.cat/@dredmorbius/111849495055957910>
Clearly Apple knows about this, they make no effort to fix it?
diskutil disableJournal “/Volumes/Macintosh HD” (or whatever volume is)
https://support.apple.com/en-gb/guide/disk-utility/dskuf8235...
Otherwise, go to Recovery mode, mount the disk in Disk Utility, and then open Terminal and rm some shit.
$ > large-file
It's possible that rm / unlink require working space to perform the transaction of removing the file from the directory, while truncation does not.That isn't reasonable in the sense of "this is what the filesystem should do in this situation" but if the log and user data are allocated from the same pool it is quite possible to exhaust both.
If it was a journal issue would something akin to using an initramfs (or live environment) and mounting with data=writeback enable removing files? Or maybe APFS doesn't support that?
I showed him how he could type rm -rf and then paste the files in the terminal and it took him a good 15 minutes for him to get it all organized but once he did he literally cried because I got his computer back.
Those days were pretty brutal, but there were shining moments.
Suddenly queries started failing, so I investigated and even managed to remove some files on the affected volume, but it wasn't enough. Eventually, as the volume's storage class did not allow for expansion, I reached out to the support team to move the data to a larger one.
However they're doing it, their disk space calculations are either wrong or estimates.
I'm begging the question:
In the 70s / 80s there was consumer tort litigation for all kinds of "misrepresentation", which was fair, because there was a huge and growing business of dark patterns of false claims, faulty products, and schemes for exploitation of unwitting and/or oppressing customers.
But it also became absurd, like class action suits against RCA for selling TVs with 25" inch screens where the picture only measured 24.5 inches due to the cabinet facia overlapping the edge of the tube or the raster not reaching all the way to the edge, etc.
Tort reform became a hot-button political issue because an enormous subsector of "consumer-rights" legal practice developed to milk payouts under consumer law. You still see this today for "personal injury."
So Apple has trillions in pockets and all their kit is sold with capacity specs.
Well, customers has better be able to get access to all that capacity, even if it kills their device.
I'm wondering if it's not a bug but a legal calculus, a la intro to Fight Club where Ed Norton is reviewing the share value implications of Ford paying off claims for Pintos that explode when rear-ended versus the cost of a recall?
I had a recent backup so i just reinstalled everything. I always make sure there is some space on the disks from that moment on.
As somebody that has managed several times to lock myself out of my user session by this (and had to learn how to enter again by several means); couldn't they start deleting any files with less than this size and increasing slowly the space available one logfile at a time?.
Just wondering.
There are threads on Reddit about this.
For me, flashing it through Finder with the official firmware was the only option. I lost some photos, the rest I was able to restore from iCloud backup.
But filling up an SSD should not brick your volume like that. This is a filesystem implementation bug, not a user error.
not sure it this was OS or filesystem feature, but it refused to allocate literally anything if free space reached ~ 100-500MB on system partition so it was being kept _always_ usable, even logging was denied IIRC
Does it actually do that? I.e. stop the user from writing new data when storage space is extremely low?
Yeah, this has happened to me too. It became a lot more of an issue when my disk was (forcefully, non-consensually) converted to "APFS" which, along with breaking both alternative operating systems I had installed, also seemed to have a much greater chance of entirely bricking once I ran out of space.
I was consistently able to repair it by running a disk check in recovery mode, then deleting the files from the recovery terminal, however that is only accessible by a reboot, which by its very nature, necessarily loses all work that I had open, as it's impossible to save.
I never had this issue with HFS+. Everyone who says it can technically lock up is missing the fact that in practice it was still far more resilient than the newer and supposedly "better" APFS.
it gobsmacked me that I needed space on a filesystem to make space on a filesystem.
I think the whole stack of operating systems and tools that assume that this is possible get in trouble when it's not possible. I don't want my computer to become a locked down sandbox but it seems like this is where we are headed.
The only solution is to copy all the files you want to keep to another (not full) disk, then reformat and copy them back, or if you don't have another disk to copy to, somehow edit the disk directly to "manually" free some space.
It's true that mounting the affected Data partition on another machine won't help.
And booting Recovery and mounting is equivalent to mounting on another machine.
But it can help, as follows.
Before resorting to a wipe, try booting into Recovery, mount the Data partition, then use rm on a large file. When it fails, unmount the Data partition and run fsck -y on the Data partition at the commandline. If it finds errors and fixes, you'll get free space.
If you can't figure out how to mount the Data partition read/write at the commandline, close Terminal and run Disk Utility, locate the "Data" partition in the sidebar and right-click to Mount. Quit Disk Utility and restart Terminal. You will find user's data files the /Users folder.
You can locate large files on the mounted Data partition with find /Users -size +100M
Use df -h /Users to verify that 100M or more are free. Find and rm large files as needed. Then boot normally and finish cleanup from comfort of normal operation.
Note— Running fsck via Disk First Aid in Disk Utility should be the same as running fsck -y at commandline, but the UI for Disk Utility can be confusing. For example if it can't unmount the drive to perform the repair, DU will misleadingly advise you that the drive has failed and cannot be repaired. DU has some other odd behaviors, so it's more effective to use fsck at the commandline.
Another fine point: It's the specific APFS "Data" partition filesystem that's locked up (user's data), so you need to repair that specific volume, e.g. disk2s2. Look up the Data partition with disutil list Repairing the drive as a whole (e.g., disk2) is not what you want; this just checks that there's a partition table, which will naturally be OK. Similarly, repairing the other APFS system partitions will not help, nor will repairing the APFS Container disk. Fix the specific "Data" partition.