Where is my file's metadata?
eclecticlight.co
eclecticlight.co
ADS may also be a security risk[3].
[0] https://learn.microsoft.com/en-us/openspecs/windows_protocol...
[1] https://learn.microsoft.com/en-us/openspecs/windows_protocol...
[2] https://learn.microsoft.com/en-us/sysinternals/downloads/str...
[3] https://blog.netwrix.com/2022/12/16/alternate_data_stream/
HPFS was followed by Silicon Graphics XFS and MS NTFS, both in 1993.
After XFS became open source, it was the first Linux file system with support for extattrs, which is why it still has its own commands for handling them, besides those that work with any other Linux file system.
Then the rest of the Linux file systems have also adopted this feature, with the main exception of tmpfs, which for many years did not support any extattrs, then it added support only for certain security-related extattrs and only now support for user extattrs has been announced, but with some limitations.
Metadata onwed by a file ought to be embedded in the file's data, so that it will be preserved if you copy the file to a different system. This is the case for images, pdf, mp3 etc.
Metadata attached to a file is useful to simplify many 3rd-party programs, but should generally not be preserved if you copy the file. Stuff like display attributes for Finder, quarantine attributes for the virus scanner, etc.
Each of these requirements deserves a different implementation approach, which is why the current state of affairs makes sense to me.
Any file copy operation should by default make a perfect copy of the source file, without losing any kind of metadata, because that is the meaning of the word. It is not advertised as a partial copy command.
Even when some metadata would be useless of the destination, the purpose of the copy may be to store temporarily the file in another place, with the intention to bring it back later to the original location, when no alteration should have happened to the file.
If a complete copy cannot be done, e.g. because the destination file system cannot store some of the metadata, e.g. it has timestamps with a lower precision or it does not store some of the extended attributes, like the Linux tmpfs file system, the copy operation should fail with an explicit error message, like when the destination does not have enough space for all file data.
Any file copy command should have options to strip some or all metadata from a file, when that is desired, but that should not be the default option.
In the past I had unpleasant surprises with the Linux coreutils or with ssh, which lost silently metadata, even with the appropriate command-line options, which should not have been needed, in the case of coreutils because they had a compilation option to ignore the extended attributes and that option had been used in the binary packages available for certain Linux distributions.
When copying over the network, I always use rsync over ssh, because it copies reliably all metadata, even between different operating systems and file systems, e.g. Linux, FreeBSD and Windows.
I agree that "rsync --archive" ought to preserve extended attributes, same with creating a tarball. But when copying a file, as opposed to creating a backup/archive, I don't think extended attributes should be included by default.
And yes I've also had to compile custom versions of rsync at one point to get a version which understood extended attributes on Mac, but I think that's mostly standardized now.
EDIT: Also most programs don't even preserve permissions/ACLs on copy, so not sure why they'd preserve attributes by default.
The first time when I have used cp, more than thirty years ago, it was a shock for me to discover that this is the default behavior, because it is not something that I can associate with a command named "copy".
Since that day and until today, I always alias the various copy commands like cp or rsync to include all the necessary options for making perfect copies.
I have never needed to use them with different options, which would make imperfect copies, but there have been countless situations when my usage would have been broken by lossy copies.
I get that this is a Mac-focused site, but isn't starting at the personal computer era missing out a big chunk of history? My understanding was that in the early days of mainframes and then minicomputers, records/files consisting of multiple fields, in a system-standard structure, was the norm, and that Unix was somewhat radical in pushing for the minimalist alternative that "a file is just a bag of bytes [in a hierarchical namespace]". I seem to recall this being discussed in the Unix Haters Handbook, in fact. So while the xattr model may not exactly match the "old way", conceptually it was something of a return to form rather than a wholly novel idea
"We went to lunch afterward, and I remarked to Dennis that easily half the code I was writing in Multics was error recovery code. He said, "We left all that stuff out. If there's an error, we have this routine called panic(), and when it is called, the machine crashes, and you holler down the hall, 'Hey, reboot it.'"
https://multicians.org/unix.html
footnote: Really the real reason unix was popular was that it was available. AT&T was banned from the operating system business. so it was very easy to get a copy(with source). Many university computer departments did just this. and all those newfangled CS students started hacking on and graduated familiar with the system and brought it to industry with them.
https://en.wikipedia.org/wiki/AppleSingle_and_AppleDouble_fo...