I think the original Mac OS fucked us all by implementing this pretty much perfectly in the 1980s, but using their weird HFS "resource forks" to do it so that it not only couldn't be implemented on other platforms, but also lent credence to the "ugh fuck it lets just mangle the file name... and then hide it~!!! wooo000!~!" argument..
¯\_(ಠ_ಠ)_/¯
UPDATE: My memory was faulty; the "what kind of file is this?" metadata was actually stored as per-file metadata, the "type" and "creator" fields in the HFS(+) filesystem.
But filesystems seem like the biggest species-level fuckup.
Windows had their shit, Mac had HFS then "HFS+" until... tbh, quite breathtakingly, like five minutes ago, When they finally introduced APFS. But like... no checksumming, just Mötley Crüe data integrity, it's all good, iCloud backups FTW amirite!?!
ZFS is good; I use it. But it didn't pan out, because a humanity-saving filesystem has to work on the systems humans actually use... (T_T)
I do not blame them one bit.
Nevertheless, Apple built something objectively worse (T_T)
That crown would be a pretty tough one to hold, there are some good candidates.
It's more a classic case of Apple's "not invented here"
ZFS took just about 5 years to mature into a production release. From its inception in 2000 to production release in 2005.
Oracle's btrfs (their attempt at competing with ZFS) began in 2008. Still isn't mature.
Once Oracle bought Sun, it was like "Which is more important, your filesystem or your OS?"
Which isn't a hard question with regards to any filesystem I've ever used except ZFS. Seems insane to use anything else for long-term storage, unless it has some kind of high corporate budget.
That’s incorrect. File type and creator code always were stored where they belong: in the file system.
Also, what’s weird about alternative file streams?
> so that it not only couldn't be implemented on other platforms
The main reason it “couldn’t” be implemented elsewhere is because of the dominance of FAT, which didn’t think it worthwhile.
In fact, I always want to see file types alongside their names, so I really don't mind that the two are together.
Separately, what `ls` displays with or without switches would be independent of how the extension is stored. Say the extension were a distinct field, `ls` could still display that as `{filename}.{extension}` by default.
It's an accident of history that the extension is part of the filename.
But that would break piping the output of `ls` to other tools that expect a filename.
https://mywiki.wooledge.org/ParsingLs
https://www.shellcheck.net/wiki/SC2010
It's brittle and the output is designed for humans, not machine parsing.
Use a glob or find.
In a world where extensions were their own distinct metadata, whether or not to show them would be a switch to `ls` (whose output you shouldn't be parsing) and how to retrieve that metadata from a filename w/o parsing `ls` output would be something like:
ext=$(stat --format="%E" -- "$filename")
You'd also have a switch to `find` to select files by extension instead of having to pattern match their names.My rule is that it's ok to use `ls` only in a directory where you know the filenames aren't weird. That rule is the same regardless of whether a human or a computer will be parsing its output. In general, reusable shell scripts shouldn't make assumptions about their environment, so they shouldn't use `ls`; but piping `ls` to another command in a one-off manner at an interactive shell is fine if you know the filenames in the directory aren't weird, just like it would be fine to use `ls` at the same interactive shell and simply read its output yourself.
While a filename extension could probably be superseded by file metadata, it's also nice when you can read the intention right in the name without additional tooling.
Especially when you consider a hypothetical world without filename extensions thus every UI shows the filetype next to the filename thus "readme.md" in our world is just "readme:md" in their world, and it's not obvious what was gained nor lost.
Which are bad because those don't survive filesystem transmissions. A filename with extension will.
Because when you copy a file to some other system (e.g. USB drive, or email attachment, etc), only the common properties survive. So having a filesystem with awesome capabilities that the others lack becomes worse, not better. (T_T)
It won't actually because paths are not portable across file systems. The big reason is case sensitivity (and yes, some file extensions are case sensitive), but there's also issues with ASCII vs UTF-8 vs UTF-16 encoding.
Meanwhile there are multiple standards for exchanging file metadata between file systems (SMB, NFS, FUSE, FTP, whatever). Most of these support arbitrary metadata on files.
For example: HTTP does not supply that metadata, so a file would lose it's "extension". Where would you store mimetype? Not in the file obviously.
In fact it's arguable that metadata is more reliable than file names, because that metadata is more standardized.
> For example: HTTP does not supply that metadata, so a file would lose it's "extension". Where would you store mimetype? Not in the file obviously.
Not sure what this point is. It's up to the HTTP serving application to determine what the mimetype of another file is however it needs to. A limited design could use file extensions, but a resilient design would just query the file system for unambiguous metadata. Like I said, there are multiple ways to do this.
There's a lot of stuff that's not "in" the file. Like it's name.
Those are not hard requirements. It's nice to have, but in doubt it will just be what the destination filesystems think those values should be.
>There's a lot of stuff that's not "in" the file. Like it's name.
Correct. And everything not in the file is in danger of being lost. The filename (with extension) is the minimal viable data that has a chance to survive. The extension even more so than the basename.
Apple type/creator was a superior metadata system for sure, but the UI was opaque for non-technical users. Paired with the modern "open with..." settings, it would still be better than how file extensions are used today.
From the latest edition (#21):
> Technical Note: The electronic edition of this magazine is valid as both PDF and ZIP. Thanks to Ange Albertini, it is also a PCAP-NG packet capture of an experiment by Yannay Livneh. See page 7.
Edit: Ah, and of course, Cosmopolitan (by @jart) that produces an amalgamation of formats bundled into one file (including ZIP) that runs across a bunch of OSes https://news.ycombinator.com/item?id=38101613
Even in the world of image uploading via browser, just checking extensions is discouraged. If the feature is that humans can be allowed to manipulate it, it must be sanitized/verified before accepting. This isn't just for text heading for a database.
Where's the best place? It seems like being able to send a file handle to an app without having to open the file nor read a sidecar/db could be quite convenient?
i'm not saying throw away extensions, as they are great hints, but trust without verifying is not just for politics.
The better solution I've seen proposed is to store file type information as metadata in the filesystem itself, but this would lead to compatibility problems with archiving tools and any other scenario where a file is moved between differing file systems.
File extensions are the best way to make determining a file type easy without parsing the whole thing while also maintaining cross and backwards compatibility.
What's the web equivalent of using a floppy disk to move data between Mac, OS and Windows, I wonder. Downloading to the local computer and uploading, presumably. I guess that is pretty equivalent. Though you can't even download the native representation of a Google (text, spreadsheet, ...) document.
Is it a flame war or a religious debate? E.g. vi vs emacs
(Seriously if you've been a vimmer for a while and wanted to give emacs a go give Doom Emacs a try and drop hlissner a couple of bucks it's definitly worth it)
Just open the program first.
People should fear untrusted programs, but should be at ease viewing data through programs that they trust.
(2) I suspect for lots of files, lots of users wouldn't know which program to open.
As for (2)... good. Asking your computer to perform actions that you understand so poorly that you can't figure out what program is going to run is a recipe for disaster. That's the behavior we should be training people to avoid, not merely engaging with suspicious data but doing so recklessly.
That change would make computers completely unusable beyond Chrome and maybe Office for 99% of people.
I find that insane, and that's basically what happens on Linux today.
[1] https://www.theverge.com/2019/2/21/18234448/winrar-winace-19...
As I understand it, a very large portion of the attacker's toolkit has to do with tricking users into running programs they've never heard of by clicking things they think are familiar.
But the real disaster is not the successful attacks, it's the culture that we're creating where users are taught to click things and trust the OS default behavior while simultaneously trained to never click things that seem out of the ordinary.
It creates a paralysis in the user when it comes to exploring their tooling and fails to create a learning gradient. This widens the gap between them and people like you and me.
That's a disaster for them because they get taken advantage of by people in the know, and it's a disaster for you and me because they end up blindly supporting bad behavior (drm, companies mishandling user data, etc) since it's bad in a dimension that they've been locked out of by our failure to pave a path towards competence.
With "open by extension" the .pdf file can only target vulnerabilities in my PDF reader, but with "open by mimetype" they could be trying to exploit any program that is configured to open files. If you're doing "open by extension" and know that Adobe's exquisite PDF reader will open when you double click a PDF, doing so is not more unsafe than opening up Reader and navigating to the file from within.
That's without mentioning the main obvious advantage of "open by extension" - the fact that I can configure programs that open file formats that haven't been blessed by whoever wrote your `file` utility.
You're correct, which makes it doubly confusing as to why Windows doesn't do it.