Who invented file extensions in file names?
retrocomputing.stackexchange.com
retrocomputing.stackexchange.com
[0] where different use cases might be satisfied by e.g.
"a text file"
"a file exclusively comprised of alphanumeric UTF16 characters"
"a CSV file"
"a file with tabular data and key-value metadata that calls itself CSV but is not spec-compliant so internally we'll call it CSVprime"
"A file with tabular data written as ASCII text that represents $this_kind_of_measurement"
... etc
If you have any resources on this topic off the top of your head I'd appreciate it if you shared themThat's what allows it to do complex things, eg identify all the flavors of ELF objects even though they share a magic number or determine if something is JSON or CSV without one.
Think about it: A JSON file can also be considered a text file. It could also be some higher level type of file, depending on whether it conforms to some application-specific JSON schema. Thus the kind of file it is has more to do with what you want to do with it; it's not some intrinsic property of the file.
what would happen to type theory if, say, the type of an object has a probability attached to it?
A Curly braces Separated file filled with Values
This is the authors website. Apparently yeah its not part of GNU utils, I had no idea, I knew it came with most Linux systems so I looked for the Debian package and found the site linked above.
Too late to edit!
% file crawl.mjs
crawl.mjs: Java source, ASCII textEdit: yep, the test for Java source is just /^import.*;$/
I wrote a Ruby library that attempts to be good at this https://github.com/okeeblow/DistorteD/tree/NEW%E2%80%85SENSA...
might wish to discuss files containing type information, and data/value information both, thus numerous but different valued files can all be of the same type.
you can get fancier if you want and talk about valid C source files or C source files which contain bugs, etc.
you could consider a file "weakly or adhoc typed" if there is only one or a pair of agents that use that type, and "strongly or standard typed" if there are numerous "different/distinguished" agents that "rely" on that file type.
this was off the top of my head
For example, for CSV files we can use .csv.txt extension, so if shell has no software for CSV files, it will treat them as a plain text, and if it has a software for viewing CSV files, but not for editing or printing them, it will add these verbs from records for txt files.
This is backwards compatible and would require minimal changes to file managers and graphic shells.
Another example: you can easily grep through a JSON file sometimes. But what if that JSON doesn't contain newlines? In that case it would be great if the system knew how to convert the data to a multiline string that makes sense for grep. This cannot be done if the data is "just bytes".
Further, JSON data can be represented with less bytes if you know the structure.
However, there was a trick that I encountered: copies made not by a CLI command or a properly written utility, say by sending a file over email, did not preserve the type in the filesystem. Then you would try to compile or assemble the code somebody sent you, and get a baffling message about your file being of the wrong type. If you had encountered this before, you would
REN xyz.asm xyz.xyz
CREATE/TYPE=ASM xyz.asm
COPY/APPEND xyz.asm xyz.xyz
DEL/V xyz.xyz
and go on with your work.Digital Research (the creators of CP/M) also didn't invent this. For more details, see the stackexchange link.
As a sidenote it's sad that WinFS failed.
Local files are even easier to lose and even harder to search than cloud.
Hell, try as they might, Microsoft still doesn't have ubiquitous OneDrive usage among Windows users.
This is such a web weenie take. Not everyone lives in the "cloud".
So annoying that we have to call it "the cloud".
Most users today don't know what a file even is[1], and even in the 90s and 00s most users still didn't understand files. Ever seen a desktop filled with icons of all sorts because the user just dumps all their files there? I'm sure you have, and most users are like that. Navigating a file system and using it is fucking pig latin nonsense to them.
Couple that also with users not taking backups because it's an inconvenience (until they wish they had one), and cloud storage solves all the problems most users have had. They don't need to manage their files anymore and they're all backed up by the cloud provider; they don't need to worry about their konpoohters anymore as they go about enjoying their lives, if they even have a konpoohter at home at all anymore.
[1]: https://news.slashdot.org/story/21/09/27/2032200/students-do...
And possibly an undervalued reason why the iphone took off. Anyone else remember this image? https://cdn.osxdaily.com/wp-content/uploads/2011/03/windows-...
Maybe I'm doing something wrong, but other colleagues have complained about the same thing, so I'm not sure...
Not if you have your own NAS, and you can use programs like everything (1) for an easy quick search, and you are in control of everything not under someone else’s policy mercy.
(1) voidtools.com
They allow the extension to be the fire emoji
Which is a shame because the people who created Mojo are definitely not amateurs.
But, when you're trying to show off your Unicode support to Californians, emojis get more reaction than umlauts.
These things seem unrelated—experienced people are allowed to have fun too. I don't see how this correlates with lack of experience at all. This just speaks to your personal bias against having fun.
Yes. Yes you can.
And God help us all.
I am literally going to put some CSS emoji easter eggs in my web app right now... o_O
Edit: I got my htm and html backwards and now comment not sense. Fixed now but obviously the folks below were right / responding to an earlier version.
That way if I change my site infrastructure from, e.g. static html, to php, to mediawiki, to cgi-bin, to asp, to ruby on rails, all my URLs can stay the same (because cool URLs don't change) throughout, without containing misleading legacy baggage.
Oh, yeah. I mean, you have to configure the server to treat extensionless files as html/cgi/php/whatever anyway, so you could just have it load extensionless URLs from .html/.cgi/.php files, but I prefer to keep my server config as "surprise-free" as possible. If `/article/foo` is served from the file `$DOCUMENT_ROOT/article/foo` on the filesystem, I find that less confusing for everyone than if it's served from `$DOCUMENT_ROOT/article/foo.php`.
It's a matter of personal preference, for sure. Just like `.htm` over `.html`. (Or `.jpg` over `.jpeg` for that matter.) I wouldn't say the way I do it is "right", but it's a trade-off that works for me.
Where the inertia of the history of 3-letter extensions in DOS and the likes is what caused this convention, of course
.rust.programming.language would make far more sense than pretending Rust will ever need to work in DOS.
You are not looking close enough. While Java source code has four characters in its file extension (.java), the output of its bytecode compiler has five characters in its file extension (.class).
perl -> pl
python -> py
rust -> rs
ruby -> rb
javascript > js
But java?
java -> ja? jv? jav?
Clearly Sun should've gone with a different name for the language entirely, I dunno, maybe they could've named it after an oak tree instead:
oak -> ok
Hmm, maybe I could get used to "jv" after all.
Then Perl tried to steal it.
No one uses the latter, though.
typical java, prioritizing verbosity at the expense of efficiency /s
Not entirely sarcastic though, looking at e.g. DescriptiveFooThingFactoryFactory
The Spelt Out File Extension Pattern?
I think the original Mac OS fucked us all by implementing this pretty much perfectly in the 1980s, but using their weird HFS "resource forks" to do it so that it not only couldn't be implemented on other platforms, but also lent credence to the "ugh fuck it lets just mangle the file name... and then hide it~!!! wooo000!~!" argument..
¯\_(ಠ_ಠ)_/¯
UPDATE: My memory was faulty; the "what kind of file is this?" metadata was actually stored as per-file metadata, the "type" and "creator" fields in the HFS(+) filesystem.
While a filename extension could probably be superseded by file metadata, it's also nice when you can read the intention right in the name without additional tooling.
Especially when you consider a hypothetical world without filename extensions thus every UI shows the filetype next to the filename thus "readme.md" in our world is just "readme:md" in their world, and it's not obvious what was gained nor lost.
In fact, I always want to see file types alongside their names, so I really don't mind that the two are together.
Separately, what `ls` displays with or without switches would be independent of how the extension is stored. Say the extension were a distinct field, `ls` could still display that as `{filename}.{extension}` by default.
It's an accident of history that the extension is part of the filename.
But that would break piping the output of `ls` to other tools that expect a filename.
https://mywiki.wooledge.org/ParsingLs
https://www.shellcheck.net/wiki/SC2010
It's brittle and the output is designed for humans, not machine parsing.
Use a glob or find.
In a world where extensions were their own distinct metadata, whether or not to show them would be a switch to `ls` (whose output you shouldn't be parsing) and how to retrieve that metadata from a filename w/o parsing `ls` output would be something like:
ext=$(stat --format="%E" -- "$filename")
You'd also have a switch to `find` to select files by extension instead of having to pattern match their names.My rule is that it's ok to use `ls` only in a directory where you know the filenames aren't weird. That rule is the same regardless of whether a human or a computer will be parsing its output. In general, reusable shell scripts shouldn't make assumptions about their environment, so they shouldn't use `ls`; but piping `ls` to another command in a one-off manner at an interactive shell is fine if you know the filenames in the directory aren't weird, just like it would be fine to use `ls` at the same interactive shell and simply read its output yourself.
But filesystems seem like the biggest species-level fuckup.
Windows had their shit, Mac had HFS then "HFS+" until... tbh, quite breathtakingly, like five minutes ago, When they finally introduced APFS. But like... no checksumming, just Mötley Crüe data integrity, it's all good, iCloud backups FTW amirite!?!
ZFS is good; I use it. But it didn't pan out, because a humanity-saving filesystem has to work on the systems humans actually use... (T_T)
I do not blame them one bit.
Nevertheless, Apple built something objectively worse (T_T)
That crown would be a pretty tough one to hold, there are some good candidates.
It's more a classic case of Apple's "not invented here"
Once Oracle bought Sun, it was like "Which is more important, your filesystem or your OS?"
Which isn't a hard question with regards to any filesystem I've ever used except ZFS. Seems insane to use anything else for long-term storage, unless it has some kind of high corporate budget.
ZFS took just about 5 years to mature into a production release. From its inception in 2000 to production release in 2005.
Oracle's btrfs (their attempt at competing with ZFS) began in 2008. Still isn't mature.
That’s incorrect. File type and creator code always were stored where they belong: in the file system.
Also, what’s weird about alternative file streams?
> so that it not only couldn't be implemented on other platforms
The main reason it “couldn’t” be implemented elsewhere is because of the dominance of FAT, which didn’t think it worthwhile.
Which are bad because those don't survive filesystem transmissions. A filename with extension will.
Because when you copy a file to some other system (e.g. USB drive, or email attachment, etc), only the common properties survive. So having a filesystem with awesome capabilities that the others lack becomes worse, not better. (T_T)
It won't actually because paths are not portable across file systems. The big reason is case sensitivity (and yes, some file extensions are case sensitive), but there's also issues with ASCII vs UTF-8 vs UTF-16 encoding.
Meanwhile there are multiple standards for exchanging file metadata between file systems (SMB, NFS, FUSE, FTP, whatever). Most of these support arbitrary metadata on files.
For example: HTTP does not supply that metadata, so a file would lose it's "extension". Where would you store mimetype? Not in the file obviously.
In fact it's arguable that metadata is more reliable than file names, because that metadata is more standardized.
> For example: HTTP does not supply that metadata, so a file would lose it's "extension". Where would you store mimetype? Not in the file obviously.
Not sure what this point is. It's up to the HTTP serving application to determine what the mimetype of another file is however it needs to. A limited design could use file extensions, but a resilient design would just query the file system for unambiguous metadata. Like I said, there are multiple ways to do this.
There's a lot of stuff that's not "in" the file. Like it's name.
Those are not hard requirements. It's nice to have, but in doubt it will just be what the destination filesystems think those values should be.
>There's a lot of stuff that's not "in" the file. Like it's name.
Correct. And everything not in the file is in danger of being lost. The filename (with extension) is the minimal viable data that has a chance to survive. The extension even more so than the basename.
Apple type/creator was a superior metadata system for sure, but the UI was opaque for non-technical users. Paired with the modern "open with..." settings, it would still be better than how file extensions are used today.
From the latest edition (#21):
> Technical Note: The electronic edition of this magazine is valid as both PDF and ZIP. Thanks to Ange Albertini, it is also a PCAP-NG packet capture of an experiment by Yannay Livneh. See page 7.
Edit: Ah, and of course, Cosmopolitan (by @jart) that produces an amalgamation of formats bundled into one file (including ZIP) that runs across a bunch of OSes https://news.ycombinator.com/item?id=38101613
Even in the world of image uploading via browser, just checking extensions is discouraged. If the feature is that humans can be allowed to manipulate it, it must be sanitized/verified before accepting. This isn't just for text heading for a database.
Where's the best place? It seems like being able to send a file handle to an app without having to open the file nor read a sidecar/db could be quite convenient?
i'm not saying throw away extensions, as they are great hints, but trust without verifying is not just for politics.
The better solution I've seen proposed is to store file type information as metadata in the filesystem itself, but this would lead to compatibility problems with archiving tools and any other scenario where a file is moved between differing file systems.
File extensions are the best way to make determining a file type easy without parsing the whole thing while also maintaining cross and backwards compatibility.
What's the web equivalent of using a floppy disk to move data between Mac, OS and Windows, I wonder. Downloading to the local computer and uploading, presumably. I guess that is pretty equivalent. Though you can't even download the native representation of a Google (text, spreadsheet, ...) document.
Is it a flame war or a religious debate? E.g. vi vs emacs
Just open the program first.
People should fear untrusted programs, but should be at ease viewing data through programs that they trust.
(2) I suspect for lots of files, lots of users wouldn't know which program to open.
As for (2)... good. Asking your computer to perform actions that you understand so poorly that you can't figure out what program is going to run is a recipe for disaster. That's the behavior we should be training people to avoid, not merely engaging with suspicious data but doing so recklessly.
That change would make computers completely unusable beyond Chrome and maybe Office for 99% of people.
As I understand it, a very large portion of the attacker's toolkit has to do with tricking users into running programs they've never heard of by clicking things they think are familiar.
But the real disaster is not the successful attacks, it's the culture that we're creating where users are taught to click things and trust the OS default behavior while simultaneously trained to never click things that seem out of the ordinary.
It creates a paralysis in the user when it comes to exploring their tooling and fails to create a learning gradient. This widens the gap between them and people like you and me.
That's a disaster for them because they get taken advantage of by people in the know, and it's a disaster for you and me because they end up blindly supporting bad behavior (drm, companies mishandling user data, etc) since it's bad in a dimension that they've been locked out of by our failure to pave a path towards competence.
I find that insane, and that's basically what happens on Linux today.
[1] https://www.theverge.com/2019/2/21/18234448/winrar-winace-19...
With "open by extension" the .pdf file can only target vulnerabilities in my PDF reader, but with "open by mimetype" they could be trying to exploit any program that is configured to open files. If you're doing "open by extension" and know that Adobe's exquisite PDF reader will open when you double click a PDF, doing so is not more unsafe than opening up Reader and navigating to the file from within.
That's without mentioning the main obvious advantage of "open by extension" - the fact that I can configure programs that open file formats that haven't been blessed by whoever wrote your `file` utility.
You're correct, which makes it doubly confusing as to why Windows doesn't do it.
(Seriously if you've been a vimmer for a while and wanted to give emacs a go give Doom Emacs a try and drop hlissner a couple of bucks it's definitly worth it)
In fact even Apple DOS 3.1 (1978) supported 8 different file types.
The SOS filesystem was reused for ProDOS (and later GS/OS as well), but a lot of the advanced features were removed for ProDOS due to wanting to fit into the smaller RAM of the Apple ][.
I’m guessing this could even be a NeXT leftover because Image Capture is ancient. The .jpeg extension was more popular on 1990s Unix.
Which is super-annoying, even though I agree with in in theory. But generally the only reason I am writing an image to .jpg is to email it to somebody or upload it to some web app and that principled .jpeg filename extension certainly isn't helping compatibility..
But later on microcomputers in CP/M and thence MS-DOS, yes.
You would not need to hide file extensions then.
file name contains the type can be read by human.
Just try to write a program that lists all your photos. You have to know all the various extensions that apply to still images (.jpg, .jpeg, .gif, .png, etc.) and do a string comparison against every file in the drive hierarchy. Might as well go to lunch waiting for it to finish if you have many millions of files to search.
This was the primary reason why I set out to build a better system that could scan through hundreds of millions of files and find all that matched a given subset in just a few seconds. https://www.Didgets.com
It was basically instant without any explicit indexing on my end.
On my Mac, the document icons for file types still correspond to whatever program I've chosen to open that extension type by default.
Which is so strange -- an .mkv file isn't a VLC document; an .mp4 file isn't a QuickTime document, an .mp3 file isn't an Apple Music document, a .jpg isn't a Preview document, and a .tiff isn't a Photoshop document. What the heck?
I wish the OS would just define reasonable, attractive, default document icons for every standard file type, that applications couldn't control/overwrite. Reserve application-defined document icons just for those apps who truly have their own proprietary file format (e.g. .psd or .docx).
The file icon would be assigned by the "Creator" Application. Double-clicking the file would cause it to be opened by the "Creator". You could choose to open a file with anything that advertised that it could work with a "File Type". There were utilities to change the "File Creator".
It's an additional piece of information, rather than just mirroring what the extension already tells you.
And the application icons are often super ugly, and often don't distinguish between the filetypes anyways.
Honestly, I'd rather do away with default application handlers altogether, if I've got multiple apps installed for a file type. Every time I double-click on a .jpeg, just ask me if I want to open it in Preview or Photoshop. When it's a .pdf, ask me about Acrobat or Preview. When it's an .mp4, ask me about VLC or QuickTime.
I have different reasons for wanting to use a different app each time (do I want to consume or edit, and edit how?), and trying to remember which one is the default handler and when I need to pick the non-default one is cognitive load I don't want to deal with.
Especially since the application-specific icons also sometimes don't even always make it clear which app it is. E.g. the IINA player assigns its own dedicated icons to video files, but you'd never guess that they open IINA.
Just give me standardized icons that are recognizable at a glance, and let me pick the application I want to use at the moment.
Aren't they? Most of these tools can open many specific types of documents, so you might want to know if a document is a movie or a picture in addition to what program will open it.
And of course there's the issue that Preview is actually an editor.