Linus Torvalds: Apple's HFS+ is probably the worst file-system ever
linuxveda.com
linuxveda.com
It was already discussed to death here: https://news.ycombinator.com/item?id=8876319
http://support.microsoft.com/kb/100625
Case insensitivity like win32's is still kind of a mess, but it's unfair to paint the developers of NTFS as idiots when they actually got this right. Above the FS layer is the correct place to implement case insensitivity, because case insensitivity is a feature that aids users in selecting files, not a feature that aids applications in opening them or creating them.
Enforcing this invariant at the filesystem layer (with well defined semantics, which HFS+ at least documents semi-well) seems eminently reasonable to me.
That assumes you somehow created two files with names differing only in case, which I'd argue would happen only if you were naming your files with such generic names. Also, I don't think the fact that two files' names look similar should be any reason to treat them as the same - otherwise, would you suggest that the letter O and the number 0 (along with the letter o) should be considered identical for filenames?
(please note the tongue firmly planted in cheek, even if a lot of people actually DO work like this. which is terrifying)
>The maximum representable date is February 6, 2040 at 06:28:15 GMT.
http://dubeiko.com/development/FileSystems/HFSPLUS/tn1150.ht...
However, does anyone know the genesis behind which this particular date was chosen? It's far enough back to not cause any issues and appears to be close to the minimum date-time value obtained using the minimum value for signed 32-bit integer and the standard Unix Epoch (which would be around 13 Dec 1901), but other than that?
The 1900 epoch needs additional code to handle the fact that 1900 was not a leap year. Why not the UNIX epoch? Well that's another question entirely.
"Explains why the year 1900 is treated as a leap year in Excel 2000"
Mac OS was written in a time of memory scarcity; 1904 is just an easier choice for the epoch than 1900. That's also why they didn't choose the Unix epoch, as it would have required supporting multiple epochs, one for disk time stamps, and one to handle birth dates (time stamps on classic Mac OS are unsigned; there were no times before the epoch. And yes, there were no Mac users over 80 :-))
How about Straße, STRASSE and strasse?
For an American audience, how about Résumé and RESUME and resume?
How about Ⅸ and IX? Hint: https://en.wikipedia.org/wiki/Unicode_equivalence#Normalizat...
The answer to all that is that it depends on the user and their locale settings. So to get this right you would have to pass the locale settings into system calls that go into the filesystem so it can do the right thing on every call. But then what happens if the locale later gets changed? Before some file names may be considered equivalent, but with the new one they are different and things break. Ok how about a system wide locale setting that can't be changed? Are you going to require a disk reformat each time a different user is on the system, or a multilingual user who travels?
The right solution is exactly what Unix did in the first place. The filesystem stores names exactly as is without having to deal with locales. User space is then responsible for doing the right thing for a user. For example a word processor can treat Résumé and Resume as the same for Americans in file selection dialogs.
User space already has to deal with all this, for example by sorting file names. Here is how hard that is http://www.unicode.org/reports/tr10/ and a simple example is if z comes before or after ö. Or when showing a list of Swedish names to a German user do you use Swedish or German sort order?
A similar analogy is dealing with time. The best approach is to store the time as UTC in the filesystem and let user space show it in whatever way is relevant for the user. On the other hand if you store local time in the filesystem, you will end up in a world of hurt trying to work around that. (UTC to local time is a lossy conversion making local time back to UTC hard to work out.)
- Filesystem implementation (actual on disk bytes)
- Filesystem compatibility (adjustments to OS semantics)
- Generic filesystem layer in the kernel
- System call interface (between userland and kernel)
- Userland low level library (eg system call wrappers)
- Higher level libraries (eg libc)
- User interface libraries
- App level modules & libraries
The problem with the filesystem being case insensitive is that is the first level above, where there is virtually no context or user information. That is the hardest place to work out if the user considers Résumé and RESUME to be the same. The lower down the list the easier it gets. The important point is that how you store things and how you present does not need to be the same.
Command line tools typically use the libc level of things. It is perfectly possible to make that level so the system appears case insensitive (another poster mentioned Windows doing that).
Your Unix command line experience is very much affected by your locale. For example the ls command does the following:
- Sorts filenames
- Displays sizes (different locales have different conventions for floating point representation, digit grouping etc)
- Displays times
Environment variables like $LANG, $LC_*, $TZ etc control this in Unix.
Individual programs can also alter behaviour. For example bash can be configured to do case insensitive filename completion.
TLDR: while case insensitivity is a nice user experience for humans, it is awful for the lower layers including filesystems. The closer a program is to a human, the easier it is to give that nice user experience. There is no requirement for a case insensitive user experience to then make filesystems also implement that. And as my examples show, they couldn't do it correctly anyway.
I disagree. I don't think a program should make any assumptions about the filesystem namespace—it's up to the user how they want to configure it.
Case sensitivity may bother users sometimes, but the filesystem is not the right place to solve that problem.
Let's take other non-latin characters, like accented letters or cyrillic. Some of these alphabets have letters that act as 'a' but are not 'a'. Some of these alphabets have glyphs that look like 'a' but are not 'a'. All of these glyphs are encoded using different codes as seen by a machine.
If you want to support 'a' == 'A', then I could argue that I would want you to support 'a' == 'A' == <letter in cyrillic that means 'a'>. It is a very slippery slope, in my opinion. Logically, I would rather have 'a' and 'A' separated to avoid ambiguity, but feel free to disagree.
That's a pretty major downside, and I think the back story for his dislike of it.
To me it smacks of a rather egregious failure to grasp Postel's law. And case-sensitivity is just bad UX for the 99% of computer users who aren't programmers or IT; perhaps it makes sense on Linux but Windows and OS X are dealing with a different audience and Linus's failure to give that fact some credit makes him come across as tone deaf.
macadam.doc might be about a type of road construction.
Given how much Siracusa and Torvalds (to an extent) advocate for better user experience over ease of technical implementation, I feel I'm missing something. Our job as programmers is to encompass the complexity of the real world in our programs, not to make our users think like machines.
If differentiating between "file" and "File" is thinking like a machine, how isn't differentiating between "my files" and "myfiles" thinking like a machine?
Should "my files" and "myfiles" be collapsed to mean the same thing? What about "files"?
On the other hand, "myfiles" and "my files" have different numbers of characters. They sound different in your head, and you can easily point to the difference. It's the same as typing "alien" versus "a lien".
If "most people" cannot tell the difference, then I believe that would be a failing of the education system.
Spaces are lexically significant in the English language because they indicate the breaks between words. "My files" are two English words that convey meaning to English speakers; "myfiles" is not a word at all and is meaningful only if we charitably insert the spacing necessary to make them meaningful.
Capitalization does not impact meaning to English speakers in the same manner; "my files" and "My files" are understood to mean the same thing. Capitalization is in this case merely an artifact of sentence positioning and the like. It's orthographic convention and nothing more.
Thus, to make the user have to understand that what is a meaningless distinction in English is a meaningful distinction to the user is, by definition, forcing the user to think like a machine.
So, what you are saying, is that there is no semantic difference between "1 mm" and "1 Mm"? Or between "my next computer" and "my NeXT computer"? Or between "acorn" and "ACORN"?
At my day job, I work with people who have barely used a computer before. I could see an 80-year old woman having three versions of a letter she typed in Pages or Word, simply because she had mixed use capital letters, and then mistakening the versions and deleting the wrong file.
I could be totally making this up though. Anyone remember hearing this and can confirm?
NB: I use a case-sensitive HFS+ install for cross-compiling when it's necessary.
Also, this can trip up in several other contexts besides headers. For instance, someone renames a file in their repository, changing just a filename case. And then you'll have to scratch your head a bit to understand a git issue that will invariably arise when you pull in the changes. Unless you are fortunate enough not to have a deadline, you won't be so quick to reformat the machine.
This actually happened to me. In the end, it was quicker to spin up a linux vm than to reformat the machine.
EDIT: Also, you appear to have missed the part that says "There is a case sensitive option, but Apple actively hides it and doesn't support it." No support is really, really, really bad. Noone will deploy OSX machines formatted like that if it is unsupported.
- Creative Suite http://helpx.adobe.com/creative-suite/kb/error-case-sensitiv...
- Steam https://support.steampowered.com/kb_article.php?ref=8601-RYP...
- Unity http://forum.unity3d.com/threads/fatal-error-case-sensitive-...
Mac apps not running on case sensitive filesystems is probably more common than not.
While tastes are certainly subjective, the exploit caused by this brain damaged design was not subjective, nor a matter of taste.
...by default.
Sure, but he doesn't need to bitch about things just to bitch about things... it's not like case sensitivity is a problem for anything but poorly designed build systems.
It's kind of impressive that Apple and MS have managed to make such balls of crap work as well as they have.
If you want a good critique of HFS+, check one of John Siracusa's many rants[1].
Also, Gruber is not the author of Byword:
Even John Gruber of Daring Fireball and author of Byword…
So your Adobe example is flawed. They should be testing on case-sensitive volumes, but the vast majority of normal users don't want a file named README and readme in the same directory, and enforcing that invariant seems reasonable to me.
Which it does indeed, pretty badly. You would be hard pressed to find a worse FS still in use, excluding FAT, which will just never die.
Except the article doesn't contribute to the discussion and the quotes from Linus are unfortunately not newsworthy. Instead of a deep critique, perhaps pointing what it should learn from Btrfs and ZFS, he acuses Apple of not shipping a consumer case sensitive FS and criticizes its string and slash encoding, which although bad, are successfully abstracted even from most developers.
It then ends stating that Apple, like the weather, pays no attention to criticism. As empty an analysis as can be.
Windows and OS X use case-sensitive file-systems. That sucks. But if you are going to release software for those systems, you need to accept that that is the reality, and write your software accordingly. This is hardly the first issue Git has had due to this, Linus's time would be better spent fixing these issues rather than stamping his feet and expecting OS X and Windows to change.