The argument in favor of newlines is the same argument in favor of spaces in filenames or filenames starting with hyphen: the filesystem should not be restricted by bash and other shell scripts continual confusion between data and instructions. We should fix/replace lazy programs, don't try to fix lazy programming by neutering the system.
I think the correct action should be to design shell scripting languages that are robust about data versus instructions. The lack of robustness in shell scripting languages has been a burden since their inception. Restricting the filesystem won't stop problems with shell scripts making the same mistakes on other kinds of data.
Are there cases where these would be desired in GUI environments? Why and what is the use case?
-- Myfile With Emphasis --.txt
Basically, users will try to use any punctuation they see on their keyboard to adorn their filenames.
I'll happily agree with the spaces, because people like filenames to look like short phrases. But I've never seen a human make filenames with a leading hyphen (plenty with hyphen within the word) nor with a newline.
I could imagine that somewhere there is a person or two that likes to make hyphen-led filenames to try and sort things, but this is catering to the outliers. I can't imagine the use case for anyone to intentionally make a filename with a newline in it, and haven't seen such a thing, human or program-made. I think it's not common or valuable enough a thing to hamstring everyone else.
Edit: solution in the traditional linux way: off by default, but you can always compile them back in if you love them that much :)
And by "once", I mean "yesterday". Now, is it reasonable to design a system that takes input flags through the names of input files? No. It is clever, yes, but it's not reasonable. The fact of the matter is however that these usecases, for better or worse, do exist. We should not change filesystems to break these systems (broken as they are already) just because bourne-related shells aren't fond of them.
The other solution is to agree not put overly weird shit in filenames.
Which seems more likely to happen in our lifetime? Which would you rather code against?
I'm baffled that people are so hell bent on what "should" be allowed that they're ignoring the fact that allowing any possible character introduces exceptional complexity while offering extremely marginal benefits. I mean, imagine trying to make a gui for a filesystem browser that allows for characters that flow both up and down along with left and right. And newlines all over the place. It'd be a mess. I mean.. why? Is filename expressivity really such a major deal?
If a human script is read top-to-bottom, then the computer should handle that without complaint. This may not be possible in all cases immediately -- there's far too much Unix braindamage weighting the world down to throw it all away at once -- but we should all be trying to liberate humans from the dead hand of New Jersey.
Meanwhile, new shells (or new versions of shells) come and go all the time. Could one of these implement a foolproof (or at least, foolresistant) way of escaping arguments? Maybe.
So, yes, if you are writing scripts that should be resilient to all malicious or accidentaly provided input, it's a lot of work, and you better have code for these common operations (input, output of user-provided filenames) in all applications that try to do useful things with filenames. And probably it means that you cannot use programs not providing a "--" parse-stopper or put the data on the end of a "ssh" command-line.
And there will always be an application confused by some random other seemingly innocuous character that the kernel "still" allows in filenames.
So the solution probably is not to limit the allowed characters, and also not to try to fix the useful and working parts of shells, but to augment them with the equivalents of html_escape() and mysql_prepare("...?...")/mysql_escape() where needed (to stay in the realm of SQL and web-apps).
If you don't put weird shit in your filenames, I won't put weird shit in mine. Let's not bother the kernel with our contract. :)
What's simple? Chinese? How about vertical scripts like Mongolian? http://en.wikipedia.org/wiki/Mongolian_script What about composition characters, where unicode is basically acting like a set of instructions just like a terminal control character?
> what possible use is there for having control characters in a filename, like newline or carriage return?
Terminal control characters are only a problem because the terminal is interpreting them due to another ancient broken paradigm, but the fact is that these filenames exist. They're present right now. You're breaking old code by changing the rules on them. Keep 'rm' from operating on them, and then I can't delete them when some stupid program goes and creates them. Keep 'vi' from operating on them and I can't look inside to see wtf caused it. That's what I meant by re-quoting Linus there. You don't get to pick the reality that you already live in, and the reality that we live in has silly messy filenames with "beep" in them.
Besides, unless you go hyper-restrictive like [A-Z0-9]{1,8}, there are going to be edge-cases. Just spaces are enough to trip up bad shell scripts. Spaces are effectively control characters as far as the shell is concerned. If you allow spaces, you already have to deal with the "bad" filenames.
I don't pretend to have a good answer, so I guess my point here is that these kinds of issues are more deep-seated than they appear at first. The fact that "ls" allows what's basically an injection attack on your terminal isn't ls's fault, or necessarily the terminal's fault, or the shell's fault, it's because of a whole generation of bad assumptions and "good enough" design that adds up across these tools. Just "limiting filenames" seems like a quick fix, but isn't going to fix it because that's not the root of the problem.
* Printable (space excluded)
* Doesn't change the flow/layout of characters (no weird reverse characters or superscript/subscripts or characters that flow up/down).
Will that make everyone happy? Of course not. But it's pretty much "expressive enough" and it makes dealing with filenames 10x simpler.
Which policy seems more realistic:
A) sensible restrictions on file names that will please the majority of people and keep most programs working, while slightly annoying a lunatic fringe that wants EVERYTHING.
B) Require that every program that handles files in some capacity can handle every single possible weird unicode and control code subtlety in existence? (Keeping in mind that even browsers, some of the most complicated pieces of software in existence, still tend to screw up the corner cases).
Simple enough for a Western language speaker, sure. Pretty inhibiting for a speaker of Kannada that they can't name their résumé file with their own name.
> Require that every program that handles files in some capacity can handle every single possible weird unicode and control code subtlety in existence?
I assert again that once you allow spaces, you're already in the situation that you need to treat filenames as sequences of bytes that need quoting.
Also, I want to point out again that almost all of these problems are only when writing shell scripts. C programs or e.g. GUI applications with file pickers tend (with exceptions of course) to not have these problems since you get the filenames into them.
So programs that show the filename in the title bar: how do they display it? Really big title bar? Align the title bar on the left or right? What if I mix english and Kannada? What then?
Or what if I start with an english word in the terminal, and then switch to kannada, what's the terminal supposed to do?
I haven't used a lot of asian language systems, but I'm going to guess two things that I'm pretty sure are true:
A) they're already implicitly dealing with these restrictions, whether the file system enforces it or not
B) they probably already have workarounds to deal with these cases anyway. (alternate scripts or whatever).
Regardless of culture, I'm going to make what I think is a pretty obvious statement: good filenames are easy to type and unambiguous. If your filename is neither of those, it's not a good filename. The point of a filename isn't to be as expressive as possible, it's to be universal, descriptive, and easy to work with. None of that requires special characters.
With regard to the space thing, I think it's overblown. Spaces are common enough that good software engineers think to handle it. Newlines in files? I've /never/ seen that, even though it's possible. Also, space doesn't require an escape character. Once you start getting to newlines, it does. Once you start introducing escape characters into it, it's an entirely new layer of complexity.
That hardly sounds like a file-system issue. That's an application and UI-design issue.
Just because you think some of these issues are hard and you don't know how to deal with them on an application level (beacuse you haven't had to yet), doesn't make that a valid reason into limiting the capabilities of the file-system.
It is a problem when you have languages with fundamentally different rules trying to share the screen, and that is also not a culturist observation; indeed one would think that denying this would be the culturist position.
Of the various points I mentioned, in modern times this is probably the least important, now that RAM, ROM, and CPU are so dirt cheap. This wasn't always true.
That's what transliteration is for. I have no problems writing my name transliterated into the Latin script.
You mean like ASCII was "expressive enough" until people whose language was not English started using computers?
Your western bias shows.
That "expressive enough" attitude got us code-pages and a million different encodings to cope with the limitations we had put upon ourselves, and the related (and sometimes impossible) compatibility and inter-op issues.
Let's not walk into that one yet again, shall we?
As you could imagine, it's a special hell supporting this language. What's happened now is... a compromise has basically been made. Nastaliq is how Urdu had predominantly been written for years and years... but with growing usage of computers, even native Urdu speakers are abandoning it, because there's such poor support for it.
This article talks a little more about this: https://medium.com/stories-that-matter/9ce935435d90
a) overhead for the whitelist is going to be an issue on some systems
b) I'm going to need a kernel update every time there's a new useful code point for example, the Indian Rupee Sign http://www.fileformat.info/info/unicode/char/20b9/index.htm
The next big problem is control characters. But really, the only people who need control characters embedded in filenames are people breaking into computers using them. It doesn't matter if you use a GUI or a CUI, people generally don't embed newlines, or escape, or tabs in filenames.
The leading "-" is probably more controversial, but the POSIX standard specifically notes such filenames as nonstandard.
Unix is tools, not policy. It's what you make of it, and that's its beauty.
As I understand it most unixes require that they are composed of valid characters (not random bytes (or more evilly bits)) and can be no longer than a file system imposed length (commonly 255 bytes).
As for making the kernel enforce a single character set for the filesystem: No. No. No. Put another way: No.
http://yarchive.net/comp/linux/utf8.html
From Al Viro:
Bullshit. It has _nothing_ to characters, wide or not. For system filenames
are opaque. The only things that have special meanings are:
octet 0x2f ('/') splits the pathname into components
"." as a component has a special meaning
".." as a component has a special meaning.
That's it. The rest is never interpreted by the kernel.Linus has talked about this in the past... because of the fact that Linux only cares about 1 "special" character (the ASCII slash), the only sensible way to do unicode for filesystem names is to use UTF-8, which is backwards compatible with ASCII and thus obeys the kernel's expectations.
It's one of those "you can have any color car so long as it's black" ultimatums... you can use whatever encoding you want, so long as it uses 0x2f for slashes, and slashes only.
This is potentially confusing but, ultimately, fine. The problem is slash (byte 0x2f, in ASCII '/') in multibyte characters.
The problems that arise elsewhere should be solved elsewhere.