I'm “still afraid to use spaces in file names” years old
twitter.com
twitter.com
Pro tip: rename your development directory (or even better: the workspace path in CI) to put a space and/or special characters in it.
Forces you to deal with this properly, and immediately ensures that every automated test checks this case without you having to remember every time. Hasn't been particularly inconvenient, since I'm autocompleting it 99% of the time anyway, and I haven't shipped a single path parsing bug since.
> Microsoft intentionally made programs install to C:\Program Files on Windows 95+ to force programmers to deal with spaces in filenames.
A former co-worker changed his name in our auth system to include an apostrophe, so that whenever we handled names wrong he'd find it.
If you consider spaces “unusual” I would say you haven’t encountered a single average user in your lifetime. Spaces in file-names is the single most common thing people have, outside programming environments.
As a x-plat developer, the only platform where I (still) regularly encounter these kind of bugs are platforms where solving problems through scripting is common, like Linux, where the primary means of operation is through stringly-typed statements getting parsed and processed in a untyped-fashion. It's not very reliable.
On Windows people more often use “real APIs” (because scripting doesn't really work as well), but then these problems just goes away.
Pros and cons, I guess.
thisismyconfig.txt vs this is my config.txt or this_is_my_config.txt
...i've forced myself to stop using spaces, character, and even cap. They are all constructs that provide minimal value for the extra complexity.
I changed my username to not contain a space because it was too annoying to deal with all the random dev tools breaking. The worst offender was probably npx on Windows [1] (resolved after four years by deprecating npx), but it was far from the only one (though the JS ecosystem was somehow the worst in this regard of all languages I worked with).
Saw a few hacks where malware authors used the RTL feature (which is baked into Windows) to obfuscate file extensions. It looked like .exe.innocuous-document.docx, but was actually .docx.innocuous-document.exe
League of legends doesn’t run until I sed files for instance.
I'm begging software developers to stop using subprocess APIs that take a string argument (system(), child_process.exec(), Process.Start(string)) and start using subprocess APIs that take an array of arguments (execvp(), child_process.execFile(), Process.Start(string, IEnumerable<string>).)
I never use / on Windows as a result.
The problem with that is that YOUR code may handle it, but your tooling may not. If my code formatter break on spaces, I'm not going to change the formatter.
Inject a Turkish 'I'. I don't know how to type or paste it here, but picture an English lower case 'i' that is upper case. It is a splendid way among many to shake out some loc bugs.
Mysterious character requirements that do not conform with Microsoft’s OS limits, limits on tbe fully qualified pathname length, etc.
Naming freedom needs a stdlib module
Enforce this in LDAP.
Strict convention is better than flexibility and predicting obscure edge cases that can fail.
This will also break any code in external tools that are called during the builds of your application and do not handle spaces correctly for whatever reason, thus making it so that you won't be able to successfully finish the build.
Then again, you probably shouldn't be relying on technologies like that, but when you're struggling to keep an old enterprise system alive, causing yourself more problems is not necessarily what you should do.
Still a good idea in most cases, though.
# Rename all files in a directory
rn() {
rename "s/ /-/g" *
rename "s/_/-/g" *
rename "s/–/-/g" *
rename "s/://g" *
rename "s/\(//g" *
rename "s/\)//g" *
rename "s/\[//g" *
rename "s/\]//g" *
rename 's/"//g' *
rename "s/'//g" *
rename "s/,//g" *
rename "y/A-Z/a-z/" *
rename "s/---/--/g" *
rename "s/---/--/g" *
}
I use this all the time, especially when I download files.In addition to CLI I use it from emacs dired-mode too:
(defun my-dired-detox ()
(interactive)
(dired-do-shell-command "detox" nil (dired-get-marked-files))
(revert-buffer))
I bind it to "_" in dired-mode.By the way: what's your beef with en dashes? I mean, if it was "everything should be 'HYPHEN-MINUS' (U+002D)", then fine, but why specifically en dashes and not em dashes?
find . -depth -name '* *' | while IFS= read -r f ; do mv -i "$f" "$(dirname "$f")/$(basename "$f"|tr ' ' _)" ; done normalize_names() {
rename "s/-/_/g" *
detox -s lower *
}
in my .bashrc.I've thrown some edge cases at it, and it handles it super well. It deals with consecutive "_", remove leading garbage, normalize unicode, and even prevents naming conflicts by opting out early.
Thanks you.
I have:
Movie Bla (2020)
Movie Bla (2020).mp4
But also: Movie_Bla_(2020)
Movie_Bla_(2020).mp4
Movie_Bla_(2020).srt
Would not like to lose files like the the srt.You should better be very afraid of using spaces in filenames.
You should do everything you can to support them but you have to know you'll invariably encounter countless cases where you'll have this or that tool that won't work properly with them.
I still live in a world where I cannot name a song from the french group L'impératrice with an eacute in the filename or my car's media system will display garbage (it's running QNX and I don't know which filesystem).
FWIW, and it should be food for thought, every single Git repository in the world contains a pre-commit hook sample (disabled by default but it's there) that enforces that every committed file in the repo is named using a subset of ASCII characters.
Every Git repository in the world has that example: let that sink in.
I use Git for documents too, not only code. Why shouldn't I use my native language?
I have an Android phone and I tell MusicBrainz Picard to save all files with ASCII-only names and Windows-compatible names for the ones that get sent over to the phone. Basically for this reason. Sometimes it's players on Android itself, but even more frequently, whatever bluetooth radio I'm connected to freaking out with non-ASCII characters.
And "path component" is not an arbitrary string either - e.g. appending a path component to the path should first require converting/parsing the string into the path component, and only if that's successful appending it to the path.
For maximum correctness, you want to turn it into a file handle as soon as possible, and do all operations through the variations of the file functions that end in "at", like: https://linux.die.net/man/2/openat
The downside of this approach is that you still technically have to carry the path around with you if you ever want to present it back to the user, because once you have a directory handle, you can get back to the root directory easily enough by following parent links and seeing what directories you end up in, but that may not be what the user "thinks" the path is, and they want to see their path, not a canonicalized one. And they're mostly right. And it's not easy to correctly track changes to their intended path from this basis either.
Basically, I don't know of a really solid, 100% correct way to handle this with any reasonable degree of effort.
I guess that depends on what you mean by "string". `open` and `fopen` need a char* path to open a file. Whatever fancy Path abstraction you use eventually becomes a char* string, because that's what the kernel needs.
Yes, paths have structure, but saying "a path is not a string" is equivalent of saying "C source code is not a string". Both are strings, and both are something else, represented by strings according to rules. Different internal representations have different advantages and disadvantages. I fully agree that for things such as "adding components" an internal sequence/list representation is better, but strings can pass arbitrary IPC or even ABI boundaries much easier for example. (And you wouldn't bat an eye for example when you see FQDNs like "www.google.com" passed as a string instead of as ["www","google","com"] because the string representation works pretty well.)
"This is `a
perfectly vali'd.\010! file name\377, despite the weirdness"
text processing is hard if you must support Unicode, and that means every Unix command line tool must implement or employ a text processor to handle input. it would be much easier if objects were passed back and forth. PowerShell got this right.
xset m 0 0
to xinput --set-prop 'pointer:Logitech USB Receiver' 'libinput Accel Profile Enabled' 0, 1
Everything seems to be going this way in Linux land. Longer names, harder to type names, camelcase names, spaces... I'm looking forward to an OS that treats command line ergonomics as a first class feature and where camelcase & spaces are verboten.We started out with 'eth0', 'eth1', etc. Which adapter was which could change when adding and removing a network card. That was bad, so that prompted the evolution.
Now we have 'enp1s0', 'enp0s31f6', 'enp13s0' and many similar variations. These are supposedly more stable across device changes. As it turns out, it wasn't.
But wait, there is more! Now we have the "predictable names" scheme that produces interface names that are even longer, and not even slightly easier to remember.
Read about the whole sorry saga here:
https://wiki.debian.org/NetworkInterfaceName
I do get that it is not an easy problem to solve, especially in the face of removable network interfaces (like USB Ethernet / WLAN). But surely this is not the best we can do.
The first one is some magical incantation.
Of course, if one doesn't write a CLI to begin with, this trade-off doesn't exist - you can have your cake and eat it too.
If you had comma-delimiting like in an algol-derived language, you wouldn't need to quote things with spaces.
edit: also, code is read more times than it is written, so optimizing for readability over brevity is generally a good move.
Ironically, even NASA doesn't like space.
https://www.nas.nasa.gov/hecc/support/kb/portable-file-names...
2021-01-01-some-important-document.pdf gives me the warm fuzzies. On the off chance that some more differentiation is needed, throw in an underscore and a whole new world opens up
2021-11-11_client_projectName.ext is also OK. But underscore separates fields, hyphens for space replacement.
I hadn't heard that before and I love it.
https://web.archive.org/web/20151007005513/http://www.911cd....
P.S. should anyone want to see/run the actual batch, a copy has been uploaded here:
http://reboot.pro/index.php?showtopic=18962&page=29#entry204...
Yes on the date format.
Saves you so much time.
2021-01-01_what-happened_who-did-it_possible-reason
https://docs.microsoft.com/en-us/windows/win32/api/processth...
The arguments are a single string. So you want to pass parameters with spaces in them? You've got to add quotes and stuff all of that into a single string. Instead of doing it in a more sane manner, like oh, the arguments to main().
Sadly such set includes loads of Java programs. If only SUN had shipped a standard way to generate isolated exe files in 1998... but they worked under the presumption that you'd have a JVM already there, because distributing that monster was difficult in dialup times, so you could just hand people a jar; and the enterprise market did not care, since they had webapp servers. Sadly it's an "optimization" that became obsolete very quickly but wasn't rectified until it was too late (java 9+).
And shells and other programs still have problems with perfectly legal characters in filenames too, like '!' or ':'.
I also love when you're using bash and you have a file with ! in the name, and you accidentally fail to correctly backslash it, you not only get "bash: !rest_of_filename: event not found", but it also fails to add that command line to the history, so you can't just hit up and fix it. You have to actually go to the mouse and copy and paste.
Without asking you to always quote and escape every file name - what alternative is there? If they tried this you'd probably find you didn't like it.
"It's nothing."
"What do you mean?"
"It's nothing... It's empty space. I never taught the computer how to read empty space!"
"I never taught Virgil how to fly."
Also back then Mac file names typically did not include an extension, because the file's type was stored as part of the metadata in its resource fork. I remember one time a friend of mine was visiting and was playing around with a paint program on my Mac. Being used to DOS, when she went to save her file, she typed a very short name, and then asked me what the proper file extension should be. I smirked and said, "That's not how you name files on a Mac. THIS is how you name files on a Mac." And then I named her file "Ailsa's Cool Picture". Her mind was blown. :-)
¹This is because the colon was the path separator. But since the classic Mac OS had no command line interface, the typical user would never type or even see a file path written out.
However, I found the lack of a command-line to be restricting.
In many languages it's a requirement. For example, in Romanian, there are 8 words that collide with „fata“ if you remove the diacritics (fata, fată, fața, față, făta, făță, fâța, fâță).
Given that we have to use diacritics, spaces don't seem like a big deal.
There is one big difference: CLI utilities don't usually care about diacritics (though encoding issues can throw a wrench in that), but they care a lot about spaces. So putting spaces in filenames requires properly quoting or escaping parameters, whereas diacritics does not. That makes one-off shell snippets and scripts a lot more annoying (though TBH I tend to shy away from those anyway, these days).
That is what context is for.
Edit: can't save, downloading works.
This can cause user confusion
sh, bash and cmd.exe are shit. The shell needs serious rethinking.
$ cat proba.sh
#!/bin/sh
echo "Using quotes:"
for i in "$@"; do echo "$i"; done
echo "No quotes:"
for i in $@; do echo "$i"; done
$ ./proba.sh "ho ho ho"
Using quotes:
ho ho ho
No quotes:
ho
ho
hoCmake doesn't support semicolons, because everything in cmake is a string, and ; is the list item separator.
PATH is separated by colons, so you can't add directories containing : to it.
"What do I use a folder for?", they ask, in the same breath that they request "some way to organize things logically".
The no-filesystem movement has worked hard to eradicate this scourge from user experiences, but I fear that this is the devils work. Computer users should know what a file is, and what its for - and they should know what a folder is for, and why they would want to create one to put their files into it ..
But yet: they don't.
It hasn't improved since the 80's. Taking away the users responsibility to understand these things, only makes computing worse. The fact that "special chars in paths" breaks things, also holds this factor into place, imho.
Is that the movement to store all your data as an amorphous pile of crap, and then provide easy-to-use search tools to actually find the content you're looking for?
On one hand, I really like the search tools that come from this. But I still like to actually organize my data, so I can browse it if I want to. Also, these search tools seem to only work well enough on macOS and fall flat on their face in Windows. (and no idea where Linux falls on this)
As you as you do anything programmatic in/out of these drives it all hits the fan. So I'd add to the original statement - "Avoid 'technical' companies with special characters in their name", it's just not right...
I'm wondering when the first generation of college students will start who have never used a physical keyboard to input text.
Also, emoji.
And even more fun is, when it mostly works, but then it doesn't and you notice too late.
https://github.com/whyboris/Video-Hub-App/issues/667
Any help would be really appreciated!
It broke on the first try on a jr hire's machine, the source checkout location was `C:\source code`.
ETA: not broken in a technical sense, but having to escape them isn't the best experience. So it's just easier for me to avoid spaces.
#!/usr/local/bin/sbcl --script
(load "~/.sbclrc2")
(require 'replace-all)
(in-package :replace-all)
(format t "file is ~s" (second sb-ext:*posix-argv*) (probe-file (second sb-ext:*posix-argv*)))
(let* ((args sb-ext:*posix-argv*)
(orig (second args) )
(newfn (if orig
(replace-all orig "(" "-")
orig))
(newfn1 (replace-all newfn ")" "_"))
(newfn2 (replace-all newfn1 " " "-"))
(newfn3 (replace-all newfn2 "&" "-"))
(newfn4 (replace-all newfn3 ":" "-")))
(when orig
(format t "renaming \"~a\" to \"~a\"~%" orig newfn4)
(multiple-value-bind (new-name old-truename true-newname)
(rename-file orig newfn4)
(format nil "new-name ~a old-true ~a new true ~a" new-name old-truename true-newname))))And it is one of the biggest reason I hate Unix shells as programming languages, it is a minefield. In fact I think that after a dozen lines, Perl is a better option. It has most of what shells are good at (i.e. running commands), but saner and more powerful.
E.g. Here's a random StackOverflow q&a about a Git pre-commit hook where the top-voted answer does not properly handle filenames with spaces : https://stackoverflow.com/questions/2412450/git-pre-commit-h...
However, the 2nd and 3rd most upvoted answers do mention "-z" option to handle spaces.: https://stackoverflow.com/questions/2412450/git-pre-commit-h...
Same goes for capitalisation. All filenames should be lowercase.
Maybe it's not strictly necessary, it can avoid headaches.
I mostly use the shell and navigating in directories with spaces is annoying, you have either to quote it or put a \ before each space. You also have to remember to quote everything, and in bash that can become complex, you start adding quotes everywhere to solve problems caused by spaces (or other special characters like *) in filenames.
So I prefer to not use them, a simple _ is as readable as a space. Only thing is that spaces gets rendered better on graphical file managers, but... that could have been solved (and can still be solved) by simply adding an option to render a _ as a space graphically if there is no ambiguity. I don't care that much since I don't use graphical file managers that much.
I'm old enough to remember working with 8.3 filenames in DOS, and while the length limitation was maddening, the space part never was. Then Windows 95 came out and all restrictions were thrown out.
Why couldn't we just have a file system that robustly supports long filenames, including variable length extensions, while prohibiting certain special characters - namely spaces, slashes or any directory denoting characters in files, and characters that have special meaning in regex context? (brackets, asterisk, etc.)
https://dwheeler.com/essays/fixing-unix-linux-filenames.html
node-gyp doesn't like it when there's a space anywhere in your working path. Stuff I was messing around with was all in ~/Code Projects at the time, and using npm install on some things just broke. Looking back, I definitely could have done a better job parsing the error messages but still...
There's an issue but it was closed in 2018 as "The workaround is to use a path without blanks" https://github.com/nodejs/node-gyp/issues/439
(defun mdy ()
(interactive)
(insert (format-time-string "%04Y-%02m-%02d")))
That inserts the "proper" date format (e.g., 2021-11-11) at the current point.Then to create a date-stamped file name:
(defun file-mdy (file-name)
(interactive "sbasename: ")
(find-file (format "%s-%s.org" (format-time-string "%04Y-%02m-%02d") file-name))
(save-buffer))
And a few others.Nobody seems to misunderstand this date format. US folks might find it annoying, but understand what it means.
It's cumbersome-ish, but can be made to work.
Then there's shell injection via files containing a newline character in their name...
The findnl script as part of fslint identifies problematic patterns, and has 4 levels of stringency, with "POSIX" being the most stringent. https://github.com/pixelb/fslint/blob/master/fslint/findnl
Spaces are used to separate parameters in the command line. There's also no real need for filenames to support spaces.
But, of course, if you mix the abstractions of metadata (filename) with location, things won't be trivial.
https://github.com/acmesh-official/acme.sh/issues/1408
This is the sysadmin equivalent of piercing your nose just to make your parents mad.
Let's test something: http://example.com/my silly webpage.html.
Hey look, HackerNews just broke a URL with spaces in it. And it's written in a Lisp dialect and all; it's not some Unix job cobbed together with shell, sed and awk. The language has a string data type, and strings are passed to functions without word-breaking interpolations taking place.
You know what else breaks on spaces? Basic everyday gui text manipulation.
Suppose that in a block of text we have the sentence:
> Please look for the Holiday Schedule 2021 file.
If you double click on any part of the name like Schedule, pretty much every text widget on the planet will just select only that word, and not the entire filename.
However, if you have:
> Please look for the holiday-schedule-2021 file.
There is at least a ghost of a chance that a semi-intelligent GUI can pick that out as a word.
There exist good reasons to keep identifiers as clump beyond just command line shells.
It's why we need encoding like %20 in URLs that never pass through a shell script.
I also don't use :, as I have ran into problems with both Bash and its completion and FAT FS. Unfortunately, I routinely have timestamps in filenames, so I need to use +%F-%H-%M-%S instead of simple +%F-%T.
One thing has improved, though: I have not run into problems with ěščřžýáíé (which my language is full of) for maybe a decade, except on OpenWRT where space seems to be scarce to support non-ascii.
Edit: I now remember one problem, getting images for a website from an OS X user, which used combining characters instead of direct code points (https://en.wikipedia.org/wiki/Unicode_equivalence#Example), but HTTP requests got normalized in some browsers, leading to strange 404s.
It's not fear that keeps me from using spaces in file names, it's habit.
If we're going to play this dangerous game, from now on I'll figure out how to use nulls (\0) in my file names, and make all the C/C++ programmers cry.
I do, however, use Cyrillic (UTF-8) in filenames, and I regularly try out if moving a file into ASCII-path will let some programs open it (half the time it's that when I am having trouble).
This might already exist, but I wonder about a terminal that was really just a multi-line repl to a language. It would be preloaded with libraries that replicated all the features of the gnu core utils, but instead of calling grep like normal, you called a function like grep("args"). The advantage would be that you had access to a full blown programming language at all times. So when you needed to do something more complicated you would still have access to all the standard language features. And when you didn't need that, your canned core utils like functions would work
Edit: thinking about it again, it might not have even been the space but the exclamation mark in my path. Or both.
Edit: replace $@ with quoted version which actually changes the behavior (I was wrong that the difference is between $* and $@).
One way to look at it is that people of a certain generation eschew spaces because the tools of their formative years simply couldn't handle spaces - but another is that the olds have learned that generally erring on the side of KISS ("Keep it simple, stupid!") isn't a bad idea.
The main culprit is GNU Make which does not cope with spaces in filenames. As far as it is concerned an array is a string separated by spaces so it gets very confused. Yes there are some partial workarounds, no none of them consistently work. You learn very quickly to check all code out in a file tree with no spaces in it, otherwise builds can randomly break in strange ways. It's not always clear up front whether Make is going to be involved somewhere in the build, so it's just easier to be safe.
How can I not be afraid of spaces if this happens like every other day with every other custom tool ...
Browsers will take http://example.com/some name.pdf and automagically turn it into http://example.com/some%20name.pdf, and deliver the goods without a problem. But having that space in the URL is still out of spec, and will cause your web page to fail validation, even though it works fine.
PS-Microsoft is horrible about stupidly named folders being created and dumped in there.
Last I checked, the standard answer for GNU make is "Spaces are expected to break the tool, that's working as intended, it will never be fixed." And because we build our towering edifices of software on the pillars of the past, I can't guarantee to you that a project of arbitrary complexity won't try to cram a list of filenames through a make script.
- Spaces in filenames get transformed to non-breaking spaces by the filesystem;
- The filesystem treats nbsp as equal to space (just as case-folding treats A=a, B=b, etc.)
Now, argument parsing, mouse double-clicks, etc. all respect filenames as "words", and the output from things like 'ls' just work.
(Yes, I'm well aware that there are case-sensitive filesystems out there. I'd forgotten that iOS was one of those).
That being said, even after all these years I sometimes need to try a few times in order to get the quoting and the escapes right when communicating names of files with spaces through multiple layers of software.
Often they are not even valid UTF8 which, when you uncork the filesystem for the first time in a decade causes the most delightful crashes. The more years the better the aroma.
But also a "use a font that has a proper capital ß" hipster.
At least in my crazy old illogical head anyway.
On Unix/Linux we've grown up with case sensitive by default but everywhere else it still seems to be a problem now and again.
I should qualify this...I'm en-US so I have no idea what the experience is like for anyone else.
[Introduction](./Introduction.md)\\
[Chapter One](./chapter one.md)\\
Crashed on trying to deal with building html when there are spaces in the file name. It is still an issue.Maybe we ought have to a different character signify the end of a name? Or signfiy a option section, or the next option section of a command?
Personally I wish console shells had chosen another delimiter than space, but here we are.
I still don't know how to process this emotionally. Either it is somehow naively really genius, or stupid.
In any case, it scares me, mostly because it is a non-IT person.
Except nowadays I worry more about user names that get fed into collaborating applications (with different edit criteria) and password characters (again for systems with differing, strange edit rules.)
Although lately I have started saving my Logic Pro files with spaces, simply because I prefer it to be the name of the song as-is.
Weirdly, my friend hates underscores. But he's a baseball fan
i still get issues with old one-off scripts, that still work, and I forgot to properly quote stuff... plus the urls are pain in the ass with the %20;s.
An old prof of mine used to send emails where the subject line was always a valid identifier in C.
Hello_dear_students_where_are_your_reports_
Damn, I feel old now :P
01 - Metallica - Metallica - For Whom the Bell Tolls.mp3
Names like that were common, and had many spaces.