Why not parse `ls` and what to do instead
unix.stackexchange.com
unix.stackexchange.com
Maybe I also don't understand shell, but as it was said before: when in doubt switch to a better defined language. Thank heavens for awk.
The tired old "stick to bash because it's already installed everywhere" argument is just as weak and misleading and pernicious as the "stick to Internet Explorer because it's already installed everywhere" argument.
It's not like it isn't trivial to install Python on any system you'll encounter, unless you're programming an Analytical Engine or Jacquard Loom with punched cards.
On top of it, shell is better than Python for many things, not to mention faster.
It's also, as you mentioned, ubiquitous.
In the end, choose the tool that makes more sense. For me, a lot of the time, that's a shell script. Other times it may be Python, or Go, or Ruby, or any of the other tools in the box.
For decades, on most Windows computers I run web browsers, there's always Internet Explorer. So do you still always use IE because installing Chrome is "wasteful"? It's a hell of a lot bigger and more wasteful than Python. As I already said, that is a weak and misleading and pernicious argument.
So what exactly is bash better than Python at, besides just starting up, which only matters if you write millions of little bash and awk and sed and find and tr and jq and curl scripts that all call each other, because none of them are powerful or integrated enough to solve the problem on their own.
Bash forces you to represent everything as strings, parsing and serializing and re-parsing them again and again. Even something as simple as manipulating json requires forking off a ridiculous number of processes, and parsing and serializing the JSON again and again, instead of simply keeping and manipulating it as efficient native data structures.
It makes absolutely no sense to choose a tool that you know is going to hit the wall soon, so you have to throw out everything you've done and rewrite it in another language. And you don't seem to realize that when you're duct-taping together all these other half-assed languages with their quirky non-standard incompatible byzantine flourishes of command line parameters and weak antique domain specific languages, like find, awk, sed, jq, curl, etc, you're ping-ponging between many different inadequate half-assed languages, and paying the price for starting up and shutting down each of their interpreters many times over, and serializing and deserializing and escaping and unescaping their command line parameters, stdin, and stdout, which totally blows away bash's quick start-up advantage.
You're arguing for learning and cobbling together a dozen or so different half-assed languages and flimsy tools, none of which you can also use to do general purpose programming, user interfaces, machine learning, web servers and clients, etc.
Why learn the quirks and limitations of all those shitty complex tools, and pay the cognitive price and resource overhead of stringing them all together, when you can simply learn one tool that can do all of that much more efficiently in one process, without any quirks and limitations and duct tape, and is much easier to debug and maintain?
The good thing about the JMESPath syntax is that it is the standard one when processing JSON in software like Ansible, Grafana, perhaps some more.
On its own, I agree. But you glossed over everything else I said, so I'm not going to entertain your weak argument.
You seem to ignore that different users, different use cases, different environments, etc. all need to be taken into account when choosing a tool.
Like I said, for most of my use cases where I use shell scripting, it's the best tool for the job. If you don't believe me, or think you know better about my circumstances than I do, all the power to you.
I have worked on projects that are extremely sensitive to extra dependencies and projects that aren't.
Sometimes I am in an underground bunker and each dependency goes through an 18 month Department of Defense vetting process, and "Just install python" is equivalent to "just don't do the project". Other times I have worked on projects where tech debt was an afterthought because we didn't know if the code would still be around in a week and re-writing was a real option, so bringing in a dependency for a single command was worthwhile if we could solve the problem now.
There is appetite for risk, desire for control, need for flexibility, and many other factors just as you stated that DonHopkins is ignoring or unaware of.
They learn. We all do.
Love it.
They clearly didn't realize that even more modern Unix kernels would require hundreds of megabytes just to boot.
That was 17 years ago!
I can do (ignoring parsing issues):
for name in $(cd subdir; ls); do echo "$name"; done
This isn't easy to do with globbing (as far as I know) for name in subdir/*; do basename "$name"; done for name in subdir/subsubdir/*; do
echo "${name#subdir/}" # subsubdir/foo
doneEdit: It turns out that Bash does substitutions in the middle of strings using the ${string/substring/replacement} and ${string//substring/replacement} syntax, for more details see https://tldp.org/LDP/abs/html/string-manipulation.html
find some/dir/here -name '*.gz'
then I could get the filenames without the "some/dir/here" prefix.It would also be nice if "find" (and "stat") could output the full info for a file in JSON format so I could use "jq" to filter and extract the needed info safely instead of having to split whitespace seperated columns.
JSON makes that a lot less fragile.
find . -name '*.hs' -exec basename {} \; cd some/dir/here
find . -name '*.gz'
cd - # changes back to previous directory (cd subdir
for name in *; do
echo "$name"
done)
But I prefer: for name in subdir/*; do
name="${name#*/}"
echo "$name"
done $ x=/some/really/long/path/to/my/file.txt
$ echo "${x##*/}"
file.txtOf course you can pass `sh -c '...'` (or Bash or $SHELL) to `find -exec` or xargs but then you easily get into quoting hell for anything non-trivial, especially if you need to share state from the parent process to the (grand) child process.
You can actually get `find -exec` and xargs to execute a function defined in the parent shell script (the one that's running the `find -exec` or xargs child process) using `export -f` but to me this feels like a somewhat obscure use case versus just using an inline while loop.
In your example, newlines and spaces in your filenames will ruin things. Better is
find … -print0 | while read -r -d $'\0'; do …; done
This works in most cases, but it can still run into problems. Let's say you want to modify a variable inside the loop (this is a toy example, please don't nit that there are easier ways of doing this specific task). declare -a list=()
find … -print0 | while read -r -d $'\0' filename; do
list+=("${filename}")
done
The variable `list` isn't updated at the end of the loop, because the loop is done in a subshell and the subshell doesn't propagate its environment changes back into the outer shell. So we have to avoid the subshell by reading in from process substitution instead. declare -a list=()
while read -r -d $'\0' filename; do
list+=("${filename}")
done < <(find … -print0)
Even this isn't perfect. If the command inside the process substitution exits with an error, that error will be swallowed and your script won't exit even with `set -o errexit` or `shopt -s inherit_errexit` (both of which you should always use). The script will continue on as if the command inside the subshell suceeded, just with no output. What you have to do is read it into a variable first, and then use that variable as standard input. files="$(find … -print0)"
declare -a list=()
while read -r -d $'\0' filename; do
list+=("${filename}")
done <<< "${files}"
I think there's an alternative to this that lets you keep the original pipe version when `shopt -s lastpipe` is set, but I couldn't get it to work with a little experimentation.Also be aware that in all of these, standard input inside the loop is redirected. So if you want to prompt a user for input, you need to explicitly read from `/dev/tty`.
My point with all this isn't that you should use the above example every single time, but that all of the (mis)features of shell compose extremely badly. Even piping to a loop causes weird changes in the environment that you now have to work around with other approaches. I wouldn't be surprised if there's something still terribly broken about that last example.
The "-r" flag allows backslash escaping record terminators. The "find" command doesn't do such escaping itself, so that flag will cause files with backslashes at the end to concatenate themselves with the next file.
Furthermore, if IFS='' is not placed before each instance of read, or set somewhere earlier in the program, than each run of white-space in a filename will be converted into a single space.
EDIT: I proved your point even more. The "-r" flag does the opposite of what I thought it did, and disables record continuation. So the correct way to use read would be with IFS='' and the -r flag.
> I think that when someone uses ls instead of a glob it means they most probably don't understand shell.
In 25 years of using Bash, I've picked up the knowledge that I shouldn't parse the output of ls. I suppose that it has something to do with spaces, newlines, and non-printing characters in file names. I really don't know.But I do know that when I'm scripting, I'm generally wrapping what I do by hand, in a file. I'm codifying my decisions with ifs and such, but I'm using the same tools that I use by hand. And ls is the only tool that I use to list files by hand - so I find it natural that people would (naively) pick ls as the tool to do that in scripts.
Is it really that difficult to add --json as a flag?
there are other tools than `ls` with their soul purpose to list files; some have "improved" features than ls, ect
similarlly from above (and well said btw, even what is not quoted): > I do know that when I'm scripting, I'm generally wrapping what I do by hand, in a file. I'm codifying my decisions with ifs and such, but I'm using the same tools that I use by hand
a lot of us do similar and we know/expect the ins and outs; and all it takes to break our scripts is some edge case we never thought of. we are fortunate that our keyboard layouts are basically ascii; other languages are less fortunate. now introduce open source community driven software where an ls escaped bash code deletes somebodys home directory (as ls was parsed and an edge case of some users files cause some obscure "fun" times). an edge case is still painful..
and finally, sometimes its better elevating said bash script to python (or awk), ect. just depends on the situation and level of complexity of logic
Look ls is one of the most basic and natural Unix commands. Make it modern and useful.
Bash gibberish is fun for gatekeeping scripting neckbeards, but it's not what a proper OS should have.
there's also using zero terminiated lines in ls with `--zero`; then piping that to a number of apps which also support similar (read,xargs,ect)
might also checkout powershell on linux which may suite your needs where instead of string manipulation, everything is a class object
ls *.jpg | awk '{print "resize 200x200 $1 thumbnails/$1"}' | bash
because I never got to the point where I could remember the strange punctuation that the shell requires for loops without looking up the info pages for bash whereas I've thoroughly internalized awk syntax.Word is you should never write something like that because you'll never get the escaping right and somebody could craft inputs that would cause arbitrary code execution. I mean, they try to scare you into using xargs, but I find xargs so foreign I have to read the whole man page every time I want to do something with it.
ls *.jpg | xargs -i,, resize 200x200 ,, thumbnails/,,
I just always define the placeholder to ,, (you can pick something else but ,, is nice and unique) and write commands like you do.
for i in *.jpg; resize 200x200 "$i" "thumbnails/$i"; endAt least on mac, the max command length is 1048576 bytes, while the maximum path length in the home directory is 1024 bytes. There might be some unix variant where the max path length is close enough to the max command length to cause an overflow, but I doubt that is the case for common ones.
xargs exists in an attempt to be able to parse command output. You could for instance have awk output xargs formatted file names to build up a single command invocation from arbitrary records read by awk. Note that xargs still has to obey the command line length limit though, because the command line needs to get passed to the program. Thus, in a situation where this for loop overflows the command line, it would cause xargs to also fail. Thus I would always use globbing if I have the choice.
EDIT: If you mean that the directory is splatted in the for loop, then in a theoretical sense it is. However, since "for" is a shell builtin, it does not have to care about command line length limits to my knowledge.
I've seen some image directories with more than a million files in them.
But when you execute a for loop in bash/sh, the 'for' command is not a program that is launched; it's a keyword that's interpreted, and the glob is also interpreted.
Thus, no, that does not fail when you hit the maximum command line length (which is 4096 on most _nix). It'll fail at other limits, but those limits exist in bash and are much larger. If you want to move to a stream-processing approach to avoid any limits, then that is possible, while probably also being a sign you should not use the shell.
$ for i in *; do echo $i; done | wc -l
1000000
I'm a little bummed that it failed in fish shell, but wouldn't begrudge the author if they replied "don't do that". find . -maxdepth 1 -name "*.jpg" -exec resize 200x200 "{}" "thumbnails/{}" \;
which works for spaces and probably quotes in filenames I am not sure about other special characters.I switched the command to a graphics magick based resize since that's the tool these days, default quality is 75% (for JPEG), but is included as a commonly desired customization. ,, is from a different comment in this thread; it seems better self-documenting than the single , I'd traditionally use.
find . -maxdepth 1 -name "*.jpg" -print0 |\
xargs -0P $(nproc --all) -I,, gm convert resize '200x200^>' -quality 75 ,, "thumbnails/,,"Also: "find -exec {} \+" will take ARG_MAX into account, and may be much faster depending on what you're doing.
In general yes globbing is better for iterating through files. But parsing `ls` doesn't necessarily mean the author doesn't know shell. It might mean they know it well enough to use the tools that are made available to them.
I finally started really using my shell after switching to it. I casually write multiple scripts and small functions per day to automate my stuff. I'm writing scripts I'd otherwise write in python in nu. All because the data needs no parsing. I'm not even annotating my data with types even though Nushell supports it because it turns out structured data with inferred types is more than you need day-to-day. I'm not even talking about all the other nice features other shells simply don't have. See this custom command definiton:
# A greeting command that can greet the caller
def greet [
name: string # The name of the person to greet
--age (-a): int # The age of the person
] {
[$name $age]
}
Here's the auto-generated output when you run `help greet`: A greeting command that can greet the caller
Usage:
> greet <name> {flags}
Parameters:
<name> The name of the person to greet
Flags:
-h, --help: Display this help message
-a, --age <integer>: The age of the person
It's one of the software that only empowers you, immediately, without a single downside. Except the time spent learning it, but that was about a week for me. Bash or fish is there if I ever need it to paste some shell commands. shopt -s nullglob
for f in *; do
…
done
But never this: for f in $(ls); do
…
done
They look similar, but the latter runs ls to turn the list of files into a string, then has the shell parse the string back into a list. Even if the parsing was done correctly (and it isn’t), this is still extra work. Looping over the glob avoids the extra work. ls | each { ... }
Another examples I don't need to explain, which would be far harder in stringly typed shells: ls | where type == file and size <= 5MiB | sort-by size | reverse | first 10
ps | where cpu > 10 and mem > 1GB | kill $in.pid
It's immediately obvious what you need to do when you can easily visualize your data: > ls
╭────┬───────────────────────┬──────┬───────────┬─────────────╮
│ # │ name │ type │ size │ modified │
├────┼───────────────────────┼──────┼───────────┼─────────────┤
│ 0 │ 404.html │ file │ 429 B │ 3 days ago │
│ 1 │ CONTRIBUTING.md │ file │ 955 B │ 8 mins ago │
│ 2 │ Gemfile │ file │ 1.1 KiB │ 3 days ago │
│ 3 │ Gemfile.lock │ file │ 6.9 KiB │ 3 days ago │
│ 4 │ LICENSE │ file │ 1.1 KiB │ 3 days ago │
│ 5 │ README.md │ file │ 213 B │ 3 days ago │
... for f in $(cd subdir; ls); do
...
done
? for f in subdir/*; do
...
done
or (
cd subdir || exit 1
for f in *; do
...
done
)
work fine. However, I must insist against using `for` loops in favor of `find`.Perhaps it would help to translate this into something more like, "what pitfalls do you run into if you parse `ls`" but it's hard to get past the initial language.
I'm pretty sure you can come up with scenarios where parsing the output of "ls" is indeed the simplest solution, but that kind of article is supposed to discourage people who don't know better from going "oh, I know, I'll just parse the output of ls". As a general advice, people should indeed be pointed towards "man find" or "man opendir 3".
I think the example of "exclude these two types of files" is a good case. I often have to write stuff like `ls P* | grep -Ev "wav|draft"` which doesn't solve a problem I don't have (such as filenames with newlines in them) but does solve the one I do (keeping a subset of files that would be tricky to glob properly).
In my experience 95% of those scripts are going to be discarded in a week, and bringing Python into it means I need to deal with `os.path` and `subprocess.run`. My rule of thumb: if it's not going to be version controlled then Bash is fine.
This uses regex to match files ending in .wav or .draft (which is what I interpreted you to want). Xargs then processes the file. You could use flags to have xargs pass the file names in a specific place in the command, which can even be a one liner shell call or some script.
So the "find <regex> - xarg <command>" pattern is almost fully generally applicable to any problem where you want to execute a oneliner on a number of files with regular names. (I think gnu find has no extended regex, which is just as well- thats not a "regular expression" at that point)
Find can even execute commands itself without using `xargs`:
find -maxdepth 1 -iregex '.\.\(wav\|draft\)' -exec echo "found file:" {} \;If for some reason you do need the "find | xargs" combo (maybe for concurrency), you can get it to work with "find -print0" and "xargs -0". Nulls can't be in filenames so a null-delimited list should work.
That being said, I would interpret "-exec printf '%s\0' {} +" as being a posix compliant way for find to output null delimited files. I say this since the docs for the octal escape for printf allows zero digits. However, most posix tools operate on "text" input "files", which are defined as not having null characters. Thus I don't think outputting nulls could be easily used in a posix complaint way. In practice, I would expect many posix implementations to also not handle nulls well because C uses null to mean end of string, so lots of C library calls for dealing with strings will not correctly deal with null characters.
Even worse, it is whitespace delimited (with its own rules for escaping with quotes and backslashes)
touch "a b" ls | xargs rm # this won't work, rm gets two parameters ls | xargs -i,, rm ,, # this will work
>[..] arguments in the standard input are separated by unquoted <blank> characters [..]
As for -i, it is documented to be the same as -I, which, among other things, makes it so that "unquoted blanks do not terminate input items; instead the separator is the newline character."
$ printf "one two three\nfour five\n" | xargs -n 1 echo
one
two
three
four
fiveI understand this caveat, but I never had a file with newline that I cared about. Everyone keeps repeating this gotcha but I literally don't care. When I do "ls | grep [.]png\$ | xargs -i,, rm ,," (yes, stupid example) there is 0% chance that a png file with a newline in the name found itself in my Downloads folder. Or my project's source code. Or my photo library. It just won't happen, and the bash oneliner only needs to run once. In my 20 years of using xargs I didn't have to use -0 even once.
P*~*wav*~*draft*
This looks a bit obscure due to lack of spaces, but it's simpler than it seems; the pattern is: [glob] ~ [exclude glob]
P* ~ *wav* ~ *draft*
The ~ being the negate operator, which can be added more than once.It's essentially the same as "grep -Ev" or "find -iregex".
It's a lot less typing than find, and also something you're likely to use interactively once you're used to it, so it feels very natural.
E.g. instead of `ls | grep -Ev 'wav|draft'`, you'd have to do something like
for filename in *; do
if grep -E 'wav|draft' >/dev/null <<< "$filename"
then : # ...
fi
done
Of course, it's more convoluted, but when you're writing scripts that might be used for a long time and by many people, it helps to know that it is possible to write robust things. Tools like shellcheck certainly help. find . ! -name . -prune \
-exec grep -qE 'wav|draft' {} \; \
-exec "${action}" \; ;
Edit: I missed the herestring in the original code, so the above is wrong as mentioned in the comments; if your find has regex, you can use it to save one grep: find . ! -name . -prune \
-regex '.*wav.*\|.*draft.*' \
-exec "${action}" \; ;
Otherwise you can call sh to printf the filename into a grep.However, the point of my post is that find can perform seek, filter and execute, and should be used for all three unless it is really impossible (which is unlikely).
Yes: bash is probably fine.
No: real programming language time.
Shellcheck's page on parsing ls links to the article the author is nitpicking on, but it also links to the answer to "what to do instead": use find(1), unless you really can't. https://mywiki.wooledge.org/BashFAQ/020
I've been using Linux since 1999 and i never came across a filename with newlines. On the other hand, pretty much all "ls parsing" i've done was on the command-line to pipe it to other stuff in files i was 100.1% sure would be fine.
The important point to get across is that pipes let us build bigger commands from the commands we already know. If needed, you can back up later to teach patterns like `find [...] -exec`, `find [...] -print0 | xargs -0 [...]`, `find [...] | while read -r file; do [...] done` and so on.
There are all kinds of prerequisites to creating files with unusual names. Those barriers tend to mean beginners won't run into file name processing edge cases for a while. The exception will be files they download from the Internet. But the complexity there will usually be quote and non-ASCII Unicode characters, not newlines or other control codes.
In teaching, the one filename complexity I would try to get ahead of, preventively, is spaces. There was a time, way back when, when newbies seemed to expect to stick with short, simple filenames. These days they the people I've helped tended to be used to using spaces in file names in Finder and Explorer for office or school work.
Not piping strings avoids this issue completely. Marcel’s ls produces a stream of File objects, which can be processed without worrying about whitespace, EOL, etc.
In general, this approach avoids parsing the output of any command. You always get a stream of Python values.
Somewhere, there has to be validation phases. Just because you have objects, doesn't mean they are well formed.
https://lexi-lambda.github.io/blog/2019/11/05/parse-don-t-va...
It turns out proper validation is way harder than parsing. There is a reason text based interfaces and formats are so pervasive.
How do you check if they are still valid files?
Files come and go. References to them go stale. Every user and tool deals with this. This isn't a "validation" issue.
--zero end each output line with NUL, not newlineUsing magic, I've renamed any files you have to remove control characters in the name and made it impossible to make any new ones. (You can thank me later.)
What can't you do now?
But for whatever reason, when it is suggested, you get many people chiming in that "filenames should be dumb bytes, anything allowed except / !"
I guess the issue is not what filenames should be, but what filenames are. In general, when interacting with files, you have to expect everything but `/` and nullbyte. Even if you forbid it on your machine, someone may mount a NFS drive and open you to the world of weird filenames. And you never know who uses your code.
And the unicode itself is weird anyway - for example you may have normalised and denormalised names which may be the same or different string depending on how you look at them[1]. And I hope you are not planning to restrict filenames to some anglocentric [a-z0-9_- ]*, because the world is much larger and you can't pretend unicode doesn't exist.
[1] https://eclecticlight.co/2017/04/06/apfs-is-currently-unusab... and many other cases
Now at least when some tool or pipeline blows up horribly, it'll be hilarious.
On the surface, it looks like I'd be giving up the decently sized ecosystem of Powershell libraries for a new ecosystem without much support?
I'm interested in knowing what Nushell does differently since I'm wanting to find a better shell.
Commands and flags are case-insensitive.
Who would have thought that little old Microsoft, purveyors of MSDOS CMD.EXE, would have leapfrogged Unix and come out with something so important and fundamental as a shell that was superior to all of Unix's "standard" sh/csh/bash/whatever shells in so many ways, all of which historically used to be and ridiculously still are touted by Unix Supremacists as one of its greatest strengths?
You see, Microsoft is willing to look at the flaws in their own software, and the virtues of their competitors' software, then admit that they made mistakes, and their competitors did something right, and finally fix their own shit, unlike so many fanatical monolinguistic Unix evangelists.
They did the exact same thing to Java and JavaScript, leaving Visual Basic and CMD.EXE behind in the dustbin of history -- just like Unix should leave bash behind -- resulting in great cross platform languages like C# and TypeScript.
Edit: that reinforces my point that taking so long to get there is a hell of a lot better than taking MUCH LONGER to NOT get there.
Maybe bash's legacy inertia is a problem, not a virtue. Is certainly isn't getting a JSON parser in the foreseeable future. The ironic point is that even Microsoft's power shell has much less legacy inertia, and therefore is so much better, in such a shorter amount of time.
> Isn't it ironic that Powershell from Microsoft is so much vastly superior than bash
I agree that powershell is now better than bash. But it took SO LONG to get there. Moreover, bash has had a 12 year head-start (ok, 30 if you count earlier unix shells). Bash has legacy inertia. Even though you can now supposedly run powershell in linux, I don't know anyone who does. Does anybody?That said, I think powershell is great for utility-knife uses on windows machines.
I do. I replaced all of the automation scripts on my rpi with pwsh scripts, and I'm not regretting it. Not having to deal with decades of cruft in argument parsing and string handling, learning little DSLs for every command, etc. is so worth it.
Basic features are still lacking from PowerShell that have been in UNIX shells since the very beginning: https://github.com/PowerShell/PowerShell/issues/3316
But hey, that's a fixable problem, right? No, because PowerShell is so suffused with arrogance about its superiority that anything, no matter how simple it was to do in a UNIX shell, has to be cross-examined, re-imagined, and bent over the wheel of PowerShell's superiority, before ultimately getting ignored or rejected anyway.
PowerShell is a language unto itself. It is not a replacement for bash/zsh/etc because nobody who knows the latter well can easily migrate to the former, and that's by design.
I won't install a new shell to generate a file list on my CI server. I won't install a new shell on remote machines. Ever.
These structured shells also require commands to be aware of them, either via some plugin that structures their raw I/O output or some convention. They solve _some_ command output structuring but not _all_ the general problem.
So, the answer is good. It promotes the idea that one should be careful when machine parsing output meant for humans.
Uh... that's on you? Why do you intentionally hinder yourself?
> These structured shells also require commands to be aware of them, either via some plugin that structures their raw I/O output or some convention. They solve _some_ command output structuring but not _all_ the general problem.
Okay. It doesn't solve literally every single problem, that is true. It's still miles ahead. And when interfacing with non-pwsh commands, you just fall back to text parsing/output.
Hinder myself? An ephemeral cloud machine would not keep my custom shell anyway. By having to install it _every single time I connect_ I just loose precious time.
I want to be familiar with tools that are _already_ installed everywhere.
The shell is supposed to be a bottom feeder, lowest common denominator, barely usable tool. That way, it can build soon and get stable real fast. That (unintentional) strategy placed it as a core infrastructural piece... everywhere.
Of course, there's scripting and using it on the terminal. But we're talking about scripting, right? Parsing ls and stuff. I want the fast, lean, simple `dash` to parse my fast, lean simple scripts. pwsh is fine for the terminal leather seats.
shopt -s failglobMost of the time it's fine to just suck in ls and split it on \n and iterate away, which I do a lot because it's just a nice and simple way forward when names are well-formed. Sometimes it's nicer to figure out a 'find at-place thing -exec do-the-stuff {} \;'. And sometimes one needs some other tool that scours the file system directly and doesn't choke on absolutely bizarre file names and gives a representation that doesn't explode in the subsequent context, whatever that may be, which is quite rare.
A more common issue than file names consisting of line breaks is unclean encodings, non-UTF-8 text that seeps in from lesser operating systems. Renaming makes the problem go away, so one should absolutely do that and then crude techniques are likely very viable again.
find ~/Music -iname 'p*' -not -iname '*age*' -not -iname '*etto*'
find ~/Music -iname 'p*' -not -iregex '.*\(age\|etto\).*'
find ~/Music -regextype posix-extended -iname 'p*' -not -iregex '.*(age|etto).*'
Not that I'm likely to ever use any of that in anger, but it's good to know if ever I do wind up needing it.And/or use their return codes to verify that something worked or didn't
Or hope your higher level programming language contains built-ins for file system manipulations
I'd love to see standard JSON output across these tools. I just don't see a realistic way to get that to happen in my lifetime.
Maybe a unified parsing layer is more realistic, like an open source command output to JSON framework that would automatically identify the command variant you're running based on its version and your shell settings, parse the output for you, and format it in a standard JSON schema? Even that would be a huge undertaking though.
There are a lot, LOT of command variants out there. It's one thing to tweak the output to make it parseable for your one-off script on your specific machine. Not so easy to make it reusable across the entire *nix world.
WRT format, I'd prefer csv.
#define _GNU_SOURCE
#include <dirent.h>
#include <fcntl.h>
#include <malloc.h>
#include <stdio.h>
#include <stdlib.h>
#include <string.h>
#include <unistd.h>
#define BUF_SIZE 32768
struct linux_dirent64 {
ino64_t d_ino; /* 64-bit inode number */
off64_t d_off; /* Not an offset; see getdents() */
unsigned short d_reclen; /* Size of this dirent */
unsigned char d_type; /* File type */
char d_name[]; /* Filename (null-terminated) */
};
int writeall(char *buf, size_t len) {
ssize_t wres = 0;
wres = write(1, buf, len);
if (wres == -1) {
perror("write");
return -1;
}
if (((size_t)wres) < len) {
return writeall(buf + wres, len - wres);
}
return 0;
}
int main(int argc, char **argv) {
if (argc != 2) {
return EXIT_FAILURE;
}
int fd = open(argv[1], O_DIRECTORY | O_RDONLY);
if (fd == -1) {
perror("open");
return EXIT_FAILURE;
}
void *buf = malloc(BUF_SIZE);
ssize_t res = 0;
do {
res = getdents64(fd, buf, BUF_SIZE);
if (res == -1) {
perror("getdents64");
return EXIT_FAILURE;
}
void *it = buf;
while (it < (buf + res)) {
struct linux_dirent64 *elem = it;
it += elem->d_reclen;
size_t len = strlen(elem->d_name);
if (writeall(elem->d_name, len + 1) == -1) {
return EXIT_FAILURE;
}
}
} while (res > 0);
return EXIT_SUCCESS;
}Your shell already provides a nice abstraction over calling readdir directly. A glob gives you a list, with no intermediate stage as a string that needs to be parsed. You can iterate directly over that list.
Every language provides either direct access to the C library, so that you can call readdir, or it provides some abstraction over it to make the process less annoying. In Common Lisp the function `directory` takes a pathname and returns a list of pathnames for the files in the named directory. In Rust there is the `std::fs::read_dir` that gives you an iterator that yields `io::Result<std::fs::DirEntry>`, allowing easy handling of io errors and also neatly avoiding an extra allocation. Raku has a function `dir` that returns a similar iterator, but with the added feature that it can match the names against a regex for you and only yield the matches. You can fill in more examples from your favorite languages if you want.
The getdents system call being used in the above program is the basis for implementing readdir.
It doesn't return a string, but rather a buffer of multiple directory entries.
The program isn't parsing a giant string; it is parsing out the directory entry structures, which are variable length and have a length field so the next one can be found.
The program writes each name including the null terminator, so that the output is suitable for utilities which understand that.
Do I really have to say this again?
It certainly would make writing Python scripts that need to interact with other programs easier. But Python doesn't desperately NEED to interact with so many other programs for such simple tasks like enumerating files or making http requests or parsing json, the way bash does.
If bash was ever actually going to get json parsing in reality, it should have done that two decades ago like all the other scripting languages, since JSON is 23 years old. So don't hold your breath.
https://news.ycombinator.com/item?id=40692698 (10 days ago, 83 comments)
You might say that people don't move or rename things while files are open, but they absolutely do, and it absolutely breaks things. Even something as simple as starting to copy a directory in Explorer to a different drive, and then moving it while the copy is ongoing, doesn't work. That's pathetic! There is no technical reason this should not be possible.
And who can forget the case where an Apple installer deleted people's hard disk contents when they had two drives, one with a space character, and another one whose name was the string before the first drive's space character?
Files and directories need to have a unique ID, and references to files need to be that ID, not their path, in almost all cases. MFS got that right in 1984, it's insane that we have failed to properly replicate this simple concept ever since, and actually gone backwards in systems like Mac OS X, which used to work correctly, and now no longer consistently do.
Explorer needs to support local drives, with a lot of filesystems, including possibly third-party ones, but also network drives, FTP, WebDAV, and a bunch of other niche things. Not all of them have IDs and might not be possible to be extended. The cost is massive, solving it everywhere is impossible, and the benefit seems negligible to me (even though I fairly recently managed to eject a disk image (vhdx) in the middle of copying files onto it…)
Similarly, things like config files would be identified by their name, not their path, because the directory containing configs was a directory the system knew about. As a result, no application needed to know the path to its own config files.
This meant there was no action that the system prevented you from doing to an open file, other than actually deleting that file. There was also no way for an installer to accidentally break your system because its code didn't take your drive, file, or directory names into account.
And, of course, there are file systems that don't use paths at all, like HashFS, a bunch of modern document management systems, or the Newton's Soup.
I get your point about interoperability with existing file systems, but I think it's perfectly acceptable to offer better solutions where possible, and fall back to paths for situations where that is not possible.
latest="$(ls -1 $pattern | sort --reverse --version-sort | head -1)"
Anyone got a better solution? latest=$(printf '%s\0' <glob> | sort -zrV | head -zn1)
or with long args: latest=$(printf '%s\0' <glob> | sort --zero-terminated --reverse --version-sort | head --zero-terminated --lines 1EDIT: The linux manpages I read were from die.net, which it looks like were from 2010, guess I'll have to avoid them in the future. I checked FreeBSD, OpenBSD, and Mac man page to make sure, and unfortunately none of them support the -z flag yet.
In that case, how about:
IFS='' read -d '' latest < <(find $pattern -prune -print0 | sort -z --reverse --version-sort)Anyway check it, you might find you have files with spaces after all. For me it's:
* /boot/System Volume Information
* /proc/irq/126/PCIe PME
* /sys/bus/platform/drivers/int3403 thermal
* /etc/NetworkManager/system-connections/Hotel Xxx.nmconnection
* /home/xxx/.cache/chromium/Default/Code Cache
* and many others.Do you really think that, say, all music streaming services are storing their songs with names allowing Unicode HANGUL fillers and control characters allowing to modify the direction of characters?
Or... Maybe just maybe that Unicode characters belong to metadata and that a strict rule of "only visible ASCII chars are allowed and nothing else or you're fired" does make sense.
I'm not saying you always have control on every single filename you'll ever encounter. But when you've got power over that and can enforce saner rules, sometimes it's a good idea to use it.
You'll thank me later.