Hacking ls -l
lemis.com
lemis.com
Rather, this is a pretty interesting look into what it actually entails to make what ought to be a very simple and straightforward change.
It turns out that these simple changes are hard! Not just in identifying the piece of code to modify, but that man pages are often incomplete or unclear. It also illustrates the complexities behind making software portable - in this case, using the nation-neutral place separator. It also reminds us that solving what is on the surface a simple problem lets one uncover all sorts of interesting and messy details underneath - including more problems to solve!
These are steps that he'd have to take no matter what the code or feature. This article is not "complexity for complexity's sake", it's illustrating the complexity of making changes to any piece of code - and that it is surprisingly difficult for something that one would think is very easy!
But I still wonder if this is easier than ls -l | sed -e :a -e 's/\(.*[0-9]\)\([0-9]\{3\}\)/\1,\2/;ta' ??
The investigations would be interesting it they were more complete, i.e. if the actual result was a change in the locale which could be appliable to other tools printing numbers besides ls (in the author TODO).
I mean, will it work with bc?
At the moment it's not better than an shell script alias giving the output to sed, but it is more complex - you have to recompile a binary for every OS you use.
drwxr-xr-x 1,653 dalke admin 56,202 Mar 5 2,012 pubchem
-rw-r--r-- 1 dalke staff 59,252 Nov 16 2,011 pubchem_10,000.fps.0.9.cluster
I don't want to see the year written "2,012", and the file name is 'pubchem_10000" not "pubchem_10,000".EDIT: seeing how it has been downvoted, IMHO hacking is all about time and effectiveness. if you believe fixing ls print formats is such a crucial problem that it requires more than seconds of your time, we have different values.
Feel free to support the argument by showing your skills and improving the example.
It is exactly about showing how hard it is to get a trivial change right in every detail. Your 10 second hack is what is wrong with 10 second hacks in general and even with most real solutions that are not carefully thought out.
It's not only that the devil is in the details it is all details. And you need to get all of them right, not just the current subset of the problem that you happen to be working on.
Someone with command line experience that is mostly a user of system utilities would think passing the output of ls through a filter is the way to go whereas someone with C experience that already has insight in how Unix is architected would do something more along the lines of what the author is doing. That's pretty much how unix got built to begin with, and probably both parties would qualify their solution as a 'hack'.
It's all perspective.
(Note that this RFC comes from the days when 'hacker' was becoming a widespread term for someone who breaks into computer systems; the RFC attempts to distinguish between a 'hacker' and a 'cracker.')
correction: sorry it will only fix the modifications to the filename - the year is still broken.
ls -l | perl -pe 'while(s/^((\S+\s+){4})(\d+)(\d{3})([^\d].*)?$/$1$3,$4$5/){}' -rw-r--r-- 1 root wheel 16,596,907,252 24 Dec 2009 boskoop.disk0.bz2
-rw-r--r-- 1 grog wheel 4,173,914,809 20 Jul 2006 boskopp.tar.gz
With your perl one-liner on my directory I get mis-aligned columns: -rw-r--r-- 1 dalke staff 3,236,397,056 Sep 13 2011 pubchem.fps
-rw-r--r-- 1 dalke staff 712,181,172 Sep 13 2011 pubchem.fps.gz
That's ugly. It should be: -rw-r--r-- 1 dalke staff 3,236,397,056 Sep 13 2011 pubchem.fps
-rw-r--r-- 1 dalke staff 712,181,172 Sep 13 2011 pubchem.fps.gzhttps://github.com/samsonjs/bin/blob/master/ls-comma
It's a disgusting hack and it works very well. I didn't find anything useful in the man page so I wrote this instead. Looks like there is something in the man page but I think my hack was probably faster.
This is why you try to re-use work when possible, rather than endlessly reinventing things, because while sure, adding a comma to the printf string is easy enough, your assumptions (English locale, compiler not trying to be clever) are going to quickly become visible as things fall apart because your assumptions aren't in line with the system's assumptions.
What this story really demonstrates is that without a clear understanding of how a system is designed and the basic assumptions it makes, just "hacking on the code" is just as likely to break things as it is to fix them.
It sounds like the author found a bunch of bugs in the process of making a simple code change. That happens pretty frequently, and doesn't mean that the author should let the priests of the cathedral deal with this UNIX thing that is too complicated for the laity to hack on. It just means there is no priesthood.
This is a great shame. I like OpenBSD's approach to man pages - incorrect documentation is a bug and can be as severe as a bug in code; correct documentation is important.
Fixing up man pages is something that non-technical volunteers could help with, except when it's hard to grok what the code actually does vs what it should do.
Another problem with that is that most tools used in this process are made with technical users in mind. A lot of people can expand documentation, but sending manpages patches in a bug tracker is a technical step.
Or they could be expert translators.
So it's a shame that man pages are not as good as they could be.
There are other forms of documentation, but it'd be nice if man pages were the best the could possibly be.
alias ls="BLOCK_SIZE=\'1 ls --color=auto"
The above is a bit hacky and not very UNIXy as it's
lumping more logic into ls, rather than splitting out
into functional units.Number formatting being a very common requirement, I've proposed a design for a new numfmt GNU coreutil
http://lists.gnu.org/archive/html/coreutils/2012-02/msg00085...
which would be used like:
ls -l | numfmt --field=5 --format=%'dExcept for someone with an ancient hard disk who thinks in blocks instead of (mega, giga, etc...)bytes, who ever needs or wants that?
alias ls="ls --block-size=\'1 --color=auto"Some people improve the area they travel through, others leave debris, and many are noops who make no difference to those who come after. If there's not enough entropy fighters like Mr. Lehey working a system, it turns to kipple.
By the way, by looking at http://www.lemis.com/grog/index.php you can see that the author uses FreeBSD, just in case you were wondering about /usr/src
Alternatively, when you cleverly figure out how to work around the warning, like the author does, you now prevent that rule from triggering even when it's right. Clearly a better unit test is needed.
It is super scary that the compiler appears to be using a different constant from printf for its format checker, that shows it probably isn't using a pattern supplied by printf.
(Or they did that and it's just a bug somewhere.)
Consider that the compiler generating that warning knows only Standard C, and in fact you could be pairing it with any C library, including those that are strictly conforming and don't support the ' extension.
(In the example, he's compiling ls with the -std=gnu99 flag, which means he's targeting GNU, not C99 or C90 or C11 or POSIX or any other standard.)
FWIW, GCC 4.7.2 parses the format string without warnings.
Or just read the docs (it should be "%'*jd "). Then no warnings. (IIRC ' is in C99 and -std=gnu99 targets c99 + gnu extensions.)
The same story with the rest. Two ways of doing things — learn & think and just do it right or twiddle until it seems to likely maybe work (possibly). The article is about the latter. Plus "blame the compiler".
http://joeyh.name/~joey/blog/entry/ls:_the_missing_options/
(Well, actually, I never got around to writing -z, but it's clear what it should do, and any ls hackers are encouraged to finish that up.)
I run into the situation more often with 'du' (when trying to find which subdirectory tree has excess junk in it), to the point that, while 'du -h' is human readable, it's not particularly sortable so:
du -hs $( du -s * | sort -k1nr,1 -k2 | head )
.... which will return the human-readable output, based on numerically sorting the full numeric output. Eyeball comparisons are easier as you're aware that results are already sorted by size.du -h | sort -h
And get properly size-sorted output.
http://lists.freebsd.org/pipermail/freebsd-questions/2012-Se...
A lot of times I catch myself in the mindset of taking a step back and saying "here are the set of tools I have at hand to accomplish a task" without realizing that I should simultaneously be taking a step "in"--so to speak--and acknowledging that the tools I have to work with are not immutable tools cast of iron; they are malleable and can be re-tooled to suit my purposes.. and that sometimes going that route can be the simplest--and in fact "best"--solution.
(that said, I do see the utility, since it gives a more obvious visual queue as to the order of size differences... but if you're doing anything with the sizes programatically, you have to remove the commas afterwards... Short version: if you're going to do this, make it a unique flag, or a new flag modifier to the -l flag... don't overload the -l flag without recourse...)
Even better if the _ separator used by programming languages were a supported locale LC=C_FOR_HUMANS :-)
$ LC_ALL=en_US.UTF8 ls -og --block-size="'1" .
-rw-------. 1 5,145,416 Oct 5 16:44 A
-rw-------. 1 5,137,692 Oct 4 14:37 B
-rw-------. 1 5,147,168 Oct 8 07:52 C
This feature is documented in the "Block size" section of the coreutils manual: i.e., you can type this to see it: info coreutils 'block size'If he doesn't send the changes off to upstream, and make a case good enough for them to be approved, then all this dooms him to maintaining his fork on all the platforms where he wants it until he gets sick of it or convinces someone else to do it for him.
However, from a computer, I (and I'm certain I'm not the only one) actually expect it to output the US notation.
I'd much rather have a computer always output the same format (and that happens to be the US format), than try to be smart with locales, when the end result is that some things will do this, others that. Makes stuff harder to use, and when programming, harder to parse.
I've once had to touch Excel on a Windows machine configured for a non-US language, and it refused to import a CSV file that had commas, even though CSV means comma separated values. It required semicolons due to the locale settings of Windows. This stuff should not happen. A CSV is meant for computers, and to be interchangeable, not to use different types of commas and refuse to work with other types depending on user locale settings...
Of course, when publishing or printing, that's a whole different matter, and there it better get the locale of your country perfect. But this here was about output in the console, which is often meant as input of other scripts etc...
Makes you wonder whether anyone ever considered the problem before...
This way, I quickly run lsd to only look for directories.
Alternatively, do you know about ls -lhrS? It will print size in human formats and reverse sort the files by size - ie the bigger will be at the end of the list
I usually consider the portability of the solution. I have linux i386 and x64 machines, my arm n900, an osx latop, etc.
Recompiling ls (or, heavens forbid, cross compiling!) for each machine may be a bright idea.
Adding a line to your profile that will take advantage of the existing tools like sed is closer to hacking in my definition, because it tries to think about the bigger problem - but still I wouldn't dare calling the following "hacking":
echo "alias lll=\"ls -l | sed -e :a -e 's/\(.*[0-9]\)\([0-9]\{3\}\)/\1,\2/;ta'\"">> ~/.bashrc
As I pointed out elsewhere, your sed code doesn't work because it changes too many numbers - including filenames and dates - in the output. Also, if the exercise is to understand localization then your sed code isn't appropriate because it hard-codes "," when some locales use a "." as the thousands separator.
And for completeness, your alias can't then mix and match other flags, like "ls -lRt". It's a single command with strange side-effects if you use it incorrectly:
% lll -art
sed: illegal option -- r
usage: sed script [-Ealn] [-i extension] [file ...]
sed [-Ealn] [-i extension] [-e script] ... [-f script_file] ... [file ...]Instead of complaining about an obvious flaw in the example, maybe you could be more constructive and fix it.
Hints:
- if you want to pass other flags, make it a function and use $@
- if you want to respect the filenames and dates, either fix the regex or write it in perl.
Shouldn't take you long - back from 1999 in perl FAQ:
http://www.perlmonks.org/?node_id=653
But now it's qualifies as hacking. I guess that's due to inflation.
This is so not hacker news.
Ask a perl golfer to make you a one-liner you can copy-paste in your profile if you absolutely need some working code.
My constructive criticism is that your approach is wrong, should not be done, and cannot be easily fixed. One should never attempt to process the general output of ls. It's doable - I lived through the years of processing the "list"/"ls" output from random ftp servers - but it's nasty. Sure, you can add '$@' but then you have to worry about, say, "-i", which shows the inode number as a new leading column or "-n" which shows user/group ids instead of names. How does your alias/perl script figure out which column is the one which needs the commas?
You'll either end up with a very fragile system (producing the erroneous output as your 1-liner does) or you'll end up trying to understand most of the ls command-line arguments and/or heuristics to guess based on the output. The well-known "BUGS" section of the Unix man page says "To maintain backward compatibility, the relationships between the many options are quite complex." You're in for a long slog if you go this route.
Yes, if you want a one-off solution for a specific set of outputs then your approach would work. That would be also be boring and trivial. The linked-to article, on the other hand, was interesting.
To me, this article gets to the crux of what I find particularly delightful about hacking (tinkering?): unraveling layers of complexity underneath. I feel like I have a little better understanding of what's happening when I punch in `ls`, and I think that particular delight and knowledge is the kind of thing that appeals to tinkerers (hackers? :p) like myself. So - in my view - entirely appropriate for this crowd!
But you have a fair point; there's nothing particularly out-of-the-ordinary of this code or process, and in that sense, isn't newsworthy to hackers.
(I did not downvote you incidentally; I think it's interesting to get a sense of peoples' different thresholds for what constitutes "hacker." I play the saxophone, and an instructor once told me that people always came up to him and said "I want to be a musician. How can I do that?" Well it turns out that the moment you play "hot cross buns" on your instrument, you are indeed a musician. Perhaps not a skilled one, but you have in fact made music. I think of hacking in a similar way, and freely admit that it is a loose use of the word!)
But, just like you, I consider that not newsworthy to hackers, yet at the moment it is the #1 item on HN and it kinda makes me sad especially because of the threshold - the idea that some people do consider that hacking - here of all places - is chilling :-/
Worse - #2 item is "more people should write". I beg to differ- more people should code, so that fixing a printf and recompiling ls wouldn't be newsworthy.
I'm sorry if it was interpreted as being rude- it was not the point - I just wanted to present alternative approaches to the problem, because reconsidering the problem is sometimes the right thing to do, especially when it escalate quickly in complexity.
As edw519 has said, a one line code change takes six days to implement.