Modernizing “less”
garrett.damore.org
garrett.damore.org
I wonder what other nasty things would appear after a serious code review of the GNU and BSD core code bases.
I do agree that the code has been read probably hundreds of times and executed millions of time, but I doubt there have been many formal attempts at improvement and the overall methodoloy resembles brute forcing to me.
I could be wrong.
less is an exception to the rule, as it's a small, standalone program and not as critical. Its initial development start seems to predate the official announcement of the GNU Project. I don't think most people really expend much thought on something like a pager.
EDIT: Brute forcing is also a very pragmatic approach. Ken Thompson uttered his famous adage for a reason, even if it was somewhat humorous.
Why? Because that was easier for the programmer to do.
From a user's perspective, it seems pathetic.
(Yes, I know that there's sed and awk, but why force the user to switch tools so early?)
Do you wish your screwdriver had a pliers attachment?
Are you familiar with the Unix Philosophy?
IMO making "cut" be better at its one job is more unix-y than using the more complex tools to do conceptually simple things.
</wikipedia>
&pattern
Which will: Display only lines which match the pattern; lines which do not match the pattern are not displayed. If pattern is empty (if you type & immediately followed by ENTER), any filtering is turned off, and all lines are displayed.I'd prefer this followed a syntax closer to mutt's filters, and that the patterns were editable (e.g., typing '&' during a filter would show the currently extant filter for modification), but it's handy.
If you type '&' then the up arrow, you'll get the previously entered pattern(s), which you can then edit.
Which is why I share tips like this -- it's almost always a win.
This is the kind of thing that always astonishes me to see in a codebase: why reinvent something rather than just finding and including a compatibility implementation? Just grab an appropriate getopt.c and compile it in if the platform doesn't have one, then let the rest of the code pretend every platform has one. (Preferably an implementation of getopt_long; a quick search turned up some licensed under 3-clause BSD.)
My guess is nobody bothered to replace these parts, since options were only added gradually; if at all. Coincidentally i'm in the same predicament with a tool a coworker of mine (initially) wrote, convoluted option parsing to say the least; but I'm too busy with fixing other parts or adding proper functionality to it than to replace it.
Sometimes, 'less is more'.
"It's right in the manpage, actually". "No." "Yes, I'll send it to you."
And so I did, with the subject line "man less".
She sat at a desk right in front of mine, and I detected a somewhat painful silence as the email arrived. And realized I'd just inadvertently commented on her social life (confirmed through later conversations).
> less clunky and avoids the duplication that getopt_long results in.
Duplication because of the ugly flag/val logic, which in practice is typically passed as either NULL, 's' (long option for short option 's') or NULL, OPTION_FOO (long option with no short option, OPTION_FOO > 255)?
Yeah, that does seem silly. I think they did that to simplify the setting of boolean parameters, by passing &some_flag, 1, but that seems woefully insufficient when you need to handle arguments. I'd love to have a C library as capable as Python's argparse, instead.
And yes, that's pretty much what I'm getting at. The issue really stems down to the fact that getopt_long is just bolted on to getopt, so there's still the string short-opt syntax like "sf:t", and then you duplicate the options in the array of long opts with the chars or some random number > 255 otherwise. The string is really the most annoying part because there's no good way to generate it via the C preprocessor that I could come-up with, leading to some duplication between the long-opts and short-opts. It also can't easily generate help text for you from your arguments, which bugged me. My argument parser basically works off of a single xmacro header that holds all the argument information for getopt (Which gets organized via a few macros into an enum and an array). It's dead simple to add new arguments and there's no duplication or separate strings or etc. that you have to update at the same time besides adding code to handle that argument.
Personally, I wrote my argument parser specifically because I couldn't find any that I was happy with after looking around. They were either clunky to use (getopt_long), or were full libraries and seemed like it would be a hassle to get it integrated into my code. I'd love to see a 'standard' argument parser that works with long options well and doesn't just feel like an afterthought like it does with getopt_long, but I think the chance to make such a thing has been missed. Personally, if I would ever use such a thing then it needs to be a single or just a few .c and .h files that I can stick directly into my program. I wouldn't want to have to bet on the distribution having it or not, and I'm not going to add a separate dependency just for argument parsing.
I don't have it as a separate project specifically for the argument parser which is why I didn't link to it. But the meat of the parser is in ./common/arg_parse.c, with the header for it in ./include/common/arg_parse.h. If you want to throw it into your own project you'll want to take a quick look through the ./arg_parse.c and modify it to suit your project (The help text specifically is for my program, so you'll want to rewrite the text in that part).
You can see an 'example' usage in the same repo. The files listed below parse the arguments into a single struct with few flags inside, and also look for filenames in the arguments to load into the emulator:
./cmips/args.c
./cmips/args.h
./cmips/args.x
./cmips/args.x is a xmacro header. It mostly just contains the contents of the struct arg array for this program, but it's also used to create enum entries which index the struct arg array. The 'parse_args' function is fairly similar to what you'd do with getopt, just with the 'arg_parser' function instead.The argument parser code could probably be improved. Just looking back at it, it's got a bit to much logic going on in that single function, it could probably be split up a bit. I'm gonna be working on this project again pretty soon I think, so I might fix up this parser along with it.
However, the getopt.h-file in the actual gnu libc-package (libc6-dev) is lgpl.
Interestingly, the getopt.c/.h in gnulib (gnu portability library) appear to both be gpl, not lgpl.
I've seen a few getopt.c implementations that are MIT I believe.
Personally I'm guessing that those getopt implementations are probably all different (Just because it hardly takes any time to write one). AFAIK gnulib is GPL itself, so the getopt inside was just licensed GPL too, same thing for util-linux. libc is LGPL though, so that getopt.c was licensed as LGPL. It is kinda curious though, I'm personally just surprised that there are so many implementations of the same thing.
I stopped using it, like 4 years ago, when someone did told me on IRC that it was unmaintained and did have some bugs.
Note that I did never notice any 'bug' in my time as 'most' user, even if they may exist, and indeed, it's still available at least in all Debian versions.
It's curious that distributions are able to "maintain" a package (maybe even with custom patches) which is not maintained or updated upstream, for years.
It's encouraging to see actions like the one performed by this IllumOS developer. If a program is opensource, we can fork, improve and share. Or as users, we can take a look at the code when making choices, lots of people forget this in favor of search engine recommendations.
https://bugs.debian.org/cgi-bin/pkgreport.cgi?which=maint&da...
Related to a discussion I've been having recently on systemd.
Being a GNU package just means you agree to behave in a GNUly way. It's a very informal and easily-granted qualification.
Couldn't be that hard to do a boyer moore for non RE substrings.
http://lists.freebsd.org/pipermail/freebsd-current/2010-Augu...
edit: Eg, with warm cache:
:~/tmp/riak/riak-2.0.0pre5/deps $ time (find . -type f -exec cat '{}' \; |wc -l)
2765699
real 0m5.021s
user 0m0.144s
sys 0m0.792s
:~/tmp/riak/riak-2.0.0pre5/deps $ time (find . -type f -exec cat '{}' \; |grep -E 'Some pattern' -v -c)
2765700
real 0m5.133s
user 0m0.264s
sys 0m0.852s
:~/tmp/riak/riak-2.0.0pre5/deps $ time (find . -type f -exec cat '{}' \; |grep -E 'Some..pattern' -v -c)
2765700
real 0m5.144s
user 0m0.400s
sys 0m0.768s
# "%% " used for leading comment lines in some of this code:
:~/tmp/riak/riak-2.0.0pre5/deps $ time (find . -type f -exec cat '{}' \; |grep -E '^%% ' -c)
27535
real 0m5.597s
user 0m0.520s
sys 0m0.788s
:~/tmp/riak/riak-2.0.0pre5/deps $ du -hcs .
405M .
405M total
:~/tmp/riak/riak-2.0.0pre5/deps $ time (find . -type f -exec cat '{}' \; |ag '^%% ' >/dev/null)
real 0m5.735s
user 0m1.480s
sys 0m0.876s
#actually find/cat is pretty slow -- I guess both GNU grep and ag
#use nmap to good effect:
$ time rgrep '^%% ' . > /dev/null
real 0m0.539s
user 0m0.404s
sys 0m0.128s
:~/tmp/riak/riak-2.0.0pre5/deps $ time ag '^%% ' . |wc -l
27500
real 0m0.252s
user 0m0.284s
sys 0m0.068s
:~/tmp/riak/riak-2.0.0pre5/deps $ time rgrep -E '^%% ' . |wc -l
27553
real 0m0.535s
user 0m0.396s
sys 0m0.140s
Note that grep clearly goes looking in more files here (more mathcing
lines). Still, I guess ag is indeed faster than grep in some cases (even
if it might not be apples to apples depending how you count -- of course
the whole point of ag is to help search just the right files). :~/tmp/riak/riak-2.0.0pre5/deps $ time rgrep -E 'Some pattern' . |wc -l
0
real 0m0.266s
user 0m0.128s
sys 0m0.132s
:~/tmp/riak/riak-2.0.0pre5/deps $ time rgrep -E 'Some..pattern' . |wc -l
0
real 0m0.338s
user 0m0.212s
sys 0m0.120s
:~/tmp/riak/riak-2.0.0pre5/deps $ time ag 'Some..pattern' . |wc -l
0
real 0m0.111s
user 0m0.100s
sys 0m0.076s
I guess ag is indeed faster, even if it might not be due to fixed string
search...[edit2: For those wondering that's an (old) ssd, on an old machine -- but with ~4G ram the working set should fit, as soon as some of my open tabs in ff are paged to disk...]
In the course of checking out ag (again) I also learned about gnu id-utils[2].
[1] http://geoff.greer.fm/ag/ [2} http://www.delorie.com/gnu/docs/id-utils/id-utils_1.html
I think most of the slowdown you're seeing with "find -exec | cat" is forking at least two processes (ag and cat) for each file. Also, each process has to be run sequentially (to prevent garbled output), which makes use of only one CPU core most of the time. I've tried to keep ag's startup time fast so that classic find-style commands still run quickly. (This is why ag doesn't support a ~/.agrc or similar.)
Just FYI, you can use ag --stat to see how many files/bytes were searched, how long it took, etc. I think I'll add some stats about obeying ignore rules, since some of those can be outright pathological in terms of runtime cost. In many cases, ag spends more time figuring out what to search than actually searching.
Many thanks for not just writing and sharing ag as free software, but for the nice articles describing the design and optimizations!
At least this brief benchmarking run convinced me that I should probably try to integrate ag in my work flow :-)
Since it was only a small hack to scratch the itch I was having at the time, I never really completed that project. For example, backwards line counting is not sped up, which can sometimes be noticeable.
If you feel like working on less-speedup issues, feel free to drop me a line.
When you were experiencing long wait times, did you turn off line numbering?
less -n logfileOverall I think less is one of those tools where it's really valuable to spend 10 minutes a day in the man page for a week, which should be enough to learn essentially all of its functionality.
Markers are also very useful, particularly paired with the functionality to pipe data to another file or shell command. E.g. to extract the instance of a server error plus some lines for context from an otherwise unwieldy log file. :) I use markers rarely enough that I invariable need to reread the man-/help-page, but being aware of the functionality is half the battle. :)
Another tip: within less, press -S to toggle line wrap. (Works for most other command line options, too.)
Isn't that just grep -C?
Honestly I'd take a stab at myself if I had the time. Maybe I should start a kickstarter or something like that.
You should never invoke "more" directly in a script anyway, but rather observe the PAGER environment variable, and fall back on a plain "more" only if that isn't set. (Speaking of which, PAGER isn't described in POSIX, oops!)
If the user wants the pager to exit when the last line is reached, the user can specify the necessary option in PAGER, if their pager supports it. PAGER just has to be properly expanded: treated as a command, not a command name.
I've looked at it before, and I agree less is a mess.
Can anyone point to work that starts with the combination of the two following propositions:
1. User interface elements invented since the 1970s are a pretty neat thing.
2. Text-based shells and the command line are also a pretty neat thing.
I'm not crazy about every last aspect of his design there - but it's a start.
http://en.wikipedia.org/wiki/Acme_(text_editor)
I have never tried to use it, and I'm not sure it's mixing the ingredients in the way you are thinking.
And less has built-in tailing which you can start and stop at any time. That's its killer feature for me.
On the other hand, vim/view can have some nice syntax highlighting for syslog format log files. I haven't found that enough to switch though.
# ubuntu 14.04
alias less='/usr/share/vim/vim74/macros/less.sh'
# ubuntu 12.04: vim73, if I recall right
alias less='/usr/share/vim/vim73/macros/less.sh'
This behaves like `less` in many ways, and uses your syntax highlighting from vim. My only complaint is that some things with escape codes for colors are not flattened, but instead you see the escape codes. (Diffs seem to work fine, at least.) It also appears to read the whole thing into vim, which is likely not what you want for large files.I also use the 'vimcat' script from the vimpager project [0] as well. (I'm not sure why I haven't just used the whole thing. I must not have realized there was more when I first grabbed it.)
I've never been using a program and wished for better functionality. Even when this was all new to me, it was never a problem figuring these out, using them, and I was always satisfied with them..
So...here's the question. I don't think these are broken, so what are you fixing?
>So, now it:
>Speaks terminfo (X/Open Curses) instead of ancient BSD termcap
>Uses glob(3C) instead of a hack involving the shell and a helper program (lessecho, which I've removed from my tree.)
>Functions properly as /usr/bin/more, both with and without -e (even on broken xterms)
>Is fully ANSI C (or ISO C, if you prefer)
>Passes illumos' cstyle code style checks
>Is lint(1) clean
If that were improved I'd never have to use tail -f ever again.
Someday, I'll look into it.
First, it's not just about something not working. It's about creating tools that are extensible and understandable and hackable. Open Source is not just about "working", it's about being modifiable by the end user. All this cruft (a mess of 200 obsolete architectures, dead code and deprecated library support that nobody used since 1988) works against that goal.
Second, there are things that would be essential for some people, like international users (e.g proper multibyte support) that cannot be added due to dependancy of some custom methods of handling encodings. That's not some wishy washy magical unicorn feature request, it's essential for the main operation of what less does for those that have to deal with these encodings.
Third, there's nothing wrong in taking pride and crafting finely your tools. UNIX is supposed to be made of things that "do one thing and do it well". Less having its own utf-8 support breaks this division of responsibility. We have libaries for that. Same for getopts vs it's custom options parsing.
I know many people that never got a BSOD on Windows. Does this mean that there is nothing that could be fixed about Windows? Or is it possible that different people have different experiences?
This is what the post talks about. But in short, Iluminos (An operating system derived from OpenSolaris) needed a posix compliant pager(/usr/bin/more), ported the less program to their OS and in the process found many issues which they cleaned up and fixed.
-----
Downvotes for clarifying an ambiguity, really?
The lack of any context for HN posts (also reddit link shares) is a significant disadvantage of both sites. I've always been partial to Slashdot's link summaries, and wish that style were more widely used. See also Jakob Nielsen and microcontent.