Unix V5, OpenBSD, Plan 9, FreeBSD, and GNU Coreutils Implementations of Echo.c
gist.github.com
gist.github.com
Also, the comment…
/* This utility may NOT do getopt(3) option parsing. */
… which appears in two of them certainly cries out for an explanation. Was it true in NetBSD? Is it still true? Will people call us names if we do?Update: bash does not handle errors. zsh might, it checks the result of fclose() or if stdout, fflush(). I'm not sure that is sufficient after ignoring all the fwrite() and fputc() errors. Plus it's in an 850 line function and my monitor is only 20" tall, so I'm not following all of it.
Update2: asveikau is correct. The only ones to attempt error handling miss short reads. FreeBSD also misses EAGAIN/EINTR handling. Plan9 doesn't need that, but its man page explicitly states that short writes should be considered errors by the caller.
I guess it's official. The unix read/write system calls are too complicated to be used by experts in the operating system. (I've rather thought that for some time now.)
Both only check for negative return codes. If stdout were piped/redirected to something where write(2) may return a short byte count (let's say a socket) neither of them will detect it.
> /* This utility may NOT do getopt(3) option parsing. */
I noticed this too. My guess, totally not backed up by anything but by instinct, is getopt(3) tried to do too much which doesn't make sense for echo. For example if it tried to parse "echo -foobar" as anything other than "you spit out the string '-foobar'" I could see it being problematic.
Arguably GNU echo is the only one to actually get this correct. Which is kinda amusing given how the comments on github are all about how ugly and bloated it is and how elegant the Plan9 and UnixV ones are.
The talk is worth a watch if you care about these things. It kind of re-enforces your point about read/write being to complicated to use correctly. Meyering found these kinds of bugs in lots of other programs like perl, python, rsync and emacs.
[1] http://www.irill.org/events/ghm-gnu-hackers-meeting/videos/j... (slides: http://www.gnu.org/ghm/2011/paris/slides/jim-meyering-goodby... )
Anyways, what's the issue with the System V version, other than not handling newlines?
Also, what is the purpose with the Plan 9 loading the thing into a buffer first, other than perhaps to get the benefit of an all-or-nothing write? I suspect weirdness with the Plan 9 file handling but am not certain.
Complexity is inherent to any task in some degree. That complexity must be expressed in full in any program which performs that task correctly, whether in terms of language/library constructions or the ways in which those are combined. When the tools with which you implement a task (here, C and the core Unix API) are structured substantially differently than the high-level concepts of which that task is composed, there is additional complexity inherent to mapping those concepts.
In light of that, I don't understand what you mean by, "complexity itself is a sin," unless you're referring specifically to the strictly unnecessary complexity introduced by a sub-optimal mapping of high-level semantics to core constructs. I don't think that's the kind of complexity we're seeing here. I think we're seeing the complexity inherent to mapping shell-level semantics to C semantics and the semantics of the Unix API, specifically error-handling semantics.
I agree with jws that what this code demonstrates is that the core Unix API is difficult to use for writing precisely correct high-level programs, even ones as simple as echo. I feel that this is exactly why it's important that core high-level programs like echo encapsulate as much of the complexity of mapping to C as possible (that is, in the style of GNU rather than SysV): so that one can easily write reliable programs with high-level tools like the shell without worrying about that complexity.
The Plan 9 and the Echo:
http://9fans.net/archive/2001/09/54
(for the history buffs: yes, that is dennis ritchie posting in that thread)
It references _The Unix Programming Environment_, Kernighan & Pike (1984), pp. 77-79, "A digression on echo" which I found in PDF and is an immensely enjoyable and enlightening read :)
[1] http://books.cat-v.org/computer-science/unix-programming-env...
This simplifies quite a bit of code that reads and writes commands into ctl files.
In other words, writev on plan9 isn't a system call -- it does the same copying that echo.c does. Also, echo.c probably predates writev.c.
http://cvsweb.openbsd.org/cgi-bin/cvsweb/src/bin/echo/echo.c...
http://en.wikipedia.org/wiki/Echo_(command)
I wonder if, like a "Hello world" program, some of these simpler utilities might not even meet the minimum level of creativity to be eligible for copyright, so complexity was added - true.c and false.c are the other two examples I can think of; the "true" utility is basically generated by the default code template provided by many IDEs.
There's also a fun story floating around called "The UNIX and the Echo" about what it should do without any arguments, or whether or not it should accept options or escape sequences:
http://stackoverflow.com/questions/3290683/bloated-echo-comm...
The official POSIX definition leaves it implementation-defined, with XSI forbidding options and requiring escape sequences:
http://pubs.opengroup.org/onlinepubs/9699919799/utilities/ec...
It's amazing how much variety can be present for functionality that seems so simple at first glance.
/* echo.c, derived from code echo.c in Bash.ALso surprised how many library's the FreeBSD and GNU flavours want to pull in.
[1]: https://github.com/freebsd/freebsd/blob/master/bin/echo/echo...
[2]: https://github.com/freebsd/freebsd/commit/cbf5708f4336b824b8...
Using 1000 arguments does not seem at all a typical use case for echo.
My guess is that somone had a weird use case for echo where the old version was slow. He benchmarked, profiled, made some changes, and submitted a patch. "The source code is too long," is a crappy argument, so might as well use the faster, longer code.
https://github.com/uutils/coreutils/blob/master/src/echo/ech...
POSIX sed supports exactly 3 options (-n, -e and -f), the latest FreeBSD sed supports 10 (adding -E, -a, -I, -i, -l, -r and -u — the last two being GNUism compatibility options not necessarily available on older versions, they are not on my OSX machine) and GNU sed supports 9 short and an additional 4 long options. And that's not counting the extensions to the sed command set.
Definitely less convenient, but would make building cross-platform scripts much easier.
/* System V machines already have a /bin/sh with a v9 behavior.
Use the identical behavior for these machines so that the
existing system shell scripts won't barf. */
(BSD systems supported suppressing newline with -n; UNIX System V systems supported \ escape sequences; POSIX came later and allowed either unless the XSI option is supported).I mean... in main() there's a labeled goto for "just_echo" - even that section alone has something like 6 or 7 indentation levels. If I was writing satire code of GNU style I'd probably wouldn't have gone so far. Hehehe.
@astro Yes. An XSI compliant echo actually supports no options at all. See here for how echo is supposed to work: POSIX. 2008 §echo
the link goes to http://pubs.opengroup.org/onlinepubs/9699919799/utilities/ec... ; here I read OPERANDS
The following operands shall be supported:
string
A string to be written to standard output.
[...]
The following character sequences shall be recognized on XSI-conformant systems within any of the arguments:
[...]
\c
Suppress the <newline> that otherwise follows the final argument in the output. All characters following the '\c' in the arguments shall be ignored.
wait, what???http://www.meetup.com/Classical-Code-Reading-Group-of-New-Yo...
I know those guys were pretty smart, so I'm legitimately curious.
i==argc? '\n': ' '
part of the print statement cleaner.
I meant --no-nl or similar.