http://www.freebsd.org/cgi/cvsweb.cgi/~checkout~/src/bin/cat...
http://git.savannah.gnu.org/cgit/coreutils.git/tree/src/cat....
http://www.freebsd.org/cgi/cvsweb.cgi/~checkout~/src/bin/cat...
http://git.savannah.gnu.org/cgit/coreutils.git/tree/src/cat....
As always, if you want to read C code as written by the same people who invented C, the Plan 9 source code (http://plan9.bell-labs.com/sources/plan9/sys/src/) is a great resource.
1) The GNU version has MUCH more verbose commenting 2) The GNU version reads input in blocks in cooked mode, rather than a character at a time as the BSD cat does. This is MUCH faster, but it leads to a lot more complexity, and thus more code. It also carefully calculates the optimal block size to use for this; this is also quite complex and carefully commented (see, eg, the 20 line comment at line 736). 3) The GNU version has to be portable to multiple unixes, and thus has a number of places where it has to test for multiple error codes, and/or missing features 4) The GNU version has a very verbose --help output, which consumes a good page or so of code by itself.
I don't really see any 'bloat' there. Sure, it's okay to have a simple cat, but it's not a bad idea to optimize a tool that's used so frequently. And the verbose --help output, verbose commenting, and portability are all part of the GNU coding standards. You can argue about whether you want to spend all that effort on it, but I don't think the sheer volume of code is a good measure for whether the code is good or not.
As the poster above said, one man's bloat is another man's essential feature.
Back to the subject: reading those old sources really learns you why, back in the seventies, people found Unix so appealing. even ignoring the feature growth/creep (or whatever you want to call it), you do not have to wade through a zillion copyright header lines, option parsing that goes on for ages, locale-specific stuff, etc, before getting to the meat of the program. Disadvantage is that some code dives into assembler fairly quickly (for example, printf is mostly assembly in the system I refer to above)
Now, what do you think is more appropriate for educational purposes. Source code with plenty of comments or source code with nearly no comments at all?
I routinely gloss over comments when reading code anyway; the most accurate documentation is found via reflection & introspection on the system itself, not comments.