246 karma · joined September 8, 2013
Brian Kernighan mentions in the book that awk provides "the most bang for the programming buck of any language--one can learn much of it in 5 or 10 minutes, and typical programs are only a few lines long" [p. 116, UNIX: A History and Memoir]. Also keep in mind Larry Wall's (inventor of Perl) famous quote/signature line: "I still say awk '{print $1}' a lot."
More background on awk from Brian Kernighan in a 2015 talk on language design: https://youtu.be/Sg4U4r_AgJU?t=19m45s
A pointer to the study that you're probably thinking of is research done by Paul Nation and Robert Waring (vocabulary researchers). They cite their own 1985 study and a 1989 study with the following quote: "With a vocabulary size of 2,000 words, a learner knows 80% of the words in a text which means that 1 word in every 5 (approximately 2 words in every line) are unknown. Research by Liu Na and Nation (1985) has shown that this ratio of unknown to known words is not sufficient to allow reasonably successful guessing of the meaning of the unknown words. At least 95% coverage is needed for that. Research by Laufer (1989) suggests that 95% coverage is sufficient to allow reasonable comprehension of a text. A larger vocabulary size is clearly better." [1]
1. http://www.fltr.ucl.ac.be/fltr/germ/etan/bibs/vocab/cup.html
perl -nae'printf("%s:%s %s\n",$F[2],$F[0],$F[1])'Brian Kernighan himself, in this 2015 talk [1] on language design, states that Awk was primarily intended for one-liner usage (he mentions this at 20:43).
[1. https://www.amazon.com/Introduction-Theory-Computation-Micha...]
[1] https://www.amazon.com/Good-They-Cant-Ignore-You/dp/14555091... [2] https://www.youtube.com/watch?v=qwOdU02SE0w [3] https://www.youtube.com/watch?v=IIMu1PGbG-0
But if a person is merely looking to bang out a compiler without getting overwhelmed with how to convert NFAs to DFAs for lexing, etc., some good alternative books are:
A Retargetable C Compiler: Design and Implementation, by Hanson and Fraser (http://www.amazon.com/Retargetable-Compiler-Design-Implement...). This book constructs and documents the explains the code for a full C compiler with a recursive descent approach (no flex/lex or bison/yacc). I have some experience augmenting this compiler, so I can vouch for the book's ability to clearly convey their design.
Compiler Design in C, by Allen Holub (http://www.holub.com/software/compiler.design.in.c.html). Downloadable PDF at that link as well. A book from 1990 in which Holub constructs his own version of lex and yacc, and then builds a subset-C compiler which generates intermediate code.
Practical Compiler Construction, by Nils Holm (http://www.lulu.com/shop/nils-m-holm/practical-compiler-cons...). A recent book which documents the construction of a SubC (subset of C) compiler and generates x86 code on the back end.
If you only spent a few weeks working on that new project, then you probably didn't spend near the amount of time Bill Gates spent writing, thinking about, and reviewing his Altair BASIC code. Even though he whipped up his code in less than a month or two prior to the first MITS demo, he likely spent weeks or months after that demo modifying and polishing the BASIC interpreter for subsequent releases. You didn't mention your experience level, but Gates' years of prior programming experience likely benefited him as well, providing him with a nice cognitive framework to which lots of these facts could "stick."
One additional thing: when Gates says he still knows the source code for Altair BASIC by heart, it probably doesn't mean "completely, line-by-line" by heart. I'm guessing it means he still remembers some snippets by heart, or that he believes he could re-write it from scratch from memory (which would still be exceedingly impressive).
Some useful resources on memory and learning:
Memory and Learning: Myths and Facts (http://www.supermemo.com/articles/myths.htm)
Want to Remember Everything You'll Ever Learn? Surrender to This Algorithm (http://archive.wired.com/medtech/health/magazine/16-05/ff_wo...)
Make It Stick: The Science of Successful Learning (http://www.amazon.com/Make-It-Stick-Successful-Learning/dp/0...)
http://www.amazon.com/Why-Programs-Fail-Systematic-Debugging...
Babies and pre-pubescent children do indeed have a distinct advantage when it comes to learning sounds, tones, and pronunciation for a language [1][2]. It's also been observed that children are more likely to sound native (in language 2) if they move to another country and learn that second language prior to puberty (see again, Nagle [2]). So, gbog, if your child speaks/hears French, Chinese, and English regularly before hitting puberty, he/she will likely sound native in all three languages. But this advantage doesn't nullify the claim that adult learners can still become fluent in a foreign language. It just means that adult learners might never sound exactly like a native speaker, even though motivated learners can achieve near-native pronunciation in their target language [3].
[1] Babies are born with the ability to hear/distinguish between the sounds of all world languages, but this ability starts to diminish around the age of 10 months (pp 39-40, The Bilingual Edge, King and Mackey, http://www.amazon.com/The-Bilingual-Edge-Second-Language/dp/...).
[2] "Critical period research clearly demonstrates that the probability of a near-native L2 phonology rapidly diminishes as we age", Charles Nagle, "A Reexamination of Ultimate Attainment in L2 Phonology: Length of Immersion, Motivation, and Phonological Short-Term Memory", http://www.lingref.com/cpp/slrf/2011/paper2913.pdf
[3] "The results of this study highlight the fact that learners appear to be able to achieve near-native pronunciation without significant formal instruction in pronunciation which perhaps also evidences the role of implicit learning in L2 phonology...", Nagle, http://www.lingref.com/cpp/slrf/2011/paper2913.pdf
http://mitpress.mit.edu/sicp/full-text/book/book-Z-H-11.html...