Donald Knuth – The Patron Saint of Yak Shaves (2017)
yakshav.es
yakshav.es
Though fully documented and available via anonymous FTP on the Arpanet, all of these prototypes were experimental proofs-of-concept and were completely discarded, along with the languages they implemented, with the “real” versions rewritten from scratch by Knuth (each first by hand on legal pads, then typed in and debugged as a whole.)
And you missed one more obscure but instructive example: To get the camera-ready copy for ACP Vol 2, Knuth purchased an Alphatype CRS phototypesetter. And, unhappy with the manufacturer’s firmware, he rewrote the 8080 code that drives the thing. Eight simultaneous levels of interrupts coming from various subsystem: the horizontal and vertical step-motors that you had to accelerate just right, while keeping synchronized with four S100 boards that generated slices of each character in realtime from a proprietary outline format, to be flashed into the lens just as it was passing the right spot on the photo paper. (Four identical boards, since you had to fill the second while the first one did its processing; times two for two characters that might overlap due to kerning). Oh, and you had to handle memory management of the font character data since there wasn’t enough RAM for even one job’s worth (see our joint paper).
Fun times. I did the driver on the mainframe side, so got to be there for the intense debugging sessions that used only the CRS’s 4x4 hex keypad and 12-character dot-matrix display.
Thanks for the DVI shout-out, btw.
If you are referring to the output of TeX, according to Joris van der Hoeven it is matched and surpassed by TeXmacs
(Six bits for each upper-case alphanumeric character means a 6-character directory name neatly fits into a 36-bit word on the DEC PDP-10. And file names were limited to 6 characters for the same reason; extensions were 3 characters, with the remaining 18 bits for file meta-data, I think.)
Thank you for creating it! Is there some article or something which talks about the original project or motivation for the DVI format? If DEK implemented it as part of TeX in 1982, what was DVI being used for before that?
(For those unfamiliar, https://en.wikipedia.org/wiki/Device_independent_file_format)
Well, imagine you're an unaccomplished grad student, and you're going to tell Prof. Donald Knuth that he's wrong-headed, that his approach wouldn't scale, and that the right thing to do instead is to have a simple, intermediate format, so that nobody would have to muck around in TeX's internals to get a new output device going. Quite the rush when he gave the go-ahead.
So, the first DVI format was purpose-built for proto-TeX (and underwent a revamp for the final cut, just like everything else). It's really just a slightly compressed form of display-list (move to x,y; put down character c in font f; move; put down a rule; done with page) and nothing to write home about. While it was important in the early days to help us and others to get various devices going, and allow for Tom Rokicki's fabulous(ly important) DVI-to-PostScript program, given the ubiquity of PDF today, it's appropriate that we've come full circle, and PDF output is built into most TeXs that people now use, leaving DVI mostly a historical footnote.
Presumably without the source code, so he RE'd it too? In that case, I'd say he captured the spirit of right-to-repair/read/modify even better than Stallman (who, if I remember correctly, founded his GNU empire entirely on complaining about not getting the source code for, ironically enough, also a printer.)
The key issue was that their firmware insisted on loading proprietary font outline info from the included 5.25" floppy drives, while we needed to be able to have individual characters downloaded dynamically for caching in the machine's 64K (or less?) of RAM, intermixed with the per-line display list info, such that once you started the stepper motor that controlled the horizontal worm-screw that moved the lens over the paper, everything would be ready as it moved along; having to stop and restart a line would inevitably lead to jitter among the characters, as mechanical systems are subject to hysteresis -- you can never get back to exactly where you (thought you) were. I vaguely recall Knuth grousing that the standard firmware couldn't even be proved to be able to handle an arbitrarily complex line of text, and thus his full rewrite.
Ultimately, the output was not completely satisfactory. Tall parenthesis and integral signs are made up of multiple pieces (as that's how TeX and Metafont handle them), and because of the alignment issues mentioned above, plus the fact that the individual pieces were "flashed" onto photographic paper separately, so where they met there was a bit of multi-exposure that lead to some "spread" of the blackness. You can see this if you look closely at integrals in ACP Vol 2, 2nd Ed.; they're a little chubby right along the centerline.
I mention all this as it's yet another case of throwing away a prototype, this time software plus hardware. The Alphatype was replaced by an Autologic APS Micro-5, which didn't use a moving lens, and was generally more digital, and thus had no problem with characters made up from pieces. And no firmware needed replacing, though we did have to run it in a completely unintended mode where each character was sent in run-length format each time it was typeset (meant for occasional one-off logos); this was wildly inefficient and slow, but as the APS used long rolls of photographic paper rather than single large sheets like the Alphatype, it could run unattended for hours to produce many dozen pages at a go. At least Knuth didn't have to rewrite any firmware for it (nor software; I handled it, and avoided telling him how grossly inefficient it was, as I thought it would bother his soul).
Yak Shaving is the frustrating work we do because of impedance mismatch in the system. Libtool is a huge yak shave, for example. But so is writing a custom CSV parser because an input file is just not quite accessible by the standard, and you could just use python to make the translation, but your environment won't let you use python in that spot, etc. Days later nothing is really accomplished, you hate programming, and you can finally start work on the real problem after losing the motivation you had when you started.
But I like the idea of a patron saint for this, and Knuth is as good as any. Perhaps Dijkstra could be the patron saint of Unused Better Solutions. I'm sure there's a litany of things that could use saints in this field.
Maybe Knuth could have used Fortran and not Pascal, but that doesn’t solve the literate programming problem ... he would still want something like WEB.
Seems like Fortran is more maintained today than Pascal, but Knuth couldn’t have known that. The community yak shave to translate the code arose out of other forces in the industry.
I don’t want to go too far into it cause the website is a treasure trove and exciting to explore on your own but if you follow some of their latest stuff, they’ve built a stack based virtual machine, and a whole set of software around it and it truly feels like they are just exploring every avenue that interests them without rushing themselves. Truly the embodiment of my grow a beard and learn Haskell dreams.
Iirc that should be part of the taocp plan.
Syntactic Algorithms, in preparation.
9. Lexical scanning (includes also string search and data compression)
10. Parsing techniques
...
And after Volumes 1--5 are done, God willing, I plan to publish Volume 6 (the theory of context-free languages) and Volume 7 (Compiler techniques), but only if the things I want to say about those topics are still relevant and still haven't been said."
Given what he did for computer algorithms (and doc type setting) this makes me really curious about his views on compilers.
If you look into any great pre-TeX book like Rudin [1], it looks much uglier and more difficult to read. I'd actually pay a significant amount of money for new editions of Rudin or Halmos typeset in TeX and Computer Modern. They'd be much easier to go through.
I must also add Latin Modern is a slight update to Computer Modern that is slightly thicker and looks much better on screens.
[1] https://web.math.ucsb.edu/~agboola/teaching/2021/winter/122A...
Caveat, I like computer modern.
That is, is this a blending of bad subjective experience with otherwise good objective experiences?
Knuth has chosen Computer Modern to be a so-called "Modern" typeface, which is a name applied to a certain style of typefaces that became popular at the beginning of the 19th century and which was frequently used for mathematics books of that time.
Of all styles of typefaces, the "Modern" typefaces are the least appropriate for computer displays, because no desktop computer display has a resolution high enough to render such typefaces.
On computer displays, Computer Modern and all typefaces of this kind (with excessive contrast between thick and thin lines) are strongly distorted to fit the display pixels, so they become much uglier than when rendered correctly on a high-resolution printer (i.e. at 1200 dpi or more).
Therefore you cannot compare Computer Modern with more recent typefaces that have been designed to look nice on computer displays or with old typefaces from other families, which were suitable for low-resolution printing.
Computer Modern was strictly intended for very high resolution printing on paper and not for viewing documents on computers.
I typeset 2 books with TeX and read far more printed pages typeset with TeX than I ever read online.
Makes me wonder what you think of Comic Sans ...
For many things, I'm a Palatino Person, but I am aware that it is aging poorly. I also love the european public display use of Helvetica, and for novels and other continuous text flow, I'm a cheap date, anything will do.
Try the closely related Aldus, or in extremis just condense Palatino by a few percent. (The latter suggestion may get me yelled at, somewhat understandably, but I really do think it's an improvement if you're setting substantial quantities of text.)
For me it's the exact opposite. I find the font to be beautifully designed and perfect for mathematical typesetting (e.g., it properly distinguishes ν and v for instance). However, it's terrible for on-screen reading and printing using laser printers because of how spindly the glyphs are. Indeed, Knuth designed CMR for the printers of the 1980s [1].
The most carefully designed math fonts IMHO are the MathTime fonts [2] designed by Michael Spivak to use in conjunction with Times. It isn't a free font however. (For a sample see this paper [3] published in the Annals of Mathematics.) MathTime is also a perfect example of yak shaving. Spivak was dissatisfied with most available fonts when he set out to update his book on calculus so he basically learned type design from the scratch [4]. I also really like the Fourier fonts (meant to be used with Adobe's Utopia) designed by Michel Bovani and I recently came across a book on statistical mechanics [5] that uses them, which I thought was typeset beautifully.
[1]: https://tex.stackexchange.com/a/361722
[2]: https://pctex.com/mtpro2.html
[3]: https://annals.math.princeton.edu/wp-content/uploads/annals-...
[4]: https://tug.org/pracjourn/2006-1/spivak/spivak.pdf
[5]: http://www.unige.ch/math/folks/velenik/smbook/Statistical_Me...
https://news.ycombinator.com/item?id=16526151
Today, major TeX distributions have their own Pascal(WEB)-to-C converters, written specifically for the TeX (and METAFONT) program. For example, TeX Live uses web2c[5], MiKTeX uses its own “C4P”[6], and even the more obscure distributions like KerTeX[7] have their own WEB/Pascal-to-C translators. One interesting project is web2w[8,9], which translates the TeX program from WEB (the Pascal-based literate programming system) to CWEB (the C-based literate programming system).
The only exception I'm aware of (that does not translate WEB or Pascal to C) is the TeX-GPC distribution [10,11,12], which makes only the changes needed to get the TeX program running with a modern Pascal compiler (GPC, GNU Pascal).
...
I may write a blog post on this since it's relevant to how https://www.oilshell.org/ is written in a set of Python-based DSLs and translated to C++.
https://tex.stackexchange.com/a/576314/2148
Although TeX has many years of development effort behind it, the core functionality of converting macros to math glyphs is reasonably straightforward and can be accomplished in a few months---especially given that high-quality math fonts are freely available. Here's a Java-based TeX implementation that provides the ability to format simple TeX equations:
https://github.com/DaveJarvis/JMathTeX
The code supports neither vectors/matrices nor bracket sizing (PR welcome!):
Increase your price by 25% every year or so, until you start losing customers, and ask your previous customers if you can use them as a reference.
It's a good gig because it actually requires something that is somewhat rare: lots of experience, so it's a natural match for the older hands-on tech people.
I understand what you mean with this in context, and it's true, but it's still somewhat funny to talk about Donald Knuth and "focusing on one thing". It's more like he focused on everything in programming, and somehow does it well.
In the case of my example, note that when you load this paper in a browser, the tab will say "main.dvi". So I'm guessing the paper was typeset in LaTeX, published as DVI, and when PDF came along they converted it, but the DVI -> PDF conversion algorithm was really bad.
[1] https://courses.cs.duke.edu/spring02/cps296.1/papers/H-SIGMO...
Bitmap version of the font, I guess.
https://legacy.cs.indiana.edu/~dfried/mex.pdf
(I only know about this paper because someone linked to it on HN several years ago.)
I'd be interested to know more about how the conversion to PDF ended up like this.
These are examples of files that have gone the route DVI → PS → PDF, where the PS file contained Type 3 (bitmap) fonts without hinting instructions for on-screen viewing. If you have the original PS file you can often fix them with Heiko Oberdiek's amazing `pkfix` (and `pkfix-helper` if needed) tool [1].
(You can zoom the PDF to 500% to see the font's glyph shapes down to actual pixels; this pixelation is not inherently a problem as these PDFs look fine when printed on a typical high-resolution printer: try it!)
In more detail for your example PDF [2]: running `pdfinfo` gives:
$ pdfinfo H-SIGMOD1999.pdf
Title: main.dvi
Creator: dvips 5.58 Copyright 1986, 1994 Radical Eye Software
Producer: Acrobat Distiller Command 3.0 for Solaris 2.3 and later (SPARC)
CreationDate: Thu Jan 24 16:43:42 2002 PST
So presumably:• `main.tex` has been typeset with TeX into `main.dvi` at some point (the paper has a "September 1998" and "SIGMOD 1999" in the title, so presumably it's from then),
• This `main.dvi` has been converted into a `.ps` file at some point, using dvips 5.58 from 1994 (dvips 5.70 was in 1997).
• This `.ps` file has been converted into `.pdf` using Acrobat Distiller on SPARC, presumably in 2002.
Now, with some searching online we can actually find the original PS file: it's at [3] (and dated "24-Sep-1998 19:15"). The dvips version 5.58 is too old for pkfix to run its magic directly, but by running pkfix-helper first (which does guessing based on font metrics, and in this case seems to have guessed mostly correctly: though the superscript font for footnotes is wrong), and then pkfix, and then converting to PDF, we get this equivalent I just made (compare with your example [2]): [4]
[1]: https://en.wikipedia.org/w/index.php?title=Pkfix&oldid=10194... and https://ctan.org/pkg/pkfix-helper?lang=en
[2]: https://courses.cs.duke.edu/spring02/cps296.1/papers/H-SIGMO...
[3]: https://www1.icsi.berkeley.edu/ftp/global/global/pub/techrep...
[4]: https://shreevatsa.net/post/2022-pkfix-etc/tr-98-033_pkfix-h...
I’m still giggling at this one. Amazing write-up.
In 1968 a book like TAOCP needed to cover assembly programming. Using an existing assembly language for an existing machine would not work well unless you wanted the book to be "TAOCP for CDC 6600 Programmers" or "TAOCP for Burroughs B5500 Programmers" or "TAOCP for IBM/360 Programmers" or similar. There were a lot of significant differences between different machines, and so if your book used any particular one for its implementations it would be harder to use for readers who used a different machine.
MIX allowed Knuth to capture the important points that you needed to learn about assembly programming, without getting you tied to any particular machine.
I took a 1-week summer class at Stanford with Knuth. At lunch, he told us the story of a Chinese restaurant that served food in the style of the province of Hunan. They had chopsticks made with a typo, said "Human cuisine".
But how did it go from preparing for yaksmas to the current meaning?
He started with wanting a digital version of Monotype Modern 8A (what became Computer Modern). His first idea was to go to Xerox and use their scanning equipment to digitize the font, but they would only let him use it if they would have copyright on the resulting fonts, so he started looking deeper at the fonts himself. That gave him the idea of describing the font shapes from scratch (with equations) instead of simply scanning them, so he came up with METAFONT. To use his own digital fonts, he would need a typesetting system: TeX. He did implement TeX in SAIL, but when it became clear he would need to rewrite it in a portable way (Pascal), and Hoare suggested publishing the program as a book, he came up with literate programming and WEB.
And as mentioned in another thread (https://news.ycombinator.com/item?id=29862858), TAOCP is itself a "yak shave", Knuth's response to being asked to write a book about compilers, in his first year of grad school in 1960 (given his compiler exploits earlier, like http://ed-thelen.org/comp-hist/B5000-AlgolRWaychoff.html#7). The publishers who were hoping to sell a book to aspiring early-1960s machine-language compiler-writers never got their wish, but they got something much more.
I wish he would update his volume on Searching and Sorting with modern search engine techniques.
And I would expect coverage of NLP techniques.
Of course, a search engine makes use of distributed computing techniques, so I'd also expect a volume on that topic :)
I’m curious on if this helped or hindered AoCP? Did it improve the work because the great theorist is also practicing on the side? Or did it slow down creation of AoCP? If the answer is “both”, which effect is stronger?
This is written with all due respect to this great polymath.
Being in total control of your tooling is very nice for work that will span decades.
I wrote some macros in TeX, over 100 of them, and in them used programming language constructs including if-then-else, do-while, dynamically allocated typed variables, and file reading and writing. Sounds like a "programming language" to me.
Sorry, TeX was not written primarily for programmers or computer scientists. Instead the target audience was people who use a lot of mathematical notation. In my opinion, the main glory of TeX is how it helps position the mathematical symbols in mathematical expressions, including some really complicated ones.
I know a guy, a good mathematician, who had a really tough time understanding the purpose of TeX. He kept evaluating TeX in terms of what it did for the future of word processing taken very generally, maybe all the way to video as in some Hollywood movies. E.g., TeX is not promising for generating video of a Darth Vader light saber battle. Then he noticed that TeX is not really that future. I finally explained to him that TeX was not trying to be the future of some generalized word processing, thus, was not looking ahead, and instead was looking back and at something he knew well -- the literature of advanced math as in math journals such as published by the AMS (American Mathematical Society or some such). So, TeX was to ease the word processing needed for pages of mathematics as in the math journals and textbooks.
That friend kept asking me to write a converter that would convert a file of TeX to a file of HTML. I kept telling him that such a converter was impossible because TeX was a programming language and HTML was not. I did explain that at least in principle could write a converter to convert TeX output, that is, a DVI (device independent) file to HTML. There are converters, heavily used, to convert DVI to PDF (portable document format or some such).
I like TeX; it is one of my favorite things, and I use it for all my higher quality word processing, the core, original math for my startup, business cards, even business letters. My last published paper (in some mathematical statistics) was in TeX, and using TeX was liberating because I could just go ahead and do the math and not worry about how I was going to get the word processing done, did not have to bend the math and reduce the content to make the word processing easier.
Future of TeX? The fraction of the population that wants to typeset complicated math expressions seems to be tiny, and there are lots of alternatives for others. So, my guess is that TeX will be like, say, a violin -- won't change much in hundreds of years.
The OP mentioned LaTeX: For people new to TeX, no, you don't have to learn LaTeX. The approach of LaTeX is different. In an analogy, LaTeX wants you to state if you are building a bicycle, motorcycle, car, truck, boat, or airplane, and then lots of lower level details are handled for you. With TeX, never decide what vehicle type are building and, instead, work with the parts and pieces -- yes, with a lot of help.
There is Knuth's book on TeX, The TeXbook, and also the books on LaTeX. In comparison, Knuth's book is a lot shorter than the books on LaTeX. So, I got the books on LaTeX, looked at them, and decided that it was easier just to stay with TeX and the macros I could write for TeX. So, for people new to TeX, don't really have to get and read the books on LaTeX.
Getting math typeset was a big problem. TeX is a good solution. Problem solved. We can move on!
The collection of TeX macros I wrote has over 100 macros
Being able to seamlessly transform the language on such a fundamental level is one of the reasons I don't think we'll ever see a next major TeX - all it would need to be is a macro atop of TeX. And then everyone has their own feelings, so it'd probably stay as is.
According to Massimiliano Gubinelli (please take a look at the following comments thread https://news.ycombinator.com/item?id=27820466, in particular at https://news.ycombinator.com/item?id=27822662), the reason why a converter is impossible is that the _syntax_ of TeX is Turing complete. If one had kept syntax and programming constructs separated, then a converter would have been possible. (I understand the argument only superficially but I trust Gubinelli enough to report this here).
> Future of TeX?
is TeXmacs, worth trying
[1] https://archive.org/details/AnEssayTowardsARealCharacterAndA... , last paragraph, he also mentions other books on letters and parts of letters.
[2] http://www.perseus.tufts.edu/hopper/text?doc=Perseus:text:19...
I have seen of the scrawnier ones in labs, they probably would only feed one or two people so I am not sure the utilitarian approach is a winner ? Maybe if we fed them a diet of Mountain Dew and Pizza it might improve the flavours?
You could spend a couple of days chasing down the larger specimens, but they tend to put up more of a struggle, and in the time spent, you could have a dozen scrawny engineers already in your larder or stew.
Fascinating discussion :-)
If you want to make sure you're doing something sufficiently differentiated, find an area where you still need to shave a few yaks.