tex.web – Version 3.141592653
ctan.math.utah.edu
ctan.math.utah.edu
Pascal code is extracted from the WEB. Then the Pascal is translated into C. Afterwards, years of development have accumulated on that C source to handle utf-8, modern fonts, PDF output, etc.
Most work nowadays happens on top of LuaTeX that replaces and extends pdfTeX.
There are shockingly few people maintaining that core of the TeX environment.
The source control is mirrored here: https://github.com/TeX-Live/luatex/tree/experimental
LuaTeX was in turn an extension of pdfTeX, but I think at this point it would be better described as a fork rather than an extension. (Though it's not clear to me exactly how to distinguish these concepts.)
Anyhow, I believe the 3.141592653 updates have already been ported to all these engines.
You're right, it never caught on, but a few years ago I read a statement by Knuth that said he still writes code that way. Today, about half of the time my Emacs configuration is written in org-mode babel[1] which supports literate programming. In fact, this a common example of literate programming that I see on the internet along with Jupyter Python[2] notebooks widely used by scientists.
Although literate programming seems to work okay for Emacs configurations (often reaching over a thousand lines of code), I've never felt like the approach helped me enough to employ in other programs.
Jupyter was inspired by Wolfram Notebooks, which uses "literate programming" as a term directly in their marketing materials. https://www.wolfram.com/notebooks/
The bugs that have been found in the past couple decades have been rather esoteric. The one that was most likely to actually occur in real use was raised if someone responded e (edit the source at the error location) when no file had occurred (I think it had to happen on the first line of input as well). Most related to extreme edge cases that would be difficult or impossible to raise in a production TeX system.
> A Markdown-formatted document should be publishable as-is, as plain text, without looking like it’s been marked up with tags or formatting instructions (Gruber)[...] forks of Markdown, such as GitHub Flavored Markdown and MultiMarkdown extend standard Markdown in various ways, and are targeted for their particular use cases. GitHub’s implementation for code blocks, for example, looks like something a group of programmers would want to have; it looks more like code than text <http://shindoisshin.net/blog/2014/9/6/standard-markdown-cont...>
Consider this file. Those who are expected to interact with it in this form includes Knuth and his collaborators—the few who have sent in corrections—and nobody else. You aren't expected to read a literate program like this. In an ideal world, where we had computing better figured out, Knuth wouldn't interact with it like this, either.
The second thing to address would be the use of "static" target languages like C and Pascal. A literate programming work should aim to be more like the experience of programming by configuring program entities (I hesitate to say "objects")—like Self, except powered by the written word. Among the negative ways to react to the last remark would be to say, "'Self, except powered by the written word'? So pretty much the opposite of Self, then. Basically, what we already have." But this is too superficial. Consider Jef Raskin on "the woes of IDEs": <https://queue.acm.org/detail.cfm?id=864034>
Second establishing shot: Self and Smalltalk are notoriously image-based systems. You grab hold of objects and then configure them, relying on point-and-click for much of it. Some aspects of these systems will by their nature be exposed in part as text, e.g. textual representation of methods, but you're still interacting with live objects. I'll assume that you're familiar with this, but for completeness including for the benefit of those who aren't, consider the way the way the CSS style sidebar in your favorite browser's Web developer tools work. Only now imagine that instead of just using it for debugging and tweaks, suppose when you needed to make a permanent edit to the style sheet on your homepage, you opened up the CSS viewer, made the edit, and the result persists—not just in your browser, but by changing the very style sheet itself. Here's David Ungar on this experience (around 1 hour in):
> I try to explain to people that the notion of compiler is broken. Of course I learned this from Smalltalk, but what we want to build is experiences—artificial realities that convince you that your source code is real. It's directly executed. There's no lag between editing and running[...] The environment stresses things in your program, not tools—which is another rant I have. It's this whole idea that we want to put you in an artificial reality—I got that from Randy [Smith]—in which it's easy and natural and low-cognitive-burden to get the computer to do what you want it to do, rather than running language translators that turn weird strings of text into bits the machine can run <https://vimeo.com/594563690>
Now I want to attack the traditional experience by doing the opposite and leaning in on the text part. The point-and-click part of object configuration in Self and Smalltalk is not readily reproducible (hence the proliferation of image-based designs for runtime implementations), and texts deliberately designed for consumption are pretty robust as far as transmission goes with respect to both physical transmission as well as a matter of transmission of knowledge—one of the best ways to "get up to speed" on something is to have someone just tell you about it.
A good LP system would work the way Self and Smalltalk do, but with the point-and-click eliminated. Interactive object configuration should be replaced by written exposition. Despite mainstream programming being overwhelmingly text-based, none of them really work this way. (After all, if being text-based were sufficient, then there would be no meaningful difference between mainstream filed-and-compiled source code vs LP systems.)
I suppose what I'm advocating for could be thought of as object-oriented literate programming, as treacherous and unfashionable as the term OO is.
Neither REPLs nor notebooks really capture this in the right way, being rooted still too far away from LP and and more akin to traditional programming, but it's a start. I'm sure that interactive fiction could contribute something here, but with the open sourcing of Inform7 being two years overdue (and itself being implemented in C++), there's a dearth widely available prior art.
On the other hand, ironically, Knuth recommended getting familiar with the program by picking up one particular part and "navigating" the program to study just that part. (See https://youtu.be/D1jhVMx5lLo?t=4103 at 1:08:25, transcribed a bit at https://shreevatsa.net/tex/program/videos/s04/) He seems to find using the index (at the back of the book, and on each two-page spread in the book) to be a really convenient way of "navigate" the program (and indeed randomly jumping through code, as you said), and he thinks that one of the convenient things about the "web" format is that you can explore it the way you want. This is really strange (to us) as the affordances we're used to from IDEs / code browsers etc are really not there, but it makes a bit of sense when you consider that books are Knuth's life, and he must find things like flipping pages with a bookmark in the index page at the back of the book etc really second-nature.
(BTW if you know someone who has the super-obscure interest in making WEB programs actually navigable with hyperlinks etc, do consider https://github.com/shreevatsa/webWEB/discussions or if you know some better way of finding such people let me know :P)
I think that splitting your software into well-documented modules and APIs is a superior approach to lp. Whereas not much software is well-documented these days, where "well-documented" means for me:
a) high-level narrative of the why and how of the code, including descriptions and (informal) correctness proofs of the important algorithms in the code (can be just a doi to a paper)
b) detailed documentation of parameters and return results of the functions that comprise the module(s)
It does allow straightforward, short procedual/structured programs to become very readable and easily understandable - but for bigger "piles of code" - it's probably not that good a fit in practice.
I guess https://github.com/daly/axiom is both an argument for this being true (I seem to recall there was an effort to get away from lp) - and against (proof of existence: it's a big system, it's old, it seems to not be dead).
Then there's the other thing - I don't recall who's quote it is - but it is along the lines of: "There are few good programmers, there are few good writers of prose/technical documentation - therefore the subset of people that are both great programmers and great writers are tiny - and that is the subset for whom literate programming is a great fit".
I do think there's a middle ground though, and "notebooks" for "executable, repeatable" research papers is one such middle ground (or: to write a great cs paper your team need to have both skills anyway).
But there are certainly great programmers that can't write documentation on how to escape a wet paper bag.
So the following correction is not intended as disagreement but simply to share some history / facts:
> Consider this file. Those who are expected to interact with it in this form […] You aren't expected to read a literate program like this. In an ideal world, where we had computing better figured out, Knuth wouldn't interact with it like this, either.
Actually, even in the world we live in, no one is expected to read a literate program like this (Knuth has comments to that effect somewhere), and everyone is "supposed" to read the typeset program (today that would be tex.pdf or whatever). Knuth reads it off the printed book. He navigates the program using the indexes in the book, and as recently as this year he raved about how good his system was (https://tug.org/TUGboat/tb42-1/tb130knuth-tuneup21.pdf):
> While I was preparing this round of updates, I was overjoyed to see how well the philosophy of literate programming has facilitated everything. This multifaceted program was written 40 years ago, yet I could still get back into TeX’s darkest corners without trouble, just by rereading [B] and using its index and mini-indexes! I can’t help but ascribe most of TeX’s success to the fact that it has enabled literate programming.
This works for him, because he started programming in the days when program "listings"—printouts of source code—were the most common way to read programs, and it's also how he wrote TeX (https://news.ycombinator.com/item?id=10172924: “Knuth wrote the entirety of the first version of TeX on yellow legal note pads, and then typed it all in, and then started debugging” — he worked full-time on the program, writing it with pencil on paper, over several months, before typing a single word of it into a computer). And of course, he's a professional book-reader and book-writer (and book-tweaker, as his books keep getting newer editions and updates), so it comes naturally to him.
Sorry if it wasn't clear, but that was the point that I was making. The file that susam has linked to here is nominally "plain" "text", but no reader is expected to consume it in this form, or really interact with it in this form at all. This form serves exactly one set of humans: Knuth, et al, and only while they are actively editing it.
To reiterate my earlier claim: a more perfect LP system would be designed in such a way that the plain text input format would be a suitable way to read programs written in that system. (And when I say "suitable", I don't merely mean "possible". I mean designed expressly for that purpose, like how Markdown was _intended_ to be used before the GitHub contingency got their hands on it.)
> LP-related markup or whatever in the comments
That's sort of the opposite of what I'm after! Please do check the Raskin article I linked to.
About the more perfect LP system, I agree it would be nice if the input format was designed to be the suitable way to read programs: in other words, there would be no "weave" step. [Aside: should this input format be plain text? Maybe notebooks like Wolfram's or "nbdev" count?]
The idea of having the markup in the comments is instead related to eliminating the "tangle" step. (I think there are some LP-related tools that do this; one I know is Harold Thimbleby's "warp": http://www.harold.thimbleby.net/cv/files/warp.pdf — but the structured comments are in (ugh) XML. See also Visual Studio's XML comments…. He also has a later "relit" that does "reverse" LP, extracting source code out of the intended-for-publication format.)
(These goals are indeed sort of opposites, but can be reconciled: I guess we can agree that the ultimate would be a format that eliminated both the tangle and weave steps!)
(My reason for wishing eliminating the "tangle" step (and/or enabling going back from the tangled document to the LP input format) is simply an acknowledgement of network effects: LP is unlikely to ever catch on enough to be the majority, so there needs to be a way for a random programmer using their preferred IDE/editor to edit a "literate" program without getting too annoyed.)
I read Raskin's article and it has good complaints about IDEs, and (although I don't see a close connection) I think I understand what you're suggesting as the ideal language/programming environment. But for me, because of the "LP will never be mainstream" belief, I'm still thinking of targeting mainstream languages, with "code" and "comments". :-) Anyway, this part from Raskin's article is particularly inspiring and quotable:
> …the result was very readable and maintainable. To achieve this, we had regular code-reading sessions in which one programmer read and commented on each piece of code written by another programmer. In addition, an expert documenter/writer worked alongside the programmers. If he did not understand a piece of code, he would interview the programmers until he could write a cogent explanation.
> Incidentally, many bugs and conceptual errors were discovered during this process, which more than paid for itself in decreased debugging time. Some programmers balked at the procedures at first, but all came to be enthusiastic when the project did not slow or founder as completion neared. The project was on time, on budget, and bugless (meaning that the software was released commercially to tens of thousands of users and produced no bug reports).
(Actually, a while after that I got more used to the style and I can now mostly read the program without the difficulties that used to bother me earlier, and I haven't gone back and updated the page. But I think chronicling my early difficulties arising was useful, as I suspect other new people may face them too: and I wouldn't be able to do it now: the "curse of knowledge".)
More recently, I tried to connect some people who may be interested in making this program more readable, converting it into other languages/formats, etc: https://github.com/shreevatsa/webWEB — please join if you're interested or can help or have already done something along these lines.
https://www.docdroid.net/9OBjfK8/tex-pdf
sudo apt install texlive
# maybe download other files from that dir
weave tex.web
dvipdfm tex.dvi
# output is tex.pdfhttps://tex.stackexchange.com/a/576314/2148
I forked and optimized a version of JMathTeX to provide plain TeX rendering in Java:
% A reward of $327.68 will be paid to the first finder of any remaining bug.
2^15, for those playing along at home.> At the time of my death, it is my intention that the then-current versions of TeX and METAFONT be forever left unchanged, except that the final version numbers to be reported in the “banner” lines of the programs should become
> TeX, Version $\pi$
> and
> METAFONT, Version $e$
> respectively. From that moment on, all “bugs” will be permanent “features.”
@* \[1] Introduction. This is \TeX, a document compiler intended to produce typesetting of high quality. The \PASCAL\ program that follows is the definition of \TeX82, a standard @:PASCAL}{\PASCAL@> @!@:TeX82}{\TeX82@> version of \TeX\ that is designed to be highly portable so that identical output will be obtainable on a great variety of computers.
The main purpose of the following program is to explain the algorithms of \TeX\ as clearly as possible. As a result, the program will not necessarily be very efficient when a particular \PASCAL\ compiler has translated it into a particular machine language. However, the program has been written so that it can be tuned to run efficiently in a wide variety of operating environments by making comparatively few changes. Such flexibility is possible because the documentation that follows is written in the \.{WEB} language, which is at a higher level than \PASCAL; the preprocessing step that converts \.{WEB} to \PASCAL\ is able to introduce most of the necessary refinements. Semi-automatic translation to other languages is also feasible, because the program below does not make extensive use of features that are peculiar to \PASCAL.