Literate Programming: Empower Your Writing with Emacs Org-Mode
offerzen.com
offerzen.com
I've used org-mode at work and this was actually the strongest selling point. Using pandoc and massaging the org-mode reader to map better onto the pandoc markdown ast I could export to pdf, html and docx (!!!). I had production code + prose sent to the CTO as a word attachment in an email and she loved it.
The tangle comments also meant I have a workflow to share the code as regular code with the rest of the team and use git to see which code sections they changed. There wasn't enough traffic to justify writing an un-tangler that put their code back into the original document but it seems like a very interesting problem.
Emacs+evil+org+pandoc is the Swiss army chainsaw of text manipulation.
People are writing entire books in the literate style:
http://nbviewer.jupyter.org/github/rlabbe/Kalman-and-Bayesia...
The style is similar to literate classics like "Structure and Interpretation of Computer Programs" and "Structure and Interpretation of Classical Mechanics".
The only downside is version control. Jupyter uses a custom file format. It is difficult to work with colleagues on the same notebook, since we cannot use the usual code-versioning tools (git/diff/etc.) to manually merge together concurrent changes.
Anybody knows a good solution for this?
For those curious, the ipynb files are json files, and the code is stored in an array of cells, as string entries
gitconfig:
[filter "nbstrip_full"]
clean = "jq --indent 1 \
'(.cells[] | select(has(\"outputs\")) | .outputs) = [] \
| (.cells[] | select(has(\"execution_count\")) | .execution_count) = null \
| .metadata = {\"language_info\": {\"name\": \"python\", \"pygments_lexer\": \"ipython3\"}} \
| .cells[].metadata = {} \
'"
smudge = cat
required = true
gitattributes: *.ipynb filter=nbstrip_full
[0] http://timstaley.co.uk/posts/making-git-and-jupyter-notebook...It's right here where you're commenting: org-mode.
If I had more time and energy there are plenty of things I could do: refactor, make comments, change variable names or try literate programming.
But given the limited time and energy which is best?
Put another way, how can you beat thoughtful variable or function names and judicious comments?
But when it comes to writing a tutorial that you can also execute to get the final result, or interact with along the way, then there is nothing better because, really, you're more interested in the prose that lends explanation to the code you're showcasing.
Although this article's about emacs and I've taken the literate approach to my own config too. That's less for the benefit of adding paragraphs of explanation to my config, but because org mode offers a fantastic way to organise code in one file (which is more often than not what an emacs config tends to be).
In either case I don't think you can do better than being thoughtful when writing code, thinking about the bigger picture and not just the individual functions and variables in isolation. For others and also yourself. This is also true for emacs itself, considering the idiosyncrasies it presents when most of us learn elisp by copying someone else, rather than reading the language docs.
This is not the position of someone programming in a literate fashion. Their contention is that prose and code are complimentary. Code explains what the program does, prose explains why it does it that way. They work together to paint a coherent view of your program. With just code, a newcomer reading your program and understanding it 100% is still left with many questions unanswered.
"Ignore these for now" ... import statements and setup
"Now lets talk about X" ... some code
"Remember those import statements? Lets talk about them, but scroll up to see the code because I can't repeat it." ...
"Oh yeah, and we could have done this on X, but again, can't show any code." ...
<<imports>>
code you want to show
and then later define what `imports` should be.
Here's something I wrote before, based on my own experiences. A key phrase from what follows is When the text was munged, a beautiful pdf document containing all the code and all the commentary laid out in a sensible order was created for humans to read, and the source code was also created for the compiler to eat.
People seemed to find it useful then, so maybe they will now as well:
A previous employer (a subdivision of a global top ten defence company) used literate programming.
The project I worked on was a decade-long piece for a consortium of defence departments from various countries. We wrote in objective-C, targeting Windows and Linux. All code was written in a noweb-style markup, such that a top level of a code section would look something like this:
<<Initialise hardware>>
<<Establish networking>>
and so on, and each of those variously break out into smaller chunks <<Fetch next data packet>>
<<Decode data packet>>
<<Store information from data packet>>
<<Create new message based on new information>>
The layout of the chunks often ended up matching functions in the source code and other such code constructs, but that wasn't by design; the intention of the chunks was to tell a sensible story of design for the human to understand. Some groups of chunks would get commentary, discussing at a high level the design that they were meeting.Ultimately, the actual code of a bottom-level chunk would be written with accompanying text commentary. Commentary, though, not like the kind of comments you put inside the code. These were sections of proper prose going above each chunk (at the bottom level, chunks were pretty small and modular). They would be more a discussion of the purpose of this section of the code, with some design (and sometimes diagrams) bundled with it. When the text was munged, a beautiful pdf document containing all the code and all the commentary laid out in a sensible order was created for humans to read, and the source code was also created for the compiler to eat. The only time anyone looked directly at the source code was to check that the munging was working properly, and when debugging; there was no point working directly on a source code file, of course, because the next time you munged the literate text the source code would be newly written from that.
It worked. It worked well. But it demanded discipline. Code reviews were essential (and mandatory), but every code review was thus as much a design review as a code review, and the text and diagrams were being reviewed as much as the design; it wasn't enough to just write good code - the text had to make it easy for someone fresh to it to understand the design and layout of the code.
The chunks helped a lot. If you had a chunk you'd called <<Initialise hardware>>, that's all you'd put in it. There was no sneaking not-quite-relevant code in. The top-level design was easy to see in how the chunks were laid out. If you found that you couldn't quite fit what was needed into something, the design needed revisiting.
It forced us to keep things clean, modular and simple. It meant doing everything took longer the first time, but at the point of actually writing the code, the coder had a really good picture of exactly what it had to do and exactly where it fitted in to the grander scheme. There was little revisiting or rewriting, and usually the first version written was the last version written. It also made debugging a lot easier.
Over the four years I was working there, we made a number of deliveries to the customers for testing and integration, and as I recall they never found a single bug (which is not to say it was bug free, but they never did anything with it that we hadn't planned for and tested). The testing was likewise very solid and very thorough (tests were rightly based on the requirements and the interfaces as designed), but I like to think that the literate programming style enforced a high quality of code (and it certainly meant that the code did meet the design, which did meet the requirements).
Of course, we did have the massive advantage that the requirements were set clearly, in advance, and if they changed it was slowly and with plenty of warning. If you've not worked with requirements like that, you might be surprised just how solid you can make the code when you know before touching the keyboard for the first time exactly what the finished product is meant to do.
Why don't I see it elsewhere? I suspect lots of people have simply never considered coding in a literate style - never knew it existed.
If forces a change to how a lot of people code. Big design, up front. Many projects, especially small projects (by which I mean less than a year from initial ideas to having something in the hands of customers) in which the final product simply isn't known in advance (and thus any design is expected to change, a lot, quickly) are probably not suited - the extra drag literate programming would put on it would lengthen the time of iterative periods.
It required a lot of discipline, at lots of levels. It goes against the still popular narrative of some genius coder banging out something as fast as he can think it. Every change beyond the trivial has to be reviewed, and reviewed properly. All our reviews were done on the printed PDFs, marked up with pen. Front sheets stapled to them, listing code comments which the coder either dealt with or, in discussion, they agreed with the reviewer that the comment would be withdrawn. A really good days' work might be a half-dozen code reviews for some other coders, and touching your own keyboard only to print out the PDFs. Programmers who gathered a reputation for doing really good thorough reviews with good comments and the ability to critique people's code without offending anyone's precious sensibilities (we've all met them; people who seem to lose their sense of objectivity completely when it comes to their own code) were in demand, and it was a valued and recognised skill (being an ace at code reviews should be something we all want to put on our CVs, but I suspect a lot of employers basically never see it there) - I have definitely worked in some places in which, if a coder isn't typing, they're seen as not working, so management would have to be properly on board. I don't think literate programming is incompatible with the original agile manifesto, but I think it wouldn't survive in what that seems to have turned into.
What tools would you recommend for the actual tangle process? Noweb?
The document I wrote into was turned into two documents; latex, which included all the source code and the accompanying discussion and design, ready to be turned into beautiful PDF; and Objective-C, ready for the compiler. Code review was done using the PDF document, design review had already been done before we started writing the noweb (and thus before we started writing the code) and the code review included checking that the design implemented matched the design approved.
It started when I got annoyed with how difficult it was to write a 'code' tutorial.
Essentially it allows you commit a markdown file (or any other format) alongside your source code and you can embed source code snippets / shell commands into this file - which then gets rendered into the main 'output' (the tutorial / article).
Writing why and how you did something.
The setup for org lets you write the equivalent of jupyter notebooks for every language, which means that you can show a toy implementation of what you're doing before the real code. This toy code is live, completely independent of the rest of the project and can be poked at by anyone opening the org file without screwing anything else up.
I have programs written in org-mode which tangle and weave not only the code but the devops. Chapter 1 is the setup for the system, chapter 2-N is code, chapter N+1 launches the app.
I have only really done it with python and C so far. But I managed to get Scala setup today in less time than it would take to get an ide working.
The tools are still immature, tangle especially needs much finer, and better documented control. But even in this state it's the only tool I can use to write programs which I can pickup three years later and grok in an afternoon. The only downside is that the development for the tool is stuck in the 90s with emails for bugs and patches.
I've gotten interest in running a workshop on org mode for literate programming so expect for that to be filled in the next week or two.
I keep wishing for the same button, mid line, every time I'm working around some kind of"proprietary" serialisation* I find in in house projects.
*I tried to coin the acronym OBSEC to mean security by obscurity, fifteen years ago.. but I have since figured how much pointed derogatory comment serves no enduring purpose unless it at least, like SNAFU and FUBAR releases the frustration felt by the observer.
My latest use cases are: - Extracting a list of issue descriptions (for Jira) from a Getting Started guide, that included "todo" comments. - Extracting a list of contacts from notes taken at during a conference. - Drafting Jekyll blog posts in org-mode.
For todo and contact information, org properties [3] are used to specify semantic fields, that can in turn used by the exporter.
To be fair, the work on these exporters in not completed. I am still getting my lisp up to speed. But from what I can see this way easier than writing pandoc backends or jupyter notebook exporters.
[1] Exporter Documentation: https://orgmode.org/worg/dev/org-export-reference.html [2] Example Exporter (.md): http://repo.or.cz/org-mode.git/blob/HEAD:/lisp/ox-md.el [3] https://orgmode.org/manual/Property-syntax.html#Property-syn...
For large normal programs though? I don't think you need that much prose. A program with 100k lines of code would become enormous.
Still, I do wish programming languages had better support for rich comments - why do no IDEs render comment blocks as markdown? Why can't I put diagrams in comments? I have to resort to shitty ASCII art like a commoner.
This has been my experience. Returning to code I wrote as a literate document is a joy.
For my part, I've written at least 3 "non trivial" literate programs using org-babel and describe them here: https://gist.github.com/jpf/d71453f535065a0d9281672152541386
I should mention that one of my "dirty secrets" of writing literate programs is that I always start with an "illiterate" program, tweaking, changing and updating as I go along. Then, once it's something I'm ready to "chisel in stone" I start converting the code into a literate document.
I do this by checking all of my code into Git, then re-create the code inside of org-mode. Every time I "tangle" from org-babel, I do a "git diff" and make sure that I haven't changed the code by documenting it.
I do this because, while it's easy to make changes to a literate program, but it's harder to do a major refactor.
It's a non-starter on a team where people aren't willing to switch over to it (at least to view and interact with org files).
It's difficult for anyone who has a heavy reliance on IDE features; I've spent literally probably close to a hundred hours trying to get emacs + IDE-like integrations working for a variety of languages, with only mixed success.
Yes, I understand that going with emacs is a 'different way' of doing things, and perhaps you won't need your IDE if you just do things the emacs way, but the barrier to entry is very high, between building in new muscle memory for shortcuts, the highly non-standard UX compared to every other editor today, and the hours and hours required to configure everything by hand, even with a starter config like spacemacs.
As a non-emacs user, my successful route to orgmode was with Spacemacs and evil for Vim keybindings. After several false starts, what worked for me was:
1) realizing I didn't need to go full emacs: I stick with Vim for plain text editing, Visual Studio for coding at work and PyCharm/Visual Studio Code for coding at home.
2) consciously learning one new org-mode feature at a time, and adding their shortcuts to my own cheat sheet.
Most folks say J/APL is unreadable, but I find that I can always come back and understand the code as fast as I can read the notebook.
That said when it came to formatting my Emacs init file, with descriptions, and justifications, I chose to use markdown:
https://github.com/skx/dotfiles/blob/master/.emacs.d/init.md
Markdown allows text, and code, to be mixed, and while there is no inline expansion support or other neat features it is minimal enough that it was painless to write and process.
This is a brilliant idea I feel very stupid for not having thought of myself.
Org-mode I tend to only really use for dynamic things. (i.e. "Reports" or tutorials which contain ebedded code rendered output.) Markdown seems like it is easier to demonstrate to other people - and doesn't really require introducing orgmode (which is worth learning, and which does have excellent documentation of its own).
.. however, this article appears to be advocating a style of programming that i'd be more apt to call Test Driven Development than Literate Programming:
it demonstrates a way to use naming conventions and mocks/stubs to describe a program, not a way to use natural language, math notation, etc. to describe a program.
for JS (the lang in the article), proxyquire+simple-mock is a non-emacs-centric way to do this. toss in tape or some other testing library, and you've got some amount of natural language documentation as well.
this is more what i'd consider literate programming in JS land
https://marionettejs.com/annotated-src/backbone.marionette.h...
... by way of disclaimer, i ought to say that, for all i know, i'm grossly misinformed as to the meaning of TDD as compared to Literate Programming
Also, I bet this trick would work (maybe) for formatting comments or introducing ASCII wireframes with tools mentioned at https://news.ycombinator.com/item?id=2651745. (For the ASCII art, maybe an extra hop through emacs isn't necessary.)
Precisely. Here are real world examples of this:
A program that implements on-the-fly encoding of images into SSTV audio files (3620 words): https://github.com/jpf/dial-a-cat#1-855-meow-jam-sending-cat...
An example SCIM server (7966 words): https://github.com/joelfranusic-okta/okta-scim-beta#welcome-...
An example OIDC "RP" implementation (7731 words): https://github.com/joelfranusic-okta/okta-scim-beta#welcome-...
I think they've made quite the impact on the South African dev market to get to the front page here.
Source: Saffer who got out, knows others from UCT etc.
I want to try both Haskell’s built in literate programming support and try the ideas in this article.
When I got started with Literate Programming, what I thought I wanted was the built-in literate programming support that Haskell has (.lhs files) – what I've come to realize is that the ability to re-organize (and re-use!) blocks of code in an org-babel style literate document is extremely powerful and something I would miss using the "we inverted code and comments" approach that Haskell uses.
Haskell doesn't have literate programming support. The "lhs" format merely changes the syntax of block comments, it provides no additional power over {- text here -}.