Writing my PhD using groff
jstutter.netlify.app
jstutter.netlify.app
The great thing about groff (compared in my experience with latex) is that you spend basically zero time on formatting/messing about once you have a set of macros you like, and the document production cycle is really fast so you edit with zero distractions using basically plain text (a lot like markdown) and then any time you want to see the finished product it's very quick to see it.
This is a pity because, otherwise, it is a great tool with its focus on document structure and output quality. I'm currently working on a LaTeX successor which seeks to address these issues, but it is really hard to make the right design compromises here -- what can be programmed? What is accessible through dedicated syntax? How does the structure drive presentation?
Computer typesetting is a rabbit hole, but a fascinating one. And I'm sure the last word on it has not been spoken yet :)
However, many of the documents we like to set in TeX are more complex than that: bibliographies, figure placement, special typographical flourishes.... And here is where the complexity of LaTeX and its macros adds to the inherent complexity of what we are trying to accomplish (and compile times quickly ballon again).
So, sometimes it helps...?
From memory the compile times didn't worry me at all. What did worry me was making it look like I'd put far more effort into documentation presentation than I actually had. Which worked well for me, especially big shout out to the hyperref package to get my far-too-many acronyms linked back to the acronym definition table for every mention), links to every citation, and proper links inside the document to each section /page reference.
Then on top of that my hacked together proofreading tools in emacs, and a torturous 10k word chapter 2 to sideline a difficult politicical problem, and I passed first time!
LaTeX really fails at "register-true" typesetting, though. You have to allow it to extend pages here or there by a line or be willing to fix many orphans and widows by hand. AFAIK, this has to do with the text flow algorithms which are paragraph-based and cannot do some global optimizations. (Correct me if I'm wrong, I'm not an expert.)
Btw, I cannot confirm the compile-time criticisms. A whole book takes just a few seconds on my machine for one run. I wonder what people are doing who get slow compile times.
My masters thesis was written on an old netbook with an Atom processor, plenty of graphics, the compile times got pretty ugly. But I did different files for each section, and set it up so the latex process would automatically kick off and run in the background after writing to the file in vim. Working within constraints like that is sort of fun, it forces you to get the slow operations off the critical path.
Currently I use a script like:
inotifywait -e close_write,moved_to,create -m . |
while read -r directory events filename; do
if [ "$filename" = "$1" ]; then
latexmk -interaction=nonstopmode -f -pdf $1 2> $1.errlog
fi
done
to just re-compile the .tex whenever it changes. I'm not really a bash programmer though so I guess this will probably be ripped apart by somebody here, haha (the top couple lines were probably taken from some post on the internet somewhere). https://sile-typesetter.org/
On of the design goals was to be able to achieve exactly that kind of line matching. IIRC it can ensure that lines on the front & back of a page line up exactly too; apparently this is important for bibles?Worth taking a look at. It recently acquired TeX style mathematics typesetting ability & has a small but active developer group.
Sure, that's better than nothing but I cannot help but wonder whether there could be an architecture where you cut down on repeating work and get faster recompiles that way!
[1]: https://typst.app/
To give some context, I'm a professor in theoretical computer science, so I write a lot of LaTeX documents and notes.
Some observations of my work flow.
- Writing: I'm writing the source, and occasionally look at the output. So as long as the output time is reasonable, then it is sufficient.
- Editing: I'm reading the output, and then edit the source. So going from the output to the part of the source I have to edit should be as smooth as possible.
- Typesetting is the least of my concern. I only check if there are any glaring typesetting problems right before we publish. This takes at most 1% of the total time in preparing a document.
- Live editing almost never happen. But I see why it might be useful to incorporate it into the work flow. (A cursor on the rending of the live editing would be very nice)
There are some choices on how to present source and the rendering.Typst went with the 2 panel design, with one side source, one side rendering. So I found something close to WYSIWYG is better for editing. However, full WYSIWYG is hard to get right and comes with its own problem. Currently I found there are a few common things people do with respect to source/rendering.
- WYSIWYG editors, which renders everything (word, TeXmacs, Lyx). Editing is done in the rendering. It is smooth, but takes a long time to get used to.
- The app Typora that renders everything except the part where you are editing (which shows as the source). This can be generalized to render all except the current line, or something similar. Editing is done in the source, but feels like I'm editing in the rendering. This is extremely smooth for my editing work, and is my preferred way.
- The app like Compositor https://compositorapp.com/ that renders everything, but can call out the selected part of the source.
- The source and render are in two different panels. Editing is done in the source. So usually one can click part of the rendering, and cursor jumps to the corresponding part of the source. This introduce some friction, as the eyes have to do a jump, and also a quick context switch.• SwiftLaTeX (https://github.com/SwiftLaTeX/SwiftLaTeX / https://www.swiftlatex.com / https://doi.org/10.1145/3209280.3209522 — the cool demo that used to be on their site seems to be gone, but see HN discussion: https://news.ycombinator.com/item?id=21710105)
• Texpad https://www.texpad.com/
• BaKoMa TeX (http://www.bakoma-tex.com/) — its eponymous author Basil K. Malyshev passed away recently, but the product and page still exists for now
• VorTeX (see Pehong Chen's PhD thesis from 1988 https://www2.eecs.berkeley.edu/Pubs/TechRpts/1988/CSD-88-436... — it actually discusses the issues of quiescence, etc).
• So if your LaTeX document is taking orders of magnitude more than about a millisecond a page, then clearly the slowdown must be from additional (macro) code you've actually inserted into your document.
• TeX is already heavily optimized, so the best way to make the compilation faster is to not run code you don't need.
• Helping users do that would be best served IMO not by writing a new typesetting engine, but by improving the debugging and profiling so that users understand what is actually going on: what's making it slow, and what they actually need to happen on every compile.
To put it another way: users include macros and packages because they really want the corresponding functionality (and everyone wants a different 10% of what's available in the (La)TeX ecosystem). It's easy to make a system that runs fast by not doing most of the things that users actually want[2], but if you want a system that gives users what they'd get from their humongous LaTeX macro packages and yet is fast, it would be most useful to help them cut down the fluff from their document-compilation IMO.
---
[1] Details: Try it out yourself: Take the file gentle.tex, the source code to the book "A Gentle Introduction to TeX" (https://ctan.org/pkg/gentle), and time how long it takes to typeset 8 copies of the file (with the `\bye` only at the end): on my laptop, the resulting 776 pages are typeset in: 0.3s by `tex`, 0.6s by `pdftex` and `xetex`, and 0.8s by `luatex`.
[2] For that matter, plain TeX is already such a system; Knuth knew a thing or two about programming and optimization!
I am a heavy user of the memoir class, and I have always suspected moving to plain TeX would not be that hard. However, the fraction of users doing this seems pretty slim so modern TeX workflows do not seem really well documented.
These days (in industry) I manage to use pandoc markdown to word for everything (for similar reasons), which is even more limiting than plain TeX. You learn to write around the limitations pretty quickly. :)
The TeX compiler then loads a format like plain TeX (which the above commenter uses), LaTeX, or ConTeXt. The format defines what macros are available. LaTeX adds a package system, as does ConTeXt (modules) so you can import even more macros on-demand. These TeX formats differ in scope and thus speed, LaTeX tends to be a bit heavier but what really weighs it down are the myriad of packages it is usually used with.
Many TeX distributions will define aliases like pdflatex in your path such that you can preload pdfTeX with the LaTeX format, but they are not really separate compilers.
You say "TeX is already heavily optimized", but that's only true for the layout engine. The input language is entirely based on macros and string expansion. That's fine if you're only going to use it for a bit of text substitution. But as a programming language it's inherently slow. (To be fair, I believe Knuth expected that large extensions, such as LaTeX, would be implemented in WEB.)
[1] https://github.com/latex3/latex3/blob/main/l3kernel/l3fp-bas...
If you're going to create a new system with a new way of doing things (i.e. not using the existing popular LaTeX macro packages), then you can already do that on top of TeX, by just not using those packages! (Use LuaTeX and program in Lua instead of via TeX macros, or do the programming outside and generate the TeX input, or whatever.)
What I'm proposing is that the hard/worthwhile problem is to take real users' real LaTeX documents and give them ways of profiling (what inefficiently written macro packages are making this so slow, because surely it's not the typesetting) and replacing the slow parts.
[^reference]: some text
Inline math works by adding $x= \frac{y}{z}$ or in a seperate math block by adding two $ signs before and after.The syntax of markdown is easier there, but LaTeX is arguably much more powerful, e.g. you can load tables from csv data, generate graphs, make it deal with your bibliography, draw circuit diagrams etc. And the layouts tend to look just good.
I ended up converting markdown to Indesign IDML and using this as a source in an Adobe Indesign layout where I could do all the basic typographic settings and styling once and update it on changes.
In my case, I was mainly concerned with making the resulting thesis.pdf PDF/A compliant. PDF/A is a archival compliance standard that's dedicated to the long term digital preservation of PDF files.
Predictably, I got way too carried away as well, and ended up trying to create fully-reproducible LaTeX PDFs as well. It was probably overkill for my use-case, but it did result in a fun blog post where I documented the process [1]
LaTeX does have a native way of generating PDF-A compliant documents, using the pdf-x package. It's still in beta, but it is quite stable and works very well. The advantage of enforcing PDF-A compliance using native LaTeX is that it allows you to take the further step of implementing reproducible builds. Once that is done, you can be certain that given a LaTeX source file, you will be able to generate a bit-for-bit identical document.
Additional post-processing steps will have to be at least documented, and will probably tie the output on the specific version of your post-processor.
This depends a lot. In most of the cases delay is only about 1 second on modern PCs. A bit more when you cite and build the document twice.
You can use LaTeX in many different ways. There are built-in editors and web services such as Overleaf. In the end, they all use the same workflow or dependencies for building the document, but might add an additonal delay.
I too have ended up tweaking my environment a lot. I ended up testing almost every LaTeX workflow.
I finally ended up for just using vim and zathura. Optimised docker image with LuaLatex builds the document. Second favorite would be LaTeX plugin for Jetbrains products. Overleaf is only good for collaborating.
On my desktop pc which has 16 CPU cores, there is only very little latency when compiling. But for text editing, it is a bit rare that you need such PC…
I agree that for many (or even most) documents, LaTeX's compilation delay is generally manageable. However, when it comes to documents with bibliography management, footnotes, margin-notes, and multiple figures, the compilation delay can get quite high.
In my own experience, I had a document of notes containing over a hundred citations managed by biblatex and bibmla [1]. It also had footnotes and margin-notes, requiring an additional repaint. The compilation time on that document was well over several seconds on my laptop, up to dozens of seconds when on battery-power.
> I finally ended up for just using vim and zathura. Optimised docker image with LuaLatex builds the document. Second favorite would be LaTeX plugin for Jetbrains products. Overleaf is only good for collaborating.
I'm very curious to hear about the docker image that you are using. What purpose does the docker image serve in the build pipeline? I know that for compiled software, sometimes having a build environment allows you to better define the environment variables, but to my understanding this is not a worry for LaTeX.
Using Docker brings several benefits. I allows me to share the same build environment for multiple different machines. I can even use my desktop remotely for building the documents if I want, just by sharing Docker Daemon.
Sometimes some package breaks after an update, and Docker allows me to roll back to working environment. I also can declare additional packages and fonts deterministically if I need them. Overall, LaTeX is quite complicated and huge system, and I rather keep it away from my host machine. Maybe Docker is a bit overkill, but I have never wasted time on fighting with package conflicts or installing Latex once again with extra packages for different machine.
* Actually set up a named style for every type of content you have. Creating shortcuts for the common ones doesn’t hurt * use whatever the paid version that powers the free equation editor. It was miles better about 10 years ago * use a master document sub document approach for categorizing things. You wouldn’t have a single text file that’s 100 pages long. Split up Word that way too
I’m pretty sure I got to a state where I was using the tooling as intended because I wasn’t actually fighting the wysiwyg. Now I did switch to LateX at the end because I was tired of not having easy version control. Word has it if you enable change tracking but it can’t beat normal tooling. Also I wanted to learn latek because it felt like a worthwhile investment (it was - writing formulas in latek is wayyy faster to write and easier to maintain).
So I liked LateX just fine. Prefer Markdown / wiki these days because I don’t work with math formulas.
Disclaimer: I have zero experience with the web version and have no idea how it scales. I imagine it still does quite well on large documents but maybe browser rendering is not so good.
Then I wrote two bio heavy papers. Using word. My thesis was in word too. If you have a ton of figures and not a ton of equations it’s not the best choice to use latex.
That is not the LaTeX I know.
Annoyingly, 'new' users, especially led by KOMAScript hints are drawn to use the 'total' positioning that is promised by using the H B P and other options. However, since these aren't holding their promises - at least not the misunderstood promised 'absolute' positioning, frustration is very often creeping in.
I've had too many co authors and friends ask me how to push figures to certain positions where all I saw was premature optimization in terms of positions and tons of wasted cycles (CPU and user) to get to intermediate solutions that are completely unnecessary if only they were saved until the last layouting runs. The change in document creation paradigm (to not care about the layout until the far end) is what 'manages' expectations and where the perceived errors mostly come from.
Last time I used Word for anything significant (a thesis) it was either word 2003 or 2007, and adding a table or inserting a new paragraph somewhere before a figure could mess up all the figure placement (sometimes the thing would literally disappear).
After that I switched to LaTeX and never looked back. Has this recently become better, or was I just unlucky/unexperienced?
You should try TeXmacs; it is not recent, but it has become smooth and it is superior to both LaTeX and Word under every point of view.
Previous documents ive written with word i would do things like tight layouts on images, maybe with anchors, but that's a recipe for things moving around.
When i came to compile my chapters to a final document i used master document-subdocument to pull everything together. I only had a few issues with blank pages being added when exporting to pdf and that was due to my use of page breaks and section breaks.
Ah, so the solution is to use Word as if it was a worse LaTeX, I see :^)
Jokes aside, the precise manual positioning of figures and such (e.g. figures at a certain height of a paragraph, with text flowing around it) was the only potential attractiveness of Word.
If that’s still broken I really don’t see the reason for switching to a program with worse typesetting (and not only) capabilities, given my use case and the fact that I’m quite comfortable with LaTeX by now
Well, this is exactly what I don’t want to do in a wyswyg editor, but glad that it worked out for you
Strong disagree; wysiwyg editing in Word is an exercise in endless frustration, and Word’s typesetting and fonts are so ugly that it’s painfully obvious when a paper has been written using Word.
Just write LaTeX and let it do the typesetting, figure layout, citation formatting, etc. It’s less painful than Word WYSIWYG editing, and the result is far more polished.
Word doesn't really do typesetting, though. You can make credible camera-ready output with it if you're rigorous with styles and learn how to anchor figures and images correctly, but the line and page breaks will still say "hi, I'm a word processor."
I don't really agree. You do however have to accept and become happy with one of LaTeX's ideas about what makes good figure placement though.
You can link word tables to excel and, provided your analysis updates your spreadsheet, can refresh all data instantly.
You can also refer to values in tables in the body of the text.
I know this because I worked in an environment where Word was the only option!
It's not really obvious which is "better" as the two mechanisms work in very different ways. If anything, I would say that OLE can be more general, but the complexity of a minimal program to supply OLE objects is quite high compared with LaTeX.
I was told by a friend that the Equation Editor in Word would silently accept LaTeX math-mode equation syntax and convert it automatically. Besides trying it out briefly, i never used it extensively, so I'm not sure how complete it is. Still, it's there.
Switching back to a simple Word template (no use of tables; just heading styles and bullet points) and submitting the .docx resolved these issues.
Couldn't you have done this with LaTeX anyway?
RIP Richard, your books were amazing and his son thought he was cool because his book was in Wayne's World 2
This article is incorrectly scaled for mobile. There's no padding around the text so it butts up against the edge of my display. The line widths are way too long for comfortable reading. The blog entry also starts off with an unsemantic blockquote element that quotes nothing from a source.
But yes, Pandoc is a cool piece of software.
<meta name="viewport" content="width=device-width, initial-scale=1">Surely WYSIWYG and "office" suites were a disaster for writing. Students seem to spend lost weeks and months fiddling with MS-Word only to create mediocre looking output.
Personally I's say it's hard to beat Org-mode, separate plain text files, then adding the desired exporter and style files at the last minute.
I am suprised, and keep being surprised, that people haven't yet figured out that there is an excellent tool, that is TeXmacs, that manages to make WYSIWYG the best way to write structured documents while having complete control on the output and never having to fiddle with details.
It has low discoverability, so you have to go through the manuals, the mailing list and the forum (and maybe the blog too, at https://texmacs.gitee.io/notes/docs/main.html) to figure out all that it can do, and to have complete control. On the other hand, it is quite usable with default settings.
Table output in particular was much lower quality than LaTeX with booktabs. I think I had to manually resize columns, which was tedious, and I never had to do that with LaTeX. There were a lot of similar situations I found myself in, where I wound up needing to fight TeXmacs quite a bit to get it to output what I wanted.
I prefer my LaTeX workflow where I can edit markup in Emacs, and have a preview almost instantly generated next to my editor by a filesystem watcher & makefile. TeXmacs necessitates using its own interface (which lacks my vim keybindings and Emacs customizations) and I could not find many resources on editing TeXmacs documents in external programs.
I did appreciate that the general typesetting in TeXmacs was high quality, and the ability to type TeX macros and get e.g. enumerated lists quickly was very nice. But overall, I prefer LaTeX.
TeXmacs's own interface is deeply customizable by the user via Scheme.
I think you can set it up to have vim keybindings---see experimental code at https://github.com/chxiaoxn/texmacs-vi-experiment and comments at http://forum.texmacs.cn/t/a-very-tiny-vim-in-texmacs/176 (I know that the lack of a block cursor has put off someone, but I did not find that comment in the brief search I did now).
There's a well working line of business in my Uni that consist on properly final-formatting thesis with Word.
Luckily for Microsoft "easy" products, there are a legion of people that work for free as technical support.
Turns out that really wasn't the bottleneck, and I had just spent another week distracting myself with technology to avoid writing.
Pain points include many customization points: re-creating exact document specifications provided externally, using specific typefaces, creating your own macros... Oh, and leaving ASCII (or ISO-8859-1) for multi-script characters.
Today's groff is a very fine software, if you are satisfied with its default settings and your task is in the domain it handles.
I agree though that it would be nice if the compilation (esp. from scratch) were generally faster.
I relented and went to LaTeX, and while the limitations mentioned here resonate with me, I've found it totally doable.
Perhaps posting a git repository of a sample phd thesis (with a couple of empty chapters, sample figures/images, tables) could be something that others would really benefit from.
For mom, there's pdfmom -step -k. No need to waste time and ssd space with pandoc.
I was hoping it would be via gnu ed too, but they used vim. Shame.
To begin with, it has a very simple and constrained structure: abstract, problem description, state of the art, interesting subtopic 1, interesting subtopic 2, interesting subtopic 3, results and perspectives.
Interesting subtopics are also just previous articles that you can just recycle.
IMHO, I think that's a pretty broad statement that's wrong as often as it's true. Surely it depends on the book, the thesis, and the discipline/genre.
Firstly, not all disciplines have theses broken down in the way you've outline. Theses in certain humanities often more resemble non-fiction than those of other disciplines.
Of course, some disciplines or departments or schools or supervisors will have you write a "thesis by manuscript" in which you present manuscripts you've written as chapter and write little interludes connecting them as well as a unifying intro and conclusion.
This on the face of it might seem like "just recycling" previous articles, but I think it overlooks the fact that those manuscripts must be written, at least in majority, by you. Even when you aren't writing a "thesis by manuscript", most people I know write chapters as they go along their PhD.
And finally, a bit of digression, but I don't think it's reasonable to exclude the amount of research it takes to write "books" or theses from the estimation of effort it takes to write them. It's an integral part of the process.
Instead of actually writing it, you research a million different ways how to render it, and then you write a blog post on it :)
I ended up falling way too deep into the Rabbit hole, and started using NixOS just to write the thesis itself. It did eventually result in a fun blog post, though!
https://shen.hong.io/nixos-for-philosophy-installing-firefox...
My conclusion is that writing a thesis fucking sucks and damaged me.
1. The majority of people just don't care very much. Just get LaTeX working locally, get it to run the packages you need - amsmath, etc. - and start working on your mathematics.
2. There is a large minority who dive deep into the rabbithole on their editing environment, typography, diagramming, etc.
Amusingly, you can tell how likely someone is to fall into either camp based on the amount of care they take with and over their notes (mostly hand-written in my day).
And then he continued publishing books about typography, in his spare time, for the next few decades.
You can just write the text, and write the equations between dollar signs, and it renders you book quality output by default.
(You also can tweak it as much as you want it, and spend as much time as you want it on it.)
Not so fast. You also need to try a million different SSGs for your blog, then eventually write your own.
My advice is: Remember, nobody's going to read it.
My parents' theses were about 50 pages, typewritten, equations and chemical diagrams entered by hand. The got the same grade as I did. ;-)
But once I was done, I wanted to blow off some steam and started writing a silly little tabletop RPG. I decided the rulebook would be text-only for portability with box drawing borders and ASCII tables and stuff, so I spent the first week or so writing a small ASCII typesetting engine in Prolog (because logic programmer).
And then automated the ToC and section numbering .
And then I spent more time writing a vim syntax file so I could read the glorious ASCII with syntax highlighting.
Here:
https://github.com/stassa/nests-and-insects
I'm still looking for ANSI/ ASCII art contributions btw.