Writing a Book with Unix
joecmarshall.com
joecmarshall.com
The part I find funny was that the book was about doing network server development with Haskell...on Linux.
If you're asking because you want help with some network programming you are doing in Haskell then you can PM me on some kind of social media.
Write it as a reference that even you'd find useful, with emphasis on defining the basics and getting them right, and increasingly lighter when it comes to intermediate and advanced concepts. Those can either form future books or consulting or both. Since you've started writing, it's worth it to go through the whole process and publish. I did it with a conventional big publisher some time ago, but I wouldn't do that again today. The idea a big publisher would require you use Word is very familiar to me, and I think it's ridiculous.
Google docs change tracking probably could be taught, but you are dealing with people who spent lots of time in their workflow and who are resisting to adapt for a single author. Mind that those reviewers and editors switch from book to book depending on frequency of author's feedback. The author is just one between many ...
If you look at many contemporary books typesetting is not an important topic for many.
Google Docs have a nice change-tracking and commenting system and next to no learning curve. The downsides are the formatting, which is rather unpredictable for complex-ish documents (figures are always a mess), and citations.
Overleaf helps get formatting and bibliography out of the way, but everybody must know at least some LaTeX, and you have to come up with your own comment-and-response system to keep track of what's going on. There is something baked in, but it is not satisfactory. On the other hand, it is very easy to just comment out stuff that is no longer needed in the main text but may still be useful or just serves as easy-to-see revision history.
I realize it was petty, and there are some regrets in how I handled it, I just thought the anecdote was amusing.
LaTeX strikes me as the ultimate tool, and far less intimidating than most people seem to think (Laport's book is an excellent intro), though for most purposes, Markdown is more than sufficient. In practice, I tend to use Markdown and either include inline LaTeX as needed, or convert to LaTeX and continue editing in that where finer-grained control is necessary.
I'd add pandoc to the toolkit, as well as GNU Make. With the two, I've got a standard makefile that can output a wide range of formats (I refer to them as "endpoints") ranging from ASCII text to standalone or snippets of HTML, PDFs, PS, MS Word, OSX, Mediawiki, and others. Adding a new endpoint is a simple matter of tweaking the makefile.
One piece of organisational advice: Do NOT apply your chapter numbers to your filenames. Instead, allow your principle document outline, using an include structure, to define the flow of the text. Depending on the project size and complexity, I'll either directly include chapters within that level, or have a top level of parts with chapters specified (as second-level includes) within those.
As you decide you need to re-work flow, this becomes much easier to manage and rearrange than if you'd pre-labled the files themselves.
Markdown is very nearly fully sufficient, and for almost any nontechnical work with minimal art or layout, will suffice. It falls flat in some interesting areas:
There's no underbar/underline markup.
There is no native colour markup.
There is no formula support -- not something typically encountered in most texts, but when you need it, you need it.
There is no fine-grained placement control for callouts, boxes, figures, images, etc. They simply appear where they happen to be dropped on a page.
(In several of these cases, you can revert to embedded HTML or styles, which are fine when rendering to HTML, but this won't be picked up by all Pandoc endpoints.)
Mind, if a work consists of nothing more than text, bold, italic, strikethrough, super/sub-script, lists, tables, sections, footnotes/endnotes, and images whose placement is not critical, Markdown is entirely sufficient.
But if you find you need more control, exporting to LaTeX and doing your final editing there will buy you a great deal more control.
For any serious writing, Microsoft Word is one of the worst pieces of tools out there. It starts showing its ugly side when start using anything remotely advanced.
As for "anything significant", even OP shows one file per chapter. That's very likely what most users of word processors do, I'd imagine, even though Word (for example) had master/child documents support in the late nineties.
Graphical UI? There are dozens of text editors that work with TeX and most of them provide graphical UIs, with menus and buttons similar to a word processor. In fact, as soon you want to do anything evenly mildly 'advanced', word processors end up being visibly more complicated.
In a word processor, I can click the 'bold' button, or press Ctrl-b to switch into 'bold mode', or highlight some text, and use the button/keyboard shortcut to bold that text.
In a TeX editor, I can also click the 'bold' button or press a shortcut to auto-create a LaTeX bold environment `\textbf{}` with my cursor placed in-between the {}s. Alternatively, I can highlight text, e.g 'my text', and click the bold button or press the shortcut and the editor will wrap `\textbf{}' arount the text, producing '\textbf{my text]'.
Up to this point the two approaches are equivalent. But now say that I want to make all instances of 'important phrase' bold. With the TeX/editor approach, it's just like any other search and replace, I tell the editor to replace all instances of 'important phrase' with '\textbf{important phrase}'. In the word processor, I have to figure out how to click into an advanced search-and-replace and choose something about replace/add styles etc.
In LaTeX, for something complicated, I can figure it out and write my function(s) for it, which are easily re-usable. In a word processor, what one 'knows' in the case of doing something complicated is a series of mouse clicks through menus - which is not only more arcane than an explicit function, but is likely to be disrupted by version changes.
People use spaces for tabs and newlines instead of page breaks despite there being solutions for them even on mechanical typewriters. As long as it works for them, that's ok.
\documentclass{article} \begin{document}
Your whole life story goes here.
\end{document}
That's fscking difficult.
I was using LaTeX for university reports 20+ years ago.
There's a reason that LaTeX was used 20+ years and is still the gold standard: the other options really suck.
There is a reason LaTeX is barely used out of academia.
That's simply untrue. Look at any serious programming books. And LaTeX is used as the back-end for quite a number of text-processing tools.
Professional publishing has long ago switched to DTP tooling, able to handle layouts and colouring in modern presses.
To be fair, the use cases are rather different. Regular fiction books are not so complicated in terms of layout. Magazines are generally involve much more complicated layouts - where there are questions of getting colours right, having text wrap, and so on.
But for technical work, it's insane to use anything not TeX-based (e.g. Scribble), even if you don't use TeX directly.
Why? I'm no big fan of Word and doing any sort of layout in Word is painful, but for actual writing I don't see a problem. It has good tools for outlines, TOCs, reviews and comments, tracking changes, footnotes and citations (when combined with EndNote) and just about anything else I've ever needed. Hell the equation editor even lets you use LaTeX if you're into that sort of thing. What makes Word so terrible in your eyes?
And yes I wrote my Masters thesis and most of my other university work in LaTeX, so I know what I'm comparing it to.
- Fields, or anything that uses fields: They're completely broken. They only seem to function in straight-forward case. For e.g: it's very easy for sequences to completely go off whack.
- File comparisons: this feature can only tolerate up to a certain number of pages. It will crash and burn for large documents. I sometimes wonder why this is even offered!
- Formatting: No matter how carefully paragraph spacing etc. is controlled there are always instances where things won't work as intended. A good example is cover pages.
- Hanging and freezing on large documents while Word figures out how to render them.
- Font rendering: it's never great. It never looks the same as what it would on a printer (irrespective of the printer). The Mac version does a decent job, but the Windows version, even with ClearType configured, is not great.
- Lastly, my biggest issue: feature incompatibilities between the versions of Word available on Mac and Windows. The Mac version (the latest O365) release doesn't have a lot of advanced features that are available on Windows such as document signing and style breaks.
https://github.com/github/markup
and ironically development seems to have stalled on markdown with the advent of commonmark
https://github.com/commonmark/CommonMark/issues/558
All these issues were only created recently, by one person, and are actually not in the right place for the type of discussion they're trying to prompt. CommonMark have a separate [forum](http://talk.commonmark.org/) for feature discussion, which seems quite active.
Others you might want to checkout not necessarily for writing a book but general CLI pleasantness:
- fzf (https://github.com/junegunn/fzf)
- autojump (https://github.com/wting/autojump)
Here's an alternative to fzf, for comparison's sake:
Well at least one reason is because ripgrep is faster. On simple literal queries they'll have comparable speed, but beyond that, `git grep` is _a lot_ slower. Here's an example on a checkout of the Linux kernel:
$ time rg '\w+_PM_RESUME' | wc -l
8
real 0.127
user 0.689
sys 0.589
maxmem 19 MB
faults 0
$ time LC_ALL=C git grep -E '\w+_PM_RESUME' | wc -l
8
real 4.607
user 28.059
sys 0.442
maxmem 63 MB
faults 0
$ time LC_ALL=en_US.UTF-8 git grep -E '\w+_PM_RESUME' | wc -l
8
real 21.651
user 2:09.54
sys 0.413
maxmem 64 MB
faults 0
ripgrep supports Unicode by default, so it's actually comparable to the LC_ALL=en_US.UTF-8 variant.There are other reasons. It is nice to use a single tool for searching in all circumstances. ripgrep can fit that role. Maybe you don't know, but ripgrep respects your .gitignore file.
I love the elegance and simplicity of plain-text notes.
It works wonderfully as a programmer journal since I generally have vscode open anyway (for gitlens even when I'm working in intellij) the friction is close to zero.
[1] https://marketplace.visualstudio.com/items?itemName=pajoma.v... and https://marketplace.visualstudio.com/items?itemName=Gruntfug...
#!/bin/sh
total=0
for FILE in `find . -type f -name "*.txt"`
do
wc -w $FILE
words=`wc -w < $FILE | tr -d ' '`
total=$(($total + $words))
done
printf "%'d" $total
echo " words"
all this achieves is wc -w $(find . type f -name "*.txt") | sed '$s/total/words/'
and frankly, i'm not sure the total->words substitution is worth the trouble.then there's the inefficiency of running wc twice per file. while this is not exactly bitcoin-level disaster, it rubs me the wrong way...
wc -w $(find . type f -name "*.txt") |
awk -v t=0 '
{ print; t += $1 }
END { print t, "words"; }
'
personally i'd just do this (in zsh): wc -w **/*.txt(.D)
the (.D) is two "glob qualifiers": the . (dot) limits the result to plain files, the D turns GLOB_DOTS on for the pattern.* https://github.com/DaveJarvis/scrivenvar
* https://github.com/DaveJarvis/scrivenvar/blob/master/USAGE.m...
The software provides a simple way to include variables in technical documentation. It also integrates with an R engine for editing R Markdown files, which can also use variables sourced from an external YAML file. (Editing XML documents that have stylesheets is possible, too.)
My authoring workflow involves Scrivenvar, Markdown, pandoc, knitr, and ConTeXt. As Markdown separates content from presentation, I prefer ConTeXt to LaTeX for the same reason.
I've now put my scripts on https://github.com/boazbk/tcs/tree/master/scripts in case anyone finds them useful. (This is not a "plug and play" package that you can install and use, but people that are better programmers than me might be able to adapt it and improve on it.)
I write all my books using vi (not vim!), troff and friends, make, and ghostview (gv) for the layout. Plus a couple of shell/awk/sed scripts for making the TOC, index, etc. I cannot imagine any better tools for the job. I tried LateX, which only got me into trouble, and Lout, which was fun, but too complex in the end. After 20-something books, above turns out to be the sweet spot.
Ripgrep is maybe an alternative for ag
IMO if you are using git then you should use `git grep` rather than ag/ripgrep/ack/grep etc.
wc -w $(find . -type f -name "*.txt")Two of the three comments already referenced bat. It's also what impressed me the most. It looks great. Link below:
There's no "original Unix" (except the first Unix back in the day). There's lineage of operating systems.
Heck, even the people who actually created UNIX in the 70s and early 80s don't have such stickups.