Sile: A Modern Rewrite of TeX
sile-typesetter.org
sile-typesetter.org
Code it, explain it, generate a Literate PDF containing the code.
The programming cycle is simple. Write the code in a latex block. Run make. The makefile extracts the code from the latex, compiles it, runs the test cases (also in the latex), and regenerates the PDF. Code and explanations are always up to date and in sync.
I have found no better toolset.
This is (most of) the source code for Axiom. The code is extracted from the latex for these pdfs.
I am currently a physics student and one of the main problems that I have day-to-day coding is being unable to include diagrams/equations/pictures in the source code. Moreover, because there is a fair amount of 'domain complexity' the comments are often longer than the code itself.
The package is still very much unstable and not even close to being ready for any sort production. However, I did use it for my masters thesis [2], and it was quite an interesting (in a good way) experience. (I recommend everyone to actually try to write literate code at least once).
[1] https://github.com/Skipp1/fortex [2] https://github.com/Skipp1/mono-rich
- simple things became uselessly lengthy and prolix
- complex things became too big to be easy read, while the equivalent "pure-code with comments" and a small new developers doc while fail to give an equally deep knowledge of the code base is FAR more quick and easy accessed.
Maybe it's my style, I do not know but while I do not feel LaTeX as painful, since most my docs are of the same kind so a template/class produced once with calm get reused issueless, I can't really digest literate programming...
Lisp program to extract latex chunks: https://github.com/daly/axiom/blob/master/books/tangle.lisp
C program to extract latex chunks: https://github.com/daly/axiom/blob/master/books/tanglec.c
Note that the C program is just a hand translation of the Lisp code.
The lisp code has an explanation and the necessary latex macros. The idea is to scan the latex, find each named code 'chunk', and add each one to a hash table. Then the hash table is scanned to dump the requested chunk to stdout. For example:
\begin{chunk}{part1} code for part 1 \end{chunk}
Ordinary latex code between chunks.
\begin{chunk}{part2} code for part 2 \end{chunk}
\begin{chunk}{part1} this code will be appended to the prior chunk \end{chunk}
\begin{chunk}{getall} \getchunk{part1} \getchunk{part2} \end{chunk}
Assuming the above chunks are in the file foo.tex then
tanglec foo.tex getall >getall.code
will print out the named chunk (getall). For an individual chunk use
tanglec foo.tex part2 >getsome.code
(Unfortunately, the package that lets each Org source block behave as though it was using the corresponding language's Emacs major mode - poly-org-mode - has a ton of bugs. It was part of why I stopped using literate programming entirely for later projects.)
[1] https://orgmode.org/ [2] https://github.com/jingtaozf/literate-lisp [3] https://github.com/jingtaozf/literate-elisp [4] https://contrapunctus.codeberg.page/blog/literate-programmin...
I tried orgmode. I even attended a course at CMU that used it for the "live notes". It is excellent for teaching. But it has the same flaw as the "live notebook" idea. There is no generally accepted structure to the approach.
Everyone "understands" books. They have a preface, chapters, an index, a bibliography, pictures, credits, and a table of contents. Literate programs leverage that shared understanding of the structure.
Think of a physics textbook. If you just copy every equation out of the book then you "have the code". All the rest is explanation. They belong together so the explanation and equations (code) are intermingled.
I'm a "primitivist". I work in straight text at an emacs buffer in fundamental mode.
The point of my code is to "talk to the machine". The point of my literate program is to "explain to other programmers (mostly 'future' me)". Note that this is NOT DOCUMENTATION. It is explanation, best presented in book form.
For the last two evenings, I have been revisiting my old manuscript materials because I am thinking of updating the material and also creating additional editions for more programming languages: Python, JavaScript, and maybe Swift and Prolog.
I had forgot how cool TeX and LaTex are. It was very easy to start working with again, even after a 12 year gap.
Many publisher's workflows involve a round of content editing bouncing a word file back and forth, then a period where a typesetter uses InDesign or similar to lay it all out, then it goes to press. You can't keep copy-editing after the designer takes over. With a workflow using source documents in Markdown and typesetting handled by SILE I am able to allow copy-edits to book manuscripts up until minutes before going to press.
Personally, I have given up on Latex and write documents in pandoc markdown which I can then convert to pdf through the intermediary latex.
• https://www.youtube.com/watch?v=5BIP_N9qQm4 [FOSDEM 2015, "Introducing SILE: A New Typesetting System"]
• https://www.youtube.com/watch?v=t_kk20vlamo [GRANSHAN Conference 2015, "Global typesetting with SILE"]
It is also somewhat puzzling that same people who advocate for a better grammar would also push the idea of a better markup. So which is it: a better programming experience or a pure markup language that one wants? I feel TeX strikes a near perfect balance here. Most importantly, TeX has stayed nearly unchanged at its core for almost forty years! It is its greatest strength, not weakness. I cannot see what latest fads can improve in TeX
Context-free languages are the standard for decades now. They haven't been in the very early days of computing due to a lack of formal education of programmers and because they come with a certain, but tiny, memory requirement.
What Knuth did with TeX was a violation of KISS. There's no good reason, but several downsides, to mix markup, code, and interpreter state like this.
I would also be quite hesitant to claim that Knuth, who literally pioneered modern LR parsing theory (and deservedly got a Turing award for it) somehow lacked knowledge or awareness of cutting edge parser design techniques.
Finally, there is absolutely good reasons to design TeX as Knuth did. That TeX withstood the test of time for over three decades is a perfect testimony to that.
https://tex.stackexchange.com/a/602950
I'll quote the key part here in case that link stops working.
> TeX has two programming systems, the "mouth" (which does macro expansion essentially) and the "stomach" (which typesets and does assignments). They run only loosely synchronised and on-demand.
> For programming purposes, they are a pairing of a blind and a lame system since the "stomach" is not able to make decisions based on the value of variables (conditionals only exist in the "mouth") and the "mouth" is not able to affect the value of variables and other state.
> While eTeX has added a bit of arithmetic facilities that can be operated in the mouth, as originally designed the mouth does not do arithmetic. There is a fishy hack for doing a given run-time specified number of iterations in the mouth that relies on the semantics of \romannumeral which converts, say, 11000 into mmmmmmmmmmm.
> Because of the synchronisation issues of mouth and stomach, there is considerable incentive to get some tasks done mouth-only. Due to the mouth being lame and suffering from dyscalculia, this is somewhat akin to programming a Turing machine in lambda calculus.
> TLDR: the programming paradigm of the TeX language is awful.
Finally, please do not let your fear stop you from enjoying TeX: most users do not do any serious programming and its output is aesthetically superb. If one day you decide to write some tricky macros, there is a vast TeXlore that awaits.
I personally think the adoption macro is the embodiment of non intuitive programming languages and why a modern programming languages, for example Rust is supporting macro is beyond me.
I'm not alone in this regard, D a modern successor of C and C++ does not support macro.
Hopefully one day Walter will write document on "Macro Considered Harmful" to enlighten the subject.
> SILE’s \define command provides an extremely restricted macro system for implementing simple tags, but you are deliberately forced to write anything more complex in Lua. (Maxim: Programming tasks should be done in programming languages!)
These are macros:
>C preprocessing expressions follow a different syntax and have different semantic rules than the C language. The C preprocessor is technically a different language. Mixins are in the same language.
D shows that it's possible to be a very powerful and potent programming language for DSL, etc, without all the conventional macro abuse and misuse.
If anyone asked me how come Python has becoming very popular nowadays, the answer will be it's one of the most intuitive and user friendly programming languages ever designed, and that's mainly due to the fact that it does not support macro [1].
Broadly speaking, if something can generate parametrized chunks of AST on the go, it's a macro. Which is exactly what D mixins do.
As for Python, it doesn't really need macros because everything is runtime. But if you need it, compile() and eval() are there, and they can be abused much worse than any macro facility ever could.
According to D authors Mixin is not macro because Mixin still looks like D language while supporting macro would meant the inclusion of macro (hygienic or not), and depending on the implementation, most often than not will render the language unrecognizable (become non intuitive). Again if someone is purposely writing a DSL in D for generating a new language's AST then that's perfectly fine. Essentially the output is another programming language because it is the intentional product of the exercises but the original programming language is still intuitive.
I suppose you can have runtime macro like VBA but Python language designers refused to incorporate it due to issues as mentioned beforehand.
I still don't see why D's mixins aren't hygienic macros. And the D documentation doesn't make that claim, either - it only highlights the difference with C preprocessor, not macros in general.
OTOH I can't think of anything in VB6/VBA that resembles macros, runtime or otherwise?
Mixin is not macro according to D authors. If you dig into Walter's past posts, perhaps you can find better explanations there than mine.
You can have macro in any runtime for example Ruby language, and that's the main reason for the maxim "Rails is not Ruby". That's also the reason Ruby is so powerful and to create a similar feat in Java will be close to impossible [2]. But at what cost? That's why Python now is much more popular even though Ruby is much more powerful. You can, however, create Rails clone in D language without all the macro nonsense [3].
[1]Ruby Macros Explained:
https://codeburst.io/ruby-macros-18bb67e051c7
[2]Stop Designing Languages. Write Libraries Instead:
http://lbstanza.org/purpose_of_programming_languages.html
[3]Diamond CMS:
Edit: here it is: https://tex.stackexchange.com/a/37597
>"The % (between \end{subfigure} and \begin{subfigure} or minipage) is really important; not suppressing it will cause a spurious blank space to be added, the total length will surpass \textwidth and the figures will end up not side-by-side."
The shifts in behavior with an added newline (or absence thereof) or a comment, as above, simply wouldn't happen if the language ignored whitespace outright.
Of course, requiring tags for newlines/paragraphs would make raw TeX as un-readable as raw HTML. It might be worth the trade for the predictability, but for the era in which TeX was written, one without IDEs or LyX, I suspect that Knuth made the right choice.
The TeX Book is a work of art that we should all aspire to when writing software.
But once you know how to run TeX at your "site" (as they would have called it then) all the stuff on the language (that is, practically the whole book) is still good.
> one of the things that TeX can’t do particularly well is typesetting on a grid. This is something that people typesetting bibles really need to have. There are various hacks to try to make it happen, but they’re all horrible. In SILE, you can alter the behaviour of the typesetter and write a very short add-on package to enable grid typesetting."
Alas, I don't know what that means. https://www.oreilly.com/library/view/latex-cookbook/97817843... appears to give a clue:
> For two-sided prints with very thin paper, matching base lines would look much better. Especially in two-column documents it may be desirable to have baselines of adjacent lines at exactly the same height.
but I don't have access to the full content.
TeX usually tries to adjust the space between lines so that a single typeblock is more visually appealing. Unfortunately, doing this adjustment independently for side-by-side columns leads to output where the text lines don't line up, and that is way less visually appealing.
Why is that particularly relevant to typesetting a Bible?
Less obvious is that bibles are often printed on thin paper to fit a big document in a small space, so having things line up on one side of a leaf to the other is more important than in most other books.
One of the examples I found is https://www.behance.net/gallery/91186859/Typography-of-the-B... .
Is it to have parallel translations keep the verses side-by-side, as the first image shows between German and Greek?
I did mention the thin paper example earlier. :)
Read the TeXbook. No, really, it is a beautifully written text. The program it describes is quite elegant. Most of the complexity comes from modern LaTeX packages, not from the core tool.
Knuth initially added macros to TeX just as a convenience to save some typing here and there, and the tradition of (ab)using that macro ability as a language in which to program, inaugurated in a small way by Knuth himself and later taken to dizzying heights by others (such as Leslie Lamport with LaTeX, and hundreds/thousands of assorted "package" authors since, and the LaTeX team currently), is IMO a mistake.
If you use TeX just for typesetting, it is quite a nice program that does what you want, and has very good error messages (yes) and debugging output.
That's the only thing I ever found it to be good for. Even then, it's a PITA. The world would not be a poorer place if it disappeared and was replaced something that that separated style definition, markup, and content in a reasonable fashion, and used modern language techniques (just a CFG would be a start) for the first two.
Also modern LaTeX is supposed to not use computer modern anymore, but the Latin Modern (which imitates computer modern but fixes some of its shortcomings.)
If some billionaire wanted to 'move the needle' and change society for the better I think building a truly better latex would be worth it. Each year some of the most brilliant people in the world sacrifice on the altar of latex, that could instead be spent on more productive research.
Perhaps. But I'd wager that more time is lost to the bureaucracy of universities and grant administration than to fighting LaTeX quirks. I'm not sure that "fixing" (La)TeX would result in a noticeable improvement in research that could be done.
maybe the graphical environments eliminate some of the pain around embedding images, or using wrapped figures.
but there is no way they manage to get around the fundamental lack of composibility. that sinking feeling you get when you just add one more thing or try to put an X inside of Y and the whole thing falls over in a pile.
And as someone who's even written packages: ...you're going to end up needing low level TeX, and TeX is insane, don't try to write packages. It's a world of hurt. (But then most people will never need to write their own packages)
Can you talk more about what you mean with the lack of composibility? Or of course link to some page(s) that makes that point?
We have PDF as de facto media for handling output and text files for holding the source files. Better TeX means we are not using an archaic madness to go from Txt to PDF but able to utilize modern IT tools.
TeX is just old and died decades ago. Academia is holding it as a hostage because papers...
Disclaimer: ex-academician and a low key maintainer of TikZ
Edit: TeX as a language is horrific exactly what you would you expect from a computer scientist so it's not that Knuth is to blame, it's the rest that didn't get his vision such as academia still leeching off of it instead of using taxpayer money to generate a proper tool
Equal disclaimer: low-key maintainer of ucharclasses.
The fundamentals of TeX's typesetting are amazing, but everything else is just bolted on. You'd probably want to rebuild it with its own, custom language, perhaps inspired more by modern XML/markdown rather than the old TeX language.
The biggest problem that we're going to have moving past TeX is that it's something of a standard for math representation in text. MathML never _really_ took off. Re-training millions of people skilled in typesetting math is going to be _tough_.
Any contender is going to face a daunting task of explaining why adopters should throw away 50 years of collective work of thousands of people that went into TeX packages in CTAN. Math notation is at best 1% of the reasons why people use TeX. Packages are.
I'm not sure if we're getting hung up on the word "complex", but here's 60+ pages on paragraph line breaking alone:
http://www.eprg.org/G53DOC/pdfs/knuth-plass-breaking.pdf
Digital Typography is 700 pages long (not everything is going to be TeX fundamentals, but...):
https://archive.org/details/digitaltypograph0000knut/page/n7...
The book "Digital Typography" is, somewhat misleadingly, not a book that was written about digital typography, but simply a collection (https://cs.stanford.edu/~knuth/selected.html) of several papers that Knuth wrote about TeX and Metafont and some related topics, over several decades.
Maybe the several hundred pages of the "Computers and Typesetting" series would be a better example for your point: https://cs.stanford.edu/~knuth/abcde.html
In any case, I agree with the comment you're replying to: if you use TeX itself without dragging behind you the entire ecosystem and all the packages (in particular, use plain TeX rather than LaTeX) you'll see it's not really complex, and the implementation is fine. Especially today, with things like LuaTeX scripting (and, say, opTeX) it is feasible to bypass LaTeX and packages and do things yourself — that is, unless external circumstances require you to use LaTeX and packages etc, which is still often the case (e.g. submitting papers to journals).
Feel free to come up with something new that supports Markdown as a subset or overlaps with Markdown.
Markdown as-is is lovely.
I think the best future option is going to be djot[0]. It is being created by the author of pandoc, who might possibly be the most qualified person in the world to appreciate all if the nuances of marking up text and parsing it.
The rationale page can do a better job than myself. https://github.com/jgm/djot#rationale
: echo "word**withboldtext**inside" | pandoc
<p>word<strong>withboldtext</strong>inside</p>
: echo "a*?*b" | pandoc
<p>a<em>?</em>b</p>(which is why markdown still doesn't officially have any maths support, only certain non-standard flavours do)
As far as I could get, Sile is exactly that.
I'd prefer one of these:
- A human-writable text format that's displayed in a formatted way, eg TeX or MathJax
- A binary format that represents structs and enums in code with a clear documentation of how it's packed and unpacked. That can be formatted for reading, and perhaps has a TeX-like API (or is just constructed programmatically using an API)
MathML, in contrast, provides neither the speed and small size of a binary format, while being unreadable and unwriteable.I am still fond of the Scribe[1]-like syntax, but then again, I also liked it in Texinfo or Borland Sprint.
It appears that there is no development in ConTeXt for a while now. I haven't checked it, but it seems that they are working on LuaMetaTeX.
The entirety of the web world with their rube goldberg interactions between way too many divs isn't exactly giving me inspiration that modern approaches are going to be an answer here.
Typesetting is hard. But at this point why not use semantic html and a specially crafted css?
\m[blue]int\m[] x
Of course, you would use some kind of pre-processor to
make this less of a PITA, but it's not inconceivable.Never went particularly deep into it though.
SATySFi (pronounced in the same way as the verb “satisfy” in English) is a new typesetting system equipped with a statically-typed, functional programming language. It consists mainly of two “layers” ― the text layer and the program layer. The former is for writing documents in LaTeX-like syntax. The latter, which has OCaml-like syntax, is for defining functions and commands. SATySFi enables you to write documents markuped with flexible commands of your own making. In addition, its informative type error reporting will be a good help to your writing.
The main problem is that a lot of the documentation is in japanese.
Mathematical syntax is complex, and the typesetting often becomes unreadable outside trivial cases. Using Unicode won't help much, especially because it's easy to confuse similar-looking but unrelated symbols. Proper syntax highlighting might help, but I've never seen a tool doing a good job with it.
The next obvious improvement is to allow (not require — allow) me to write α instead of \alpha, ∈ instead of \in, and ≤ instead of... whatever the code for that was.
And for the record the native SIL syntax resembles TeX at first blush but is actually vastly simpler because it is regular and uses a smaller set of possible syntax variations. If you don't like the look or feel of the syntax you can use something else for your source format and provide your own reader that generates a document AST.
As far as speed, even for simple stuff the real time formatting gets in the way constantly compared to just writing text in a latex file.
Does Acrobat have any typesetting capabilities? I thought another app did that and just finalised through Acrobat.
But they are life-savers when that's your situation.
Does it generate diffs/patches that ultimately can be used to reconstruct the document? Can I take a document and apply a patch and get exactly the same as you get? Can I reorder and edit patches? The answer is no.
(I am aware of the difference between LaTeX and TeX; I likely haven't used TeX directly but the comparisons here should still work fine.)
There is a huge difference in workflow between generating a document from a complete code file and iterating on a final result with incremental tweaks. That is why I cannot use Acrobat, Word, Google Docs, etc. for creating papers that I would use LaTeX for. It is just not the same, and it is not the workflow I desire.
But once the text is done, it's really nice to work with, and you never "accidentally break the whole document by adding a table" - at least not irrecoverably.
Also, pushing figure around to nice places can be so frustrating that people settle for less than perfect layout.
Another pro is that I can use git for tracking changes
Here's what it looks like if you're interested:
Edit: I initially said none of the upstream maintainers call it a rewrite, but I was wrong because the original author did in fact use that turn of phrase on several occasions including in the manual and in a talk.
I agree though it's not an ideal way to describe it because it has some many differences too. It's not a port of TeX to a different language. As you suggest TeX-inspired is much nearer the mark. It's a from-scratch effort at addressing roughly the same problem space. It does re-implement some of the same algorithms. One of it's input syntaxs resembles TeX (although the resemblance is not even skin deep). But no it is not a rewrite, just a new take.
I've been thrilled with TeX from my first usage to the present. Of course, when I want to write some math, I use TeX. I also have a collection of TeX macros I wrote for verbatim, cross-references, annotation of figures, foils, ordered lists, simple lists, etc. TeX is my standard for any good or better quality writing, e.g., serious letters. For my last published paper in applied math, right, I did that in TeX (I don't like to publish -- seems financially irresponsible). The journal was very happy to get my TeX source. I also included the source of the few macros of mine that I used in the paper, and the journal was also happy to receive those. For the core, original applied math for my startup, I wrote that in TeX.
Net, I really like TeX.
For "parsing" TeX as in this Hacker News thread, I have no idea what that might mean or why I would want to do that. TeX, just the way Knuth designed and documented it are just fine with me. For me, TeX solves a big problem; I'm just thrilled to have that problem SOLVED; and I have no desire to invest time or energy in another solution to the problem.
Also relevant, in praise of simple text:
Uh, my most heavily used program is my favorite text editor, Kedit. To me, the most important view, at the most important level, computing and/or computer usage is just simple text, i.e., is still like old fashioned typing. So, I have 100+ macros I wrote for Kedit, and some of those help with typing TeX.
"Simple text" for computing? Yup. Early in my startup, I decided to go with Microsoft and Windows instead of Linux or other versions of Unix.
For the 100,000 lines of code (~24,000 programming language statements and the rest comments or blank lines) for my startup (a Web site), I wrote that with just Kedit. Once I tried to use Microsoft's Visual Studio, and for just the start on just a first program in Visual Basic .NET, I got a big directory of a lot of files I didn't understand. I could see I would need to invest a lot of time and energy into getting the software for my startup to run via Visual Studio and could see no important reason why I should make that investment. I've been happy with that decision: For Visual Basic .NET, I type that in via Kedit.
In praise of Visual Basic .NET:
I know; I know; according to a lot of people and industry norms, I'm supposed to use C, C++, C#, or other programming languages in the family of C and certainly nothing called basic. Well, as I recall, the original C documentation admitted that the syntax of C was "idiosyncratic". E.g., it appeared that
i = ++j+++++k++
was legal -- increase each of j and k by 1; add them; assign the result to i; then again increase each of them by 1. Once I tried this statement on two C compilers, and they didn't agree on the results! So, it seemed that the syntax was too "idiosyncratic" even for the compiler writers! That was enough to warn me to stay away from C, and mostly I've been successful at that!
Then much of what I like and want from Microsoft is their .NET software and especially their documentation. As far as I can tell Visual Basic .NET (VB.NET) is a perfectly good way to make full or nearly so use of .NET and the CLR (common language runtime or some such), exploit their documentation, etc. And with basic I get traditional programming language syntax. To me, that was enough evidence -- decision made. Problem solved. TODO list item checked off.
Yup, at one point my VB.NET calls some C code, right, LINPACK -- apparently the way to do that is to use "platform invoke", and I did and it works fine. I'm still happy as a clam. So, right, I type in my VB.NET code with just Kedit -- been thrilled! VB.NET and Kedit -- happy as a clam!
For TeX, I type that in via Kedit. For email, Kedit. For forum posts, usually Kedit. For working with collections of files in the Windows file system NTFS (abbreviates maybe New Technology File System), Kedit. For my log of food, exercise, sleep, ..., Kedit. Recipes, sure, Kedit. Shopping lists, Kedit. A very important file, my most important, of various facts, short notes, references, and links, right, Kedit. Generally I like to use just simple text, and, thus, Kedit for as much as possible.
For this thread and its
"Sile: A Modern Rewrite of TeX"
I looked at it and could make no sense out of what it was or what it was for. There was something about hyphenation in Turkish! Looks like I should stay with TeX!
FWIW, in FF, Win the PDF fonts have aliasing issues when not zoomed in, Sumatra PDF looks much better.
A big reason why people love TeX is the quality of the output.
(unless it just released a new version of course, in which case the title should probably be about that)