What's new in TeX
lwn.net
lwn.net
You would get better mileage with tools such as LaTeXML[1] or TeX4HT[2], which cover very substantial subsets of TeX, if not yet 100% feature-complete.
My own take on "Should LaTeX survive on the web?" is "yes, if it evolves to meet the paradigm shift". I wrote a detailed version of that position at [3].
- [1] https://en.wikipedia.org/wiki/LaTeXML
- [2] https://en.wikipedia.org/wiki/TeX4ht
- [3] http://prodg.org/blog/latex_today/2015-03-16/LaTeX%20is%20De...
However, LaTeXML is not 100% feature-complete at all (sorry if this is not what you were saying). For anybody not knowing how to improve LaTeXML itself (eg. by adding support for missing LaTeX packages), the only way to get it working is by using it from the beginning to avoid using anything that would not work out of the box in LaTeXML.
I understand Knuth wrote a book, but reading through it has only made me wonder more about the bizarre choices.
As it happens, this is yet another Turing-equivalent form of computation. The closest analogy I have is programming in Lisp, but with the constraint that you can only use special forms and macros, no functions are allowed.
I've sometimes wondered if Lisp-like languages wouldn't be a much better way of writing documents than our current markup languages of choice. If I'll ever seriously try to pick up Racket, it'll probably be to give Pollen a spin and see how it is in practice.
Whereas in Latex you have a bunch of (sometimes conflicting) packages to solve various problems, Context is a coherent unit. I didn't need external package for anything, not even for complex stuff like wrapping text around figures.
The downside is, and the reason why I couldn't keep using it, is that the markup commands are incompatible with Latex. Submissions to journals accept Latex only :(
That said, however, if you are interested, a good place to start is by reading the user support list http://www.luatex.org/support.html .
But the final version, for which there is a plan, has not shipped or been written. Writing it is a very ambitious task and takes time, and a little money (through support from, e.g., the TeX Users Group). Things take time.
Two approaches I see as plausible:
1. Use TeX as a "document assembly language" of sorts, a pure back-end rendering target not exposed to the user, and do your document markup/scripting in some other front-end language. Pandoc [1] is a start at building infrastructure for this. You can then build a new renderer too if you want, to remove the dependency on TeX, but at least in the meantime you have a pathway to producing quality output right now.
2. Build something on top of HTML+CSS (maybe also +JS). Advantages of this would be that it'd make it relatively easy to target the web with the same document source, and open-source web browsers have already sunk a ton of development time into rendering. CSS even includes a bunch of print-oriented features that in principle provide markup for most of what you might need here, though browsers for obvious reasons haven't tended to prioritize those parts. wkhtmltopdf [2] is a project aiming to build a to-print or to-PDF document workflow on top of Webkit. However imo the results are still not near TeX-replacement level. I believe the gold standard currently, if you want high-quality print out of HTML+CSS, is the proprietary PrinceXML [3].
I write a lot of TeX, and I have never found that those things are the fault of the standard software; rather, they are my fault.
The absence of a better alternative in the domain, even though attempts to serve the domain by other tools (including ones using more "modern languages") have come and gone over the years.
TeX isn't theoretically ideal, its just practically very good and the expected cost to benefit ratio of a ground-up replacement is very high.
TeX is really, really amazingly powerful. It can do almost anything a typesetter could want to do, fairly easily, and it can do just about everything, one way or another. And its output is heart-achingly beautiful. Sadly, the code necessary to achieve that output ranges from…heart-achingly beautiful to heart-breakingly ugly.
There are other projects out there, of course. I do think that TeX & LaTeX are close to a local maximum, if not al the way there.
XML, in comparison, is a booger joke.
Funny you say that since TeX's version scheme (in part) is that it approaches pi. And IIRC will be pi upon Knuth's death.
TeX is incredibly powerful, but it's also incredibly idiosyncratic. I also personally think Computer Modern is an ugly font.
No, it's really not. XML is a markup language, not a data-encoding language. JSON, S-expressions, ASN.1 &c. are all data encodings; TeX, LaTeX, HTML and XML are markup languages.
> TeX is incredibly powerful, but it's also incredibly idiosyncratic.
Agreed.
> I also personally think Computer Modern is an ugly font.
Eh, it's not great on-screen, but it looks pretty good on paper. But TeX & LaTeX have supported multiple fonts since the beginning.
A dried up booger on the floor.
If you replaced TeX, LaTeX itself would need replacing, as would every single LaTeX package and class that you ever want to use, ones like microtype and stuff, would have to be rewritten.
To be honest, the multi-pass deal isn't that bad, but the macro expansion system is crazy complicated. Every once in a while after working in LaTeX I'll get the feeling I understand it, but that feeling inevitably dissipates after ten minutes or so.
Making a "new" TeX probably wouldn't be that hard - but it's also something that wouldn't be that useful. I would very much like something that's both simpler and also keeps some of the lessons learned/implemented (word spacing/splitting, page layout, page breaks etc).
As for other "tools in the same space", I do like pandoc a lot. I want to like python's ReST (Re-Structured Text) - but that's a package I feel is in need of a rewrite/redesign. Many good ideas there - but figuring out how to take a simple document and produce simple, modern (preferably somewhat semantic) html for example -- or to produce a decent looking PDF without needing all of LaTeX/Texlive on hand isn't easy.
Rewriting ReST tools would be a lot of work, but I think if one didn't try for 100% backwards (output, plugin) compatibility it might be worthwhile.
The astute reader will notice that ReST/Pandoc deals with structured documents, and not really layout for paper/screen (both use TeX/LaTeX as an output target/pipeline). I don't know of anything that comes close to TeX/LaTeX for "rasterized" output.
On the other hand, I also don't know of any package/combination that'll make TeX/LaTeX produce anything but messy, 90s-style html -- that generally looks awful. Even if you were to try and force a modern set of CSS down over the resulting mess. If anyone knows of a modern hypertext package for TeX/LaTeX or some similar tool, I'd be happy to be proven wrong.
XSL:FO at one point seemed to have aspirations in that direction...
I have a 450 page book that as part of the compile generates a 200 page answer answer manual. They are full of math, hyperlinked cross references, figures, etc. They use tons of packages including amsmath, hyperref, and many more. Compiling takes perhaps 15 seconds. This means compiling twice, to resolve references.
I was looking for a workflow that would produce PDFs and HTML, and I reached the same conclusions as the article.
I was hoping they had some sort of solution.
Examples:
- http://www.seas.upenn.edu/~cis194/lectures.html - http://www.scs.stanford.edu/11au-cs240h/notes/
I've wrote some basic tutorial about tex4ht configuring [1], lot of information can be found when you search TeX.sx for tex4ht tag [2], for example how to use Mathjax for math rendering[3] or how to include Javascript libraries and some responsive CSS[4]. You can ask tex4ht related questions on TeX.sx or on it's mailing list[5], there is also a issue tracker[6].
[1] https://github.com/michal-h21/helpers4ht/wiki/tex4ht-tutoria...
[2] http://tex.stackexchange.com/questions/tagged/tex4ht
[3] http://tex.stackexchange.com/q/265815/2891
[4] http://tex.stackexchange.com/a/239944/2891
I've had some success with ipython/nbconvert too (which in a roundabout way works like Markdown+pandoc, but doesn't use pandoc -- but outputs both reasonable(ish) html, and decent PDFs (via LaTeX).
https://ipython.org/ipython-doc/1/interactive/nbconvert.html
TeX has somewhere around 325 primatives, and one of the most important is the \def primative used to define macros. These primatives are used to define additional macros, hundreds of them, available in different so called formats. A basic format known as Plain TeX includes about 600 macros in addition to the 325 primatives. LaTeX is another format, the most widely used, but there are others, like ConTeXt, that are also very capable. Each of these extend TeX's primatives with their own macros resulting in different kinds of markup language.
TeX's primatives are focused on the low level aspects of typesetting (font sizes, text positions, alignment, etc.). LaTeX provides a markup language that is focused on the logical description of the document's components: headings, chapters, itemized lists, and so forth. The result is a system that does simple things easily while allowing very complex typesetting to be performed when needed.
In addition to the TeX core primatives and the hundreds of commands (implemented as macros) in a format like LaTeX there are additional packages, classes, and styles that are used to provide support for any conceivable document. LaTeX has a rich ecosystem of packages. Typesetting chess? There's a LaTeX package for that. Complex diagrams and graphics, there's a LaTeX package for that. Writing a paper in the style of Tufte? Writing a book? or a musical score? or building a barcode? there are packages for that. The documentation for the Tikz & PGF graphics package is over 1100 pages long! The documentation for the Memoir package is 570 pages.
The amazing thing is that all of this is built out of macros. Diving into this, and once one needs to customize the look of a document it's inevitable, you find yourself in a maze of twisty little passages.
Once upon a time, while writing assembly language for large computers, I enjoyed writing fancy assembler macros. I was facinated with Calvin Moore's Trac programming language based on macros and Christopher Strachey's General Purpose Macrogenerator. These were early (mid 1960's) explorations into the viability of macro processors as means for expressing arbitrary computations. Reader's interested in trying out macros for programming can try the m4 programming language (by Kernighan and Ritchie) found on Unix and Linux systems. m4 is used in autoconf and sendmail config files. Yet, TeX macros are in a whole other dimension.
All of these powerful macro systems have one thing in common: parameterized macros can be expanded into text that is then rescanned looking for newly formed macros calls (or new macro definitions) to expand as many times as one wants. This isn't just an occasional leaky abstraction; it is programming by way of leaky abstractions. Looking at TeX packages is some of the most difficult programming that I've done. It's unbelievably impressive what people have come up with (e.g. floating point implemented via macro expansion in about 600 lines of TeX), but it's also unbelievably frustrating to program in such an environment.
The LaTeX3 project is an attempt to rewrite LaTeX (still running on top of the TeX core). Started in the early 1990's it is still not done. I think its just that they are mired in a swamp of macros. They do have a relatively stable set of macros written, with the catchy name expl3, that are intended for use when writing LaTeX3. Here's a sample
\cs_gset_eq:cc
{ \cf@encoding \token_to_str:N #1 } { ? \token_to_str:N #1 }
This is described in the documentation as being a big improvement over the old macros and "far more readable and more likely to be correct first time". I can't wait.I think LaTeX is absolutely without peer, but I wish improving it's programming method wasn't so daunting. I keep toying with starting a project to do just that, but so many others have tried and failed. It's disheartening.
Links:
[TRAC] https://en.wikipedia.org/wiki/TRAC_(programming_language)
[GPM] http://comjnl.oxfordjournals.org/content/8/3/225.full.pdf
[m4] info pages available on Unix and Linux
[Tikz & PGF] https://www.ctan.org/pkg/pgf?lang=en
[Memoir] https://www.ctan.org/pkg/memoir?lang=en
[expl3] https://www.tug.org/TUGboat/tb30-1/tb94wright-latex3.pdf
[1] https://davidar.io/TeX.js/ [2] https://github.com/davidar/TeX.js/issues
Some example output can be seen in the manual http://manual.softcover.io/book/softcover_markdown#sec-embed...
The softcover tools (https://github.com/softcover/softcover) are based on code developed for creating the HTML for http://tauday.com/tau-manifesto and http://www.feynmanlectures.caltech.edu/. Internally Softcover uses https://www-sop.inria.fr/marelle/tralics/ to convert from LaTeX to XML/HTML.
P.S.: I'm not affiliated with softcover.io in any way.
And it looks like it comes with the standard Ubuntu installs. See man groff_mom.
In the article it is stated, that it only supports serverside rendering.
For the multiple output formats, at work the documentation teams use Dockbook and DITA.
Of course they don't write XML by hand in such cases, rather use tools like oXygen, XMLmind, CORENA Studio among others.