LaTeX was not built for the web
authorea.com
authorea.com
That doesn't mean that LaTeX is worthless on the web. The gluing algorithms that Knuth used for creating sentences and paragraphs do (for the most part) work, even in documents without pages. In all honesty, the web could benefit greatly from LaTeX, pagination aside.
I'm glad that the author seems to have a practical perspective on this: >We think LaTeX is still the best programming language to tell a computer how to place text on a page. But the TeX project started pre-web, in 1978, and its scope and function are tightly linked to the printed page, not the webpage.
He goes on to give an example of constructing a table in LaTex which requires (or assumes by default) a suggested location for LaTeX to place the table in the page; something which is currently nonsensical in all but a few web environments. Like the author, I also find CSS appealing for a potential "LaTeX" on the web solution.
However, LaTeX was designed from the very first to be the next standard in type setting and is a turing complete language. CSS is still under debate for being turing complete, and it lacks nearly all of the features that are iconic of LaTeX, except for the ability to format math equations ( and even then...).
I'm not sure what the next standard in typesetting will be, but it will be designed to be device agnostic (with book existing as a supported format), it will target the web and similar digital media, it will be turing complete, and it will be created specifically to fulfill each of those goals, not to have them retroactively attached in a ham-fisted way.
No. The output might be beautiful, but the language is most certainly not. Awful to debug.
I consider a more modern language that compiles to LaTeX and leverage its rendering engine for paper docs (while giving you sufficient control over the output) and also gives you a nice web output more pragmatic and a nicer way to approach the problem.
I assume you're referring to Knuth, but it's important to distinguish between TeX and LaTeX (the latter was written by Leslie Lamport).
The distinction between TeX and LaTeX is actually relevant here.
I know very little about both. Could you elaborate on the relevance?
I'm currently looking for exactly such a solution. Have you found any?
Sidenote: Ideal would be a language that comes with either Word or E-Pub conversion tools (in order to migrate existing word documents over). I say E-Pub because there is a relatively nice migration path from Word to there: Get on a Mac, open the Word in Pages, export to E-Pub. Can also be used to get a sane xhtml output, since it's just a zipped folder with xhtml and some images.
oh, and it has to play well with EndNote™ too. -sigh-
[1]: http://www.lyx.org/
what i really would love to see is some stable fusion of LyX and a web-collaborative-writing program like 'Gobby' (http://gobby.0x539.de/trac/) ... google-docs has been tried, but folks got scared between google-wave suddenly dying and that whole issue of corporate leaking to sinister third parties and corporate scraping
This, this is death to every alternative I look at.
MSWord has the additional problem that it can't be used together with my favourite Version Control System.
Absolutely agree but, as far as I'm aware, this "more modern alternative" does not exist.
Having some XML output seems to make it easier to add CSS rules ?
It can be done. For example Google docs supports headers, footers, and footnotes, but you are still free to decide whether to use the "print layout" or more webpage-like seamless layout for both editing and publishing ("publish to the web").
Or, in a simpler way, do you know Haskell?
So far as I'm aware, that's not the case. CSS is not Turing complete, in terms of what it can calculate in a single calculation. It is Turing complete if you string those calculations together, feeding the output of the last into the input of the next, which can only be done with some external source of events, but that is a larger system than "CSS". A UTM run for up to 1000 steps is not Turing complete.
Turing completeness is arguably a bug, not a feature, when the goal is producing something quickly.
* is a simple and logically consistent as Markdown
* has the ability to embed LaTeX formulas
* has a functional #include for larger documents
* can refer back to headings, equations and citations
* can compile to PDF or webpage
* compiles FAST
and I will give you a lot of money.LaTeX is a horrible language, inconsistent and badly designed and unpleasant to look at. TeX is good at what it does but the whole system is horribly slow to compile documents. Documents will compile fine with some frontends but not with others. Markdown is not feature complete enough.
The scientific community is unfortunately stuck with LaTeX for the foreseeable future. Seems like every day someone asks me how to do something that should be very simple, which inevitably involves loading some obscure package.
I think it's lacking the #include/#input feature, and I'm not sure whether it supports labels and internal references, but apparently it has some form of bibliography.
Pandoc does provide a way to link to section headings, but it doesn't yet have a generic system for autogenerating numbers and producing references to these, like LaTeX's \label{} and \ref{}. It does, however, have a system for creating running example lists: http://johnmacfarlane.net/pandoc/README.html#numbered-exampl.... This can be used for numbering equations and referring back to them, but it is considerably less flexible than LaTeX.
Pandoc has extensive support for automatic citations using CSL stylesheets: http://johnmacfarlane.net/pandoc/README.html#citations. BibTeX and BibLaTeX files can be used as the database, or YAML citations can be included in the markdown document itself.
LaTeX formulas can be embedded in markdown. They can even use macros defined in the document. Formulas can be converted to MathML or native Word equation objects, depending on the output format.
Pandoc can also be extended using "filters" that operate directly on the parsed AST. Here's an example of a filter that finds tikz diagrams and converts them to embedded images that can be displayed on the web: https://github.com/jgm/pandocfilters/blob/master/examples/ti....
At a minimum, it does markdown with embedded LaTeX and outputs pdf.
It admittedly doesn’t compile to HTML and the compile-time can be on the order of minutes for very large documents, but are these really deal-breakers?
If you really don’t like LaTeX, have a look at org-mode, which seems to be able to fulfil most, if not all, of your requirements (not sure about compile times).
I was doing an algorithms course with lots of mathematical formulas last year. Org was a godsend for completing the homework assignments.
Send an email to the address in my profile and I'll fast-track it.
I can see how you might not like the language itself, I don't either. But inconsistency is rather orthogonal quality. I think it is actually pretty consistent, which is respectable in its own right.
Obviously the big step is for it to be able to output LaTeX so that the equation capabilities and such will be directly exposed, but Asciidoc is pretty awesome already in my experience.
I'm not sure why you want the ability to embed LaTeX formulas. LaTeX formats math gloriously, but its syntax is not any less horrid for math than it is for anything else. I'd rather write "sum(i <- 1..10, frac(t_i, sigma^2))" than "\sum_{i=1}^{10} \frac{t_i}{\sigma^2}" but maybe that's just me.
This table command instructs TeX to put the table in the page, here, where the table is declared (h) AND at the top of the page (t).
PrinceXML, an XML/HTML + CSS to PDF renderer, can do this using a "float:top" style: http://www.princexml.com/doc/9.0/properties/float/
Also, before they abandoned their bespoke browser engine, Opera released an experimental build that could render web pages as paged media (see http://dev.opera.com/articles/view/opera-reader-a-new-way-to... ). It, too, supports floating to top (as vendor-prefixed "-o-top").
Paged media is as alive as ever (it's just moving from dead trees to tablets), so I wouldn't count this stuff out just yet.
http://andreimikhailov.com/slides/bystroTeX/slides-manual/in...
The main application was supposed to be slides, to replace the Beamer. But I think it is good for generic math-oriented web publishing. It does not work on Windows, though. Comments/suggestions/testing are welcome!
http://arxiv.org/abs/hep-th/9907164/Eq_4_4
what is supposed to happen?
This does require whoever generates the pdf to include the labels, but then so does html. It shouldn't be too hard to generate a reference to any named equation in a file.
What exactly is your problem with current PDF readers? PDF.js is pretty nice for browsers, Preview.app is standard on OS X, Linux has several that were fine and I know that Windows has many that are considered good alternatives to Acrobat Reader. Even phones and tablets do a good job, though e-readers do not.
https://github.com/amkhlv/pdfviewer
I started that fork because I could not find a viewer which would keep vertical position, and jump back after following an internal link, and have bookmarks with charhints.
There is a conceptual problem, however: as I said, I feel that PDF is getting deprecated because of insufficient demand. If this is true, then it does not make much sense to invest effort into TeX + PDF. Maybe I am wrong. Sometimes I think that I am simply allergic to TeX :)
If there was a reliable way to advise a PDF reader to ‘open document X and jump to reference Y’ and to specify both X and Y, then I am sure you would have no problem using hyperref to create a link to X#Y.
Interesting. I just discovered the excellent online book "Practical Typography"[0], which the author prepared using a system he created that he calls "Pollen," and which he wrote in Racket.
For layout, though, HTML with CSS is the way to go if your prime target is web.
I've been thinking about this for almost a year. I believe as well that webview-first is the future of technical publications. The problem is how to either 1) get traditional publishers to adopt newer and better technology (a lot of publishers are still using systems seemingly from 90s that doesn't even support features that have been stablized in TexLive for years); or 2) build new publishers that gain enough reputation fast enough, so that the academia would consider them as good communication channels.
Personally I like the second way better. It's more convenient (and easier to think out of box) to start from scratch than to change an existing system (by system I mean organizations, publishers, rather than a computer system).
Academia is somehow like a trust chain. People tend to follow reputable researchers/professors. If a platform can get most reputable researchers, it can be adopted soon.
Along with mendeley and peerj there's the makings of little cluster of companies trying to innovate in the academic publishing space here in London.
It's possible e-readers will someday obsolete print entirely, but I personally still find it difficult to read longer-form stuff on a screen, so I'd like to see better options for print layout of web documents. Technologically these are possible, e.g. PrinceXML shows quite a bit of print-layout stuff you can do with the web-technology stack (though it's unfortunately proprietary), but the bits and pieces don't quite plug together well yet.
(sorry for the bad knuth toolchain jokes)
Some time ago, I thought about integrating LaTeX into user-facing software, not to render an entire web page but to generate a few formulas and images containing formulas in real-time or close to real-time. It breaks my heart to say that it seems impossible to do so, because LaTex is soooo slow.
The codebase is written in an obscure language (that Knuth created for the purpose?) and would have to be converted to something optimizable such as C in order to become faster.
Since I won't be re-implementing LaTex soon, I'll stick with MathJax for simple formulas and must look for something else if I want to create a complicated formula-containing image.
Also, having had some bad experiences with to-C translated code (looking at you, Matlab), I'd say that knowing that Web2C exists doesn't cause me great optimism.
http://tex.stackexchange.com/questions/36/differences-betwee...
But still LaTeX sits on top of that, so it probably doesn't move you along very much.
If you're willing to eschew the luxuries that PDFTeX provides (microtype and pdf output) and drop back to DVI, it is possible to interpret the DVI yourself in real time using the IPC hack (which allows dumping of dvi on each page flush). The file format is specified in the dvitype manual[1].
Unfortunately, once you've processed the DVI output, you'll discover that many modern TeX fonts do not render, because they are virtual fonts and you will need to implement these. You'll also have to process the TeX font metric files.
Once you've done all this, you can easily get TeX to produce in excess of 1000 snippets per second. The rest is up to your rendering backend, but you'd have to be doing a lot wrong to end up with less than 100 per second.
[1] http://texdoc.net/texmf-dist/doc/generic/knuth/texware/dvity...
You almost certainly want to use a better font than Computer Modern, but that's easy.