Have you ever asked yourself “how did research get done before LateX?”
twitter.com
twitter.com
The annual NSF physics budget is roughly a quarter billion dollars. This only needs to make NSF-funded physicists 1% more productive over a decade to justify spending $20M, and the entire world would benefit.
EDIT: Lots of commenters worry that this will end up like various closed-source for-profit tools like Google Docs. But for a closer analogy they should look at some of the excellent free open-source science software being produced with philanthropic funding such as Zotero. Physicists will only adopt it if it makes their lives easier (and maybe not even then...).
EDIT 2: There are some pretty strong parallels between the commenters suggesting this can already be achieved with homegrown tools and the infamous HN Dropbox comment.
Ask Knuth to name one person who has the clearest vision for the future of TeX, and give that person a stipend to work on it for the next ten years, or for life, and you might actually do some good.
And before someone asks, yes, I feel similarly about GitHub, and wasn't the least bit surprised when it sold itself to the self-declared enemy of open source of previous decades (now reformed, of course!).
Yes, these things are convenient, they have great UX nudging you toward what serves the interests of whoever owns them. I'm not arguing against the usability of Google Docs but against transplanting that model to a community built on entirely different principles.
> evil science publishing parasites (you know the names)
Not being part of the scientific community, I don't - but I trust you that they exist!
The comment I was replying to specifically called out Google Docs as bad, so that was what I was referring to. I agree that they appear to have mostly-similar features (though, to me as someone who doesn't use either as a Power User, GD seems to be more focused on online collaboration, and MW on local editing, the output of which is then shared)
Then try to modify it as little as possible while improving things.
And add better support for concepts like page numbers, page headers, page footers, which while theoretically supported by HTML and CSS, aren't well supported in practice.
Given that many people use Python now for scientific purposes, and it's a straightforward language, to add extensibility to this Latex replacement, maybe just use Jinja templates and Python packages.
What I'm describing is not do-able by one person, and the result of refusing to coordinate a team around the project (and instead having a hodge-podge of tools built by various people) is the current LaTeX ecosystem mess.
If it ain't broke, don't fix it.
If you accept that premise, the decision has to be: which system would serve the most people the best? I'm inclined to think that a team with a budget of $20M would probably get the closest.
Oh, there are a couple of such candidates, actually, built with very different philosophies: there's XeTeX, intended to graft Unicode and OpenType support onto TeX with as little disruption as possible, such that it "just works" with minimal effort; and there's LuaTeX, which "deconstructs" TeX to give you an infinitely-flexible typesetting toolkit that can do just about anything, if you can figure out how to wield it.
Of course it's a work of love, but it's a tool from the 70s
Clinging to the past might be interesting but in the end it's that, a tool
TeX's downfall is the macro system. It makes it too easy for people to pile on poorly written abstractions that make working with any given document a nightmare because it has its own DSL. But because those packages and macros grew organically by researchers solving problems for themselves they are too practically useful to just drop them. So there are lots of annoying unintuitive quirks that you end up just having to put up with.
Edit: Someone beat me to it by suggesting VSCode as an IDE.
Edit: phone autocomplete typo
```
SRC :=$(wildcard .tex .bib)
Paper.pdf: $(SRC)
pdflatex Paper
bibtex Paper.aux
pdflatex Paper
pdflatex Paper
short: $(SRC)
pdflatex Paper
clean:
rm -f .glo .log .dvi .gls .toc .aux .ist .out .glg .pdf .bbl .blg .lof .brf *.sta
.PHONY: clean
```
I hate that you have to run things multiple times to get references right. Really it should just be `pdflatex Paper` and you should be done.
Every once in awhile I just get references that don't want to stick. Same with images. [ht!] and floatbarriers only go so far. As much as tikz is a pain (and beautiful), I'd definitely invest more time into making image placement and references easier for the average user.
EDIT: code has a lot of wildcards before every extension but I'm not sure how to escape those on HN
I don't use things like bibtext so I don't know how it handles that, but latexmk works perfectly for normal use cases, and ensures that you never have to run it more times than you need to.
This might be a niche opinion, but I think Rmarkdown & Jupyter get us a non-negligible portion of the way there. You can use LaTeX chunks when you need equations, markdown syntax when you need formatting (IIRC this is a pain with LaTeX), and executed code chunks to (e.g.) create and insert a graph. (One advantage of Rmarkdown is that you can have as many languages in a single document as knitr can handle, each its own chunk [0]).
Neither is all the way there yet, but, as Yihui Xie puts it [1]:
> I have a dream that one day all students and researchers will forget what "formatting a paper" even means. I have a dream that one day journals and grad schools no longer have style guides. I have a dream that one day no missing $ is inserted, and \hbox is never overfull.
[0]: https://bookdown.org/yihui/rmarkdown/language-engines.html
I think you're absolutely right. But people seem to value a specific flavour of aesthetics over easy to understand and reproducible research.
I guess that's maybe partly caused by the publication pressure, once you start to pump out papers for pumping out papers sake, there is little incentive to go beyond the superficial.
Rather, this language would be designed to come after writing, in the same way that flagship Desktop Publishing software (InDesign, Publisher) make the assumption that layout happens after writing. (Sure, you can write inside the layout software, but you're not supposed to. They're assuming e.g. a journalist–editor–layout–typesetter pipeline.) You'd have a "layout file" (written in the typesetting DSL) that has separate "non-typeset rich-text" files linked to it via file/URI references (just like Desktop Publishing apps do.)
The person doing layout would never modify the linked source files; if they wanted to e.g. change certain character sequences into filligrees, that'd be something stored entirely as an annotation in the layout file, not as markup in the source file.
In that world, scientists just write the "non-typeset rich text" (for which Rmarkdown / Jupyter would be a great fit); and then it's the job of science journals to produce the "layout + typesetting annotations."
And hey, then journals would actually have a reason to exist! They would be to science as physical magazines are to journalism at this point: a place to find the articles you can already read online, but typeset much better.
Figure sizes are also a function of document fonts and size.
Think of it like an archivist restoring a painting. First, they add a (removable) clear-resin barrier layer; and then they do restoration work on top of that barrier. In essence, rather than the paint they're adding becoming part of the painting, it instead becomes part of a "wrapper" for the painting, which can always be "unwrapped" back to the source.
Except this analogy isn't perfect, because you can only override a painting point-by-point; whereas a layout document can "losslessly" cut a linked source up into chunks and put each chunk in a different place, repeat chunks, insert things between chunks, etc.
A better analogy, actually, might be a DAW project. The input audio samples are immutable; the DAW project takes samples from the inputs, applies changes to those samples, and then sequences the results.
And note that a large part of the point of this workflow (in desktop publishing, at least) is that the upstream can continue iterating on the linked document as the downstream lays it out. Those changes may break the downstream layout, but they usually don't, because the downstream layout has usually been defined by constraints.
For example, in a newspaper article, if the source-text of the article gets more wordy, then up to a point, the layout engine just packs those same paragraphs tighter on the page it already decided they appear on, while keeping the paragraphs that appear on the next page spaced as before, and not allowing any lines from the paragraphs on page A to migrate to the flow-box on page B. But at a certain point, it switches strategies, moves an entire paragraph from A to B, and then relaxes the spacing of A back down. At that point, the layout may have been broken (because now B might be either too long to fit its flow-box, or unreadably tightly-spaced.) But that happens surprisingly rarely, compared to constraint-based reflow "just working" to absorb linked-document changes. (And note that either the writer or the layout designer can aid this process by inserting explicit page/section breaks that add weights to these constraints, to keep reflow from flapping lengths around.)
-----
Note that everything I'm talking about is already how a very domain-specific kind of Desktop Publishing already works: "engraving", a.k.a. sheet-music publishing. The score itself (the thing that contains information that could be converted to MIDI) and the engraving details (the way to present the score to a human) are kept separate in most engraving-tool formats.
You may be amused to learn that there's a TeX for sheet music -- Lilypond [1]. I don't believe it's commonly used in the industry, but it can generate decent-looking output.
It's an interesting idea, but a counterpoint would be: Why all this work for a dying medium? Paper is on it's way out, and there are so many different screen formats (including E-Inc) that dynamic layout is the way to go.
Eh, kinda. CSS requires the assistance of the HTML, though; HTML is "the boss" in the relationship, in that it gets to define where markup boundaries are, declare what CSS classes and identifiers apply to what elements, etc. So "I write the HTML, you lay it out with CSS" isn't really a valid pipelined workflow, if you expect to do anything fancy with your layout.
In GUI desktop-publishing programs, the layout document is "the boss." It's as if CSS could target its rules by byte-slices, rather than needing HTML to hand it ready-made selectors. You can take arbitrary bits and pieces of the linked source, style them how you please, and throw them onto parts of pages as you please.
Consider: in GUI desktop-publishing programs, a "pull quote" is not a second copy of the text in the source. It's just a second reference to the same text-slice of source-text, within the layout.
> Why all this work for a dying medium? Paper is on it's way out.
Layout/typesetting isn't just for books. Posters need layout. Billboards need layout. Business cards need layout. None of these are dying out in the least.
Heck, even for screen media, text within videos (e.g. title cards, ad copy, etc.) has static layout, and so benefits from layout/typesetting.
There will always be a place for tools that allow someone to efficiently describe (or hint) the best way to show text (on a screen, or on a page) for best readability — or best impact.
And even leaving aside the practical uses of text layout, I should note that kinetic typography is also an artform, done for its own sake—and one that often requires the assistance of a layout-constraint engine to achieve it. Do you think you can write something like House of Leaves purely using design tools like Adobe Illustrator? ;)
In a desktop-publishing program, you can just highlight some text in the linked source, and make it bold, and the layout file will remember that that part of that paragraph should now be bold (using a targeting heuristic that will mostly work unless the paragraph is severely changed.)
You can't do that with CSS.
I didn't read the rest of your post after this, because the only purpose of this is to inflame and troll.
LaTeX is suprisingly terrible at producing accessible documents. There are various systems and packages which make heroic efforts to make accessible output, but you have to use them from the start, as they work like you suggest -- they reimplement (or at least heavily modify) various LaTeX packages to make their outputs accessible.
I've had a blind PhD student and the standard of using LaTeX to produce PDFs for academic papers has been the single biggest block in his progress.
PDF and TeX are broken beyond repair in terms of accessibility. The former because it allows the text and image to be split, and very few people care about text as long as it looks alright. And the latter because it's entire document representation model is based around computation. Knuth purposefully made it Turing complete in a misguided attempt at timelessness (no need for a replacement if it's infinitely extensible right?), which means there is no proper way to view a TeX document as "just text data".
I've worked in a library digitising documents, and whenever we encountered PDFs we'd just OCR the damn thing.
We need to throw the thing out, swallow the bitter pill of sub optimal kerning, and replace it with HTML asap.
All that needs to be done is to have a accessible text section of the PDF--then Tex can just include the source of the document minus the Tex commands into the PDF, and PDF readers can have the screen reader work. The text could also just be stored per page, if the blind require to communicate where they found a particular item to a sighted person.
PDF already has exactly that. They separate they have a view layer and a text layer.
But it simply doesn't work, because people can't be bothered to get the text right, once the view layer is acceptable.
Also TeX minus the tags, is not the final text in any way, because it's not markup, it's code.
That's hardly "it's broken beyond compare." I think with accessibility formats which represent fully typeset written documents (as opposed to pre-typesetting, like Latex), it's always going to be an issue because they're mainly produced for people who are sighted initially.
We should simply start shaming programs which export to PDF wrongly, it's not really a difficult thing to get right.
Knuth is one of the most perfectionist computer scientists out there, and he still got it wrong.
The market economics are simply not in the favour of people doing extra work, and no amount of shaming will get that to change.
The alternative needs to be designed with accessibility builtin, and with decidability/markup as a conscious design choice.
We have such a format, HTML. The chances to get academics to use HTML with the coming open access wave are much better than getting them to rewrite all of their TeX templates.
It's more about the social economics --but if you can find me a tool that supports the PDF export wrongly, I'll fix it myself.
> The market economics are simply not in the favour of people doing extra work,
Most of the tools are open-source. Does Word or any of the closed source tools do a bad job of this? That would be the main stopping point I guess.
It's still super trivial for them to implement, and shaming them will get their PR team evaluating the risk of not implementing such a feature. You just have to be successful at shaming them.
An HTML file also doesn't represent a document on paper, most people writing papers wouldn't switch to a electronic-only format at the moment--each page is carefully typeset and designed, HTML is not at this level even today.
https://arxiv.org/pdf/1606.06389.pdf
A cursory inspection seems to indicate ctrl-F is fine here? Or am I missing something?
However if you try to select the text on page 2, you'll see that the layout engine didn't actually place the formula into its own section. Not only is it broken, with the sum signs missing, it is also jumbled up with the text below it, resulting in weird - text broken formula - text - broken formula - mumbo jumbo.
This would make it extremely hard if not impossible to properly follow the paper with a screen reader.
The figures on page 6 also produce weird artefacts on the text layer. In HTML for example you'd simply present the alt-text to the person with the screen reader, but TeX doesn't have such a feature in most vector drawing libs afaik.
The table on page 8 is also not correct in terms of the text layer. If you select it you can see that you get the top half of the leftmost column followed by the rightmost column, followed by the rest in some order.
If I may also direct your attention to this gem on page 9: "Then at least half of allnodeshaveh ̄=lgn =lgn−1(seethepaperforthedefinitionofh ̄),and 2√ thus a potential of Ω( lg n). Therefore, the total heap potential would be lg n). Conjecture: the same construction works for the fancy potential √ √ function, which would give a bound of Θ(n · 4 lg lg n)."
Shall I go on ^^?
I suspect we're using different PDF viewers; I'm using the one that comes with my browser.
The results are somewhat less disastrous than you describe, but still bad. (I'm seeing maybe half the problems you mentioned.) I'm a bit curious which viewer you're using.
I can describe what happens when I copy/paste the stuff you mentioned on pages 2/8/etc., and if you're interested I will, but if you're not interested let me ask a slightly different question:
Rather than try to get the screen reader to make sense of the final PDF, would it be easier to just download the original page source from arxiv and let the screen reader deal with that?
Using the original TeX source for the screen reader is significantly hindered by the fact that TeX isn't a markup-, but a programming language.
TeX is turing-complete by design, making it infinitely extensible, in order to avoid knuth ever having to re-typeset his books. After all if it's a programming language and a new system, problem, style, e.t.c, comes along you can just write a program that deals with it.
But this has horrible effects on render-ability, in order to know what the final document should look like, you need to run it, no way around that, thanks to Rice's theorem.
LaTeX users also generate a lot of their figures with TeX itself, write their own styling or bibliography rules, and write their own custom graphics rendering libraries.
The easiest path is to just give up, render the entire PDF to a 300dpi lossless image, and throw it into an OCR engine. These things contain a buttload of heuristics to generate structure and meaningful text based on visuals, and since we know that humans explicitly (and often only) care about those in TeX documents...
It's pretty darn sad, because in many ways TeX is holding scientific advances back, by eating its own children. My guess is that if scientific papers had branched off of plaintext tools like troff, we'd probably publish papers as machine readable semantically annotated knowledge-graphs by now.
(Also, if you look closely, the summation signs are not gone, they are replaced by the letter P. Which is not helpful I admit.)
Do you think there's any place here for education/advocacy? For instance, everyone who makes web pages knows to provide alt text for images.
If there was a standard package that everyone knew they had to include or else it breaks everything from ctrl-F to copy/paste to screen readers, presumably people would use it, right?
I'm less interested in speculating what would have been if troff had "won", (though it is indeed fun to speculate), and more interested in how to fix the mess we're in now, so that 10 years in the future, blind people have better choices than OCR.
(Though OCR is still an improvement over the best option in the 1980s I bet. Though I wasn't around so just guessing.)
That's not a culture thing, this is by the mechanism of <img alt. You won't get people who write TeX to change all of their workflows, and even if there was such a culture.
"If there was a standard package that everyone knew they had to include or else it breaks everything from ctrl-F to copy/paste to screen readers, presumably people would use it, right?"
Then it would still be nigh impossible because TeX commands, like all programming languages, compose rather poorly. It would be a herculean effort to produce a kinda but not really TeX that is both accessible with a focus on semantics, yet still compatible with the billions of lines of LaTeX/TeX out there.
"I'm less interested in speculating what would have been if troff had "won", (though it is indeed fun to speculate), and more interested in how to fix the mess we're in now, so that 10 years in the future, blind people have better choices than OCR."
Boycott LaTeX/TeX and PDF everywhere you can. Whenever you publish a paper, also publish it in markdown/html. Publish in OpenAccess Journals like [PeerJ](https://peerj.com/) which convert all of their papers to html in addition to pdf. Consider publishing papers in alternative forms like nextjournal.com .
We need to get our priorities straight in academia :/. This entire "but latex produces such beautiful documents", "I'm working towards getting into the most prestigious journal" culture of snobbery and vanity needs to stop. We need to go back to caring about the content, not the presentation, something TeX ironically was meant to do.
Fine, fair enough. If fixing TeX is hopeless, then so be it. I assumed it just needed a few small tweaks, maybe combined with slightly cleverer PDF viewers. Guess I was wrong.
But then what should people use for math? I suppose there's MathJax, which seems to have put some thought into accessibility.
There's still a problem though. I can't help but notice that the journal you linked to is a biology journal. In some math/CS circles which are TeX's "home turf", TeX is far more entrenched to the point where I'm not sure such things even exist. For instance, arxiv sort of supports HTML, but not really:
https://arxiv.org/help/submit_html
So there I guess step 1 is to make HTML a viable option.
I think MathJax is certainly a step in the right direction, they even support rendering to MathML.
But I agree that there is a certain lack there in terms of full semantic representations. MathJax is more accessible than TeX but it's still describing visual layout, instead of semantic meaning.
Pushing HTML to arxiv is also a step into the right direction.
I think the most important thing we can do is not be complacent with the state of the art. We need to go back to an age of computing where we didn't think we had it all figured out. We need to experiment, and not be afraid to take a step back in some aspects, like layout and kerning, in exchange for other advances like semantic representations and knowledge representation.
I think bred victor has a great talk on this: https://www.youtube.com/watch?v=8pTEmbeENF4
I think we need to experiment with things like observablehq.com or nextjournal.com or the many other that are coming into existence.
re: semantics vs visual layout of math... Wikipedia says OpenMath is a thing, but... that only solves half the problem. Once you have a format that encodes what you want, someone has to actually it.
Like, if some writes x^{-1} and f^{-1}, it's hard for a computer to figure out that the first one means "the number you get when you divide 1 by x", whereas the second one means "the function you get when you compute the inverse of f".
And if the author can't be bothered to slow down and say which is which, then the reader will have to guess.
re: HTML to arxiv: not ready for prime time, if you actually follow that link.
re: kerning: TeX's advantage here is not fundamental, I think. Just need a good font, as far as I know. (Actually that's not far; I know almost nothing here.)
re: layout: CSS is finally getting good at this from what I hear.
re: talk: looks familiar; maybe I should re-watch it.
Here the potential function is, to say the least, not very simple.
There's a similar story for the electronics design software package KiCAD - I played with that a bit in school as well, when it was more like a group of separate software packages duct-taped together with some import/export scripts. The enthusiasts weren't bothered by and seemed to actually enjoy the lack of polish. Recently, though, CERN has injected a bunch of money, programming time, and real-world use engineer's time into the project and it really shows.
I definitely agree that the NSF (or any major LaTeX-using organization) injecting a few million dollars into LyX would be an incredible boon to human productivity. I do feel that it would be better for an organization with internal consumers of the software to undertake this effort than for an outside team to just start with a blank canvas.
The latter too frequently ends up with the primary users being developers, where bug reports or feature requests will be responded to with "Just compile with -D FEATURE_ENABLE" or "For the syntax of that menu item, read the block comment and source code of the foo.cpp/bar() method" - basically, where TeX is right now. Most users of most software will install it with the Windows .exe, a few will use the .dmg/.deb/.rpm or their package manager: a vanishingly small subset are compiling it from source. And if you expect that your users should be able to read C++, your user base is going to very small.
I really doubt Latex is the bottleneck for doing research.
I've spend hours debugging other peoples references, bibliographies and journal dictated styles.
Also if you think that all of academia has MONTHS to work on papers, with our current publish or perish culture, you're wayyyyyy off.
I guess you intended emphasis on the "cost effective" part of the statement, but not everyone will read it like that (I didn't).
Already in the 80's Jeff Bezos refused to use latex due to its absurd difficulty to use. I can't believe it's been 40 years and there still isn't a better tool to write papers with.
It took me less than an hour to create a Google Docs template that looks roughly as good as a standard Latex document. The reason Latex looks beautiful is due to a set of wisely chosen defaults, nothing more.
Yet the focus of Google Docs clearly isn't scientific papers. I don't imagine it would take more than a few passionate employees at Google to expand Docs to a powerhouse of scientific writing.
Pictures of my template can be found below:
(and a blogpost about my Latex induced misery) http://mathiasbonde.com/?p=72
- smart figure/caption/table handling (no orphaned captions, no figures hanging out in the middle of the page) - easy, label-based cross referencing (not selecting each from a popup interface) - text-based equation input
All that being said, I still have an elaborate Word template I use for proposal writing. Because often I have to pass it off to others (who may not know TeX) to contribute, and sometimes I need to quickly bastardize my nice formatting to squeeze into hard page limits.
ShareLatex / Overleaf is very good, but I'd love to see slightly more ease of use as OP suggests. I'd like to see a switchable WYSIWYG / code editor, so laypeople can make simple formatting changes visually without needing to know the right commands, and it propagates correctly into the underlying LaTeX.
But none of the things you described, I think latex does particularly well either. Better than Google docs and the likes, sure, but I think it's possible to do better.
I truly think Google Docs could be so much more than it currently is.
They for example built a brilliant speech-to-text feature that is only available on desktop. Mobile is the one place I could imagine using such a feature. Imagine writing your documents with speech while hiking deep in the woods. Feeling the fresh air with a view of a beautiful brook as Google transcribes your thoughts. That's how I want to write!
Why on earth put so much effort into such a brilliant feature, and then have it only be available to users already sitting in front of a keyboard!?
Apologies for the tangent, it was where my mind wanted me to go!
Besides, the macro language of TeX is not the best programming language IMHO, but it just works, especially for the flexible control of long document.
\includegraphics{img1} \includegraphics{img2}
Or is that not what you meant?
https://commons.wikimedia.org/wiki/File:Latex_example_subfig...
You can reference either the whole figure by using the label inside the figure environment, or you can reference the subfigures using labels inside the subfigure environments.
I'm not saying it doesn't work. I'm saying that if your needs are simple enough, you can get away without it.
Either that or the other answer would be to say that’s why you have grad students .
I'd love Word if it had that one feature, disable direct formatting, to enable changing formatting only through styles.
BTW Latex with embedded Markdown is a rather pleasant combo after you have it set up properly.
For current users of LaTeX, it's dead in the water.
> other answer would be to say that’s why you have grad students .
Of which discipline?
Many journals do allow Word submissions. Nature, for example[0]
[0] "we strongly encourage you to incorporate the manuscript text and figures into a single PDF or Microsoft Word file" https://www.nature.com/nature/for-authors/initial-submission
In my field, Word is preferred for conference submissions, because the tight page limits (2-3 pages) is easier to work with in a WYSIWYG editor.
(If somebody is looking for a project for Mozilla's "fixing the internet" incubator, maybe give this a shot.)
Really wish somebody would make a meta package consisting of numpy, scipy, sympy, sage that removes conflicts and overlaps, and renames all functions with the same naming scheme.
On a related note, personally I think Overleaf or Google Docs is missing the point or the pain points. Even though both of them enabled offline editing but they are really a cloud first applications. The better approach is to have something like TeXmacs (native desktop first application) and make a seamless integration with synchronization and versioning to the cloud. With a proper use of synchronization and versioning technology like rsync and Git, together with seamless connectivity tools and protocols for examples Wireguard VPN, WebDAV or even the latest SMB over QUIC approach, I really think this is very feasible. The CONCEPT is similar to the useful and successful Watcom product or now SQL Anywhere database editing application.
And in fact, there are projects:
- https://www.mathjax.org/, and its much faster (and less feature-complete) variant https://katex.org/
- http://www.luatex.org/ - I guess this is the closest to what you mean
- https://rstudio.github.io/distill/ (using Knitr; it can generate HTML or PDFs)
etc, etc.
Even now, other libraries use BibTeX, which would be the easiest to replace in some more standardized format (JSON or YAML, so it is easy to load from web, directly).
If you're not on linux, spin a Ubuntu VM, and install TeXLive.
(Strictly speaking you don't even need GUI on the machine where pdf gets compiled. I add 'scp output.pdf mydesktop:.' to my Makefile so the file is transferred to my desktop everytime I make. On the desktop the pdf is already open on the big screen and if your pdf reader is good, it'll update the opened pdf everytime the file on harddisk gets updated)
I never use overleaf, and I had no problem finished up my thesis in record time. LaTeX wasn't even an issue.
edit: Overleaf will be helpful for collaborative editing though. I personally would use github/gitlab instead if I ever have to do collaboration.
I still think this is easier than merging the commonly used LaTeX packages into one code base and then building a proper IDE and a method to output to HTML.
Much better to start with markdown, extend KaTeX to support whatever you still need for math notation, and then work on making pandoc output to PDF (or equivalent) natively.
On of the bigger problems is bridging the gap between unpaginated HTML and paginated PDF. However if you start from the LaTeX side you'll need to deal with decades of technological debt as well.
https://en.wikipedia.org/wiki/ConTeXt
(but I still believe there is gold at the bottom of this)
Edit: woooah. I see there is now an installer for context, since april 2019! https://wiki.contextgarden.net/Installation
here we go again....
Instant feedback in one view, typing at the speed of thought should be the goal, and if it means a new paradigm I'm all for it. Imagine it was something you could actually program with; those gross inline math functions would suddenly be so much more parsable.
https://tectonic-typesetting.github.io
It fixes some of the issues that LaTex has.
There are some features in LaTeXML, which powers Engrafo, for adding some "LaTeX++" features, like embedding JS, etc.
Still could do with a lot of work, though.
The capabilities are all there - MathML, CSS for print layout, <figure>, SVG, highlight.js, LaTeX.css, ...
Maybe that could take the form of a single JS and/or CSS file that gets dropped into an HTML page and then the author can create their documents with nothing but special tag attributes/classes. But likely it would need to be even more abstracted than that.
Here's a good attempt that's being made. It's an extended version of Markdown that covers charts, equations, etc.: https://casual-effects.com/markdeep/
I want to write my paper and let LaTeX lay it out for me. Of course I'll have to wrestle with an unruly table or figure now and again but for the most part, the layout and typesetting of the document is handled for me.
I also prefer Latex but I wonder if we're talking about solving a problem nobody has?
You’ll get all the benefits of python, with the added benefits of generating formatted text.
Call it pyTeX.
LOL..
What a lot of people don't realize about LaTeX is that much of its awkwardness comes from it being a product of two very strong personalities (both Turing Award winners for unrelated work!) pulling it in two opposing directions:
• Donald Knuth, who designed TeX as a tool for typesetting primarily, allowing a meticulous author complete control over the appearance of his pages, and attempting to capture in a program (or at least making possible) the highest standards of typography developed over the centuries since the invention of printing.(†)
• Leslie Lamport, who wrote LaTeX as a macro package running in TeX, trying to hide complexity from the author and making things as convenient as possible, out of a (probably correct) belief that authors should not care about the appearance and instead only focus on the content! (https://lamport.azurewebsites.net/pubs/document-production.p...)
The result is that with LaTeX, things look “easy” superficially, but everything is implemented in TeX macros (never intended as a full-fledged programming language), so things break in mysterious ways. Not to mention the additional layers of complexity from zillions of users writing their own "packages" for everything. The error messages don't make sense to users because they often come from the bowels of TeX, and I suspect that many users even just ignore(!) warnings about overfull/underfull boxes.
Conversely, if you look at one of Knuth's own documents written in plain TeX, everything is startlingly simple: for example there is no automatic equation numbering or cross-references; for Equation 5 you just write "5" in the equation and refer to it as "(5)" later, none of that stuff with "\eqref{eq:foo}" or whatever. If macros are used, they tend to be custom for the document, rather than elaborate general-purpose ones. Consider giving plain TeX a try (a great book is A Beginner's Book of TeX by Seroul and Levy); everything makes sense, you feel in power, all the error messages are clear (often the same error messages! but now they apply to something you actually wrote/intended), and if nothing else, you'll end up with an appreciation of what LaTeX is/does.
This comment is turning into a repetition of my same tiresome comments on other sites (https://cstheory.stackexchange.com/a/40282, https://tex.stackexchange.com/a/398372, https://tex.stackexchange.com/a/386592, https://tex.stackexchange.com/a/384881, https://tex.stackexchange.com/a/518802 etc.) so I'll stop here I guess :-)
----
(†): good paragraphs, avoiding widows and orphans and hyphenation on consecutive lines and loose/tight lines adjacent to each other, all of that. By far the longest chapter in The TeXbook, the manual on how to use TeX, is called “Fine Points of Mathematics Typing” and teaches the careful reader about aspects of typography that TeX does not handle automatically — and this is not even counting such advice in other chapters such as ties for avoiding line breaks in “psychologically bad” places. See for example http://www.rtznet.nl/zink/comparison.pdf (linked from http://www.rtznet.nl/zink/latex.php?lang=en) comparing plain-text paragraphs in Word and InDesign and TeX. If you look at the “bad” typesetting that Knuth rejected as painful (https://tex.stackexchange.com/a/367133), I think you get the idea :-)
On the other hand, LaTeX is clearly designed with the vision of being a complete “document preparation system” for authors from the moment they first turn their thoughts into words. The former approach I guess is feasible today if you do most of your “document” stuff in another place (if not pencil and paper, then maybe a plain text editor or Markdown or…) and use TeX only at the end, for typesetting.
But boy do I hate the whole semi-broken macro-based system. The package management feels so outdated. Errors are hard to decipher. Package documentation is there but takes days and days if you actually want to read it.
And then you use Tikz. Same result. Beautiful neat graphs. Horrible obscure macro based system with impossible-to-understand error messages.
EDIT: To complement: I think it's obviously possible, lots of people do it. But I feel I would spend a lot of time on "keeping the structure right" as I edit it. Feels more straightforward with Latex. And then Math is kind of hard to write efficiently, but that's field specific.
I assume it would be similar if I knew as much about word as I do about latex. But there's no way of knowing, as I am not going to learn word without a much better reason than "i wonder if it's actually terrible to write 150 pages when you know how to use it."
If you want it truly beautiful, use a DTP and don't format the word documents at all.
I've edited some ~100 pages regulations (so very few pictures and tables) on Word files too. It's not a nightmare because Open Office works quite well with those.
In general, if you see a .tex and cannot compile it, you can look inside and get a good idea of what the file is about in 5 minutes; a bit more time and you can essentially read it by hand. I don't know of any other semi-popular text format except well-written html+css (which is insufficient for science) that supports this. The cleanest systems will go obsolete and become unsupported; if I want my writings to stay around for 200 years, my best bet is to write them in a format that does not strictly require any decoder. And LaTeX, as commonly practiced, is such a format.
But I hate the syntax, that seemingly simple tasks have arcane syntax or are hard to acomplish and mostly, that it's damn slow to compile. I was really surprised that you consider it fast, it's really interesting to me!
[1]: https://github.com/James-Yu/LaTeX-Workshop/wiki/Install#usin...
Have you tried adding the `-file-line-error` compilation option? The errors are still unreadable, but they become 1000% easier to debug.
TeX appeared when I was in graduate school. Before it did, we were using troff, or whatever it was called, which produced hideous output. TeX was a revelation. That’s why everyone switched to it, despite the hardships endured when trying to get it to do what you wanted. The output looked like a beautiful, hand-set book. There was nothing else that came anywhere close. I wrote my thesis in plain TeX, using a Mac Plus ($1400 with a deep student discount). When LaTeX appeared, the selling point was not formatting, but automatic handling of labels and references. You could add a numbered equation in the middle of your document, without having to re-number, and track down references to, the hundred equations that came after it. Glorious.
Before LaTeX? The previous generation, my professors, wrote things out in hand and gave the result to a secretary. Those with terrible handwriting might type the text, leaving spaces for the math, which they added with a pen. They didn’t waste time time formatting the paper; that was someone else’s job. TeX created a generation of typographical obsessives, including me. Whatever the secretary did wasn’t good enough.
I've tried writting a book using LateX, and it's been nothing but a miserable experience. I'm been coding for 15 years, and tweaking my Linux machine for longer. I know what it means to face rought edges.
But LaTeX is another level of shenanigans.
I seem to understand that researchers use latex because of the seemingless formula integration and automation of references.
Well, why not use asciidoc (https://asciidoctor.org/docs/user-manual/#mathematical-expre...) or myst (https://myst-parser.readthedocs.io/en/latest/using/syntax.ht...) then? They are an order of magnitude easier to use, they can generate websites and pdf, use markdown but allow extensions, they support references, and accept latex formulas.
Is it a case of "everybody uses it so I have to"?
I think a lot of the would-be LaTeX-disruptors lack enough experience actually writing mathematics. LaTeX worked because Donald Knuth is both a great programmer AND a good mathematician. When LaTeX gets replaced, it'll need to replaced by a person (or team) with similar cross-disciplinary strengths.
I personally consider compile times that are measured in seconds for "simple" tasks slow (here's the question of it being an inherent problem of the space).
However when writing stuff for my comp-sci courses, man what a chore. Pseudo-code was a PITA, as was getting decent tables of results where I wanted. So for a lot of those I threw my hands in the air and finished the paper using Word.
It might have been me, I try to learn as I went along, but it certainly was not as easy as math.
Compile time of course, can be fixed by reimplementing the whole compiler in a modern language.
On the other hand:
* For graphs and charts, I'm not sure how it could be done much better in a text-based programming language.
* Even for tables you have to have some sympathy because remember, the cells in the table can (and occasionally do) contain very complicated mathematical contents. Tables are hard, as many a frontend dev will assure you.
The first can be done in a straightforward manner. Can Word do the second easier than Latex?
If anyone has other examples of beautiful typography or graphing/figures throughout history, I'd love to see them! I've been replying to the original thread on twitter with some favorites: https://twitter.com/iraphas13/status/1262489387767480322?s=2...
https://phys.org/news/2018-03-math-bridges-holography-twisto...
https://theportal.wiki/images/1/11/Penrose-Rindler-Clifford-...
He also comes from the pre-ppt age of presenting with handwritten acetate transparencies - and still does afaik. Many of his slides have been captured for the infowebs.
http://cgpg.gravity.psu.edu/online/Html/Seminars/Fall1998/Pe...
http://cgpg.gravity.psu.edu/online/Html/Seminars/Fall1998/Pe...
Some of his original papers from the 1960s were not published at the time, but circulated as samizdat facsimiles of his handwritten notes, until later transcribed by professors or their students, then published in book collections:
http://math.ucr.edu/home/baez/penrose/Penrose-TheoryOfQuanti...
Penrose is famous for his visual imagination, which seems to ground many of his insights. Here is a paper where he invents a visual notation for tensors and operators:
https://en.wikipedia.org/wiki/Penrose_graphical_notation
http://homepages.math.uic.edu/~kauffman/Penrose.pdf
This has inspired recent work by Bob Coecke, as captured in his beautiful book Picturing Quantum Processes.
(arXiv example: https://arxiv.org/pdf/0908.1787.pdf)
http://journal.stuffwithstuff.com/2020/04/05/crafting-crafti...
https://arxiv.org/pdf/1809.05923.pdf
It is set using the Tufte LaTeX template:
But I do think that Markdown/Asciidoc is a viable system for very many tasks.
I was taking this Geography class as a social science elective. We had a paper assignment, and my professor was talking about how it's important to submit HW in this specific format and that they'll give sample .docx file. At the time I didn't have any word processor on my computer so I just quickly asked the professor if it's ok to submit PDF in the same format as I use latex to do my hws. She was like "what is latex"? I'm not a native speaker so I thought I just mispronounced it and described it to her, and she was like no I don't know that you should write your paper in Word. I was stunned. I asked her how do geographers submit their research papers, and she said we just use Word.
> Since version 3, TeX has used an idiosyncratic version numbering system, where updates have been indicated by adding an extra digit at the end of the decimal, so that the version number asymptotically approaches π. This is a reflection of the fact that TeX is now very stable, and only minor updates are anticipated.
There are references with a mouseover, possible to include an interactive chart D3.js, etc. Right now I use it for an internal report in deep learning.
See, for example this paper by Maxwell from 1865: https://royalsocietypublishing.org/doi/abs/10.1098/rstl.1865...
I am surprised that the LaTeX name didn't end up in a heated argument between Don Knuth and Leslie Lamport, like the infamous GNU/Linux vs Linux dispute between Stallman and Torvalds.
Maybe scientists are more civilized, maybe it helped that LaTeX already has TeX embedded in its name.
In retrospect, Lignux might not have been a bad name. The 'g' would be silent of course.
It is a truly beautiful book. Cover, typesetting, illustrations. It just feels compact, clean, and elegant in a way that marvelously reinforces Wirth's own aesthetic. They really don't make them like they used to.
Just do a rough cuneiform sketch on papyrus before you carve into the clay tablet for storage, right?
Also, rendering a beautiful word document is quite chalenge.
Rendered latex document looks stunningly clean and professional out of the box.
They also are plain text, which mean they can be diffed, used with git indexed easily, etc. They play well with anything, being text.
You can even read them without having to render them.
And finally, latex check the structure of your doc. If you change it, it will crash, preventing you from commiting dead references.
\[ ∫_Σ ∇ ⨯ 𝐅 ⋅ dΣ = ∮_{∂Σ}𝐅⋅d𝐫 \]
That will turn into a typeset version when run through LaTeX (using the unicode-math package). More details at https://lwn.net/Articles/657157/.
Also, any format except PDF is problematic. PDF is the only (widely used) format that will preserve formatting.