Show HN: TeXMe – Self-Rendering Markdown and LaTeX Documents
github.com
github.com
What LaTeX truly shines at, is typography. Yet a bunch of people seem to think that it's only a way to write math.
This is an interesting project. It just bums me out because I thought it was a TeX distribution for the web. I'd remove "TeX" from the name to avoid confusion.
p.legibility {
text-rendering: LaTeX;
}For your consideration, this is the manual for KOMA-Script: https://mirror.las.iastate.edu/tex-archive/macros/latex/cont...
That's 552 pages worth of typography that only considers page layouting. It doesn't deal with fonts or kerning at all. It's actually a fun read even if just for the author's occasional frustration with how often laymen confidently ignore centuries-old typography rules that exists for a reason.
For paper documents you can already do Markup -> LatTeX -> pdf for good quality documents.
Some of it, certainly not.
A (larger IMHO) part of typography wisdom has really been distilled from experience making text easier and prettier to read over generations, and should not be thrown away lightheartedly because "screen is different from paper". Yes it is, but human eyes and brain are the same ;-)
- no margin on the binding side
- silly colors everywhere
- mixes normal and reverse text for no reason
- uses german typographical conventions but the text is in english
- extreme abuse of monospace italics
- mixed serif/sans/monospace/boldface on every page
- nonstandard paper size (18x21 cm)
All of this could be expected, and not too worrying, for a text about any subject, except a fucking document about page layout!
The online version highlights everything you can click and has tiny margins to avoid on-screen whitespace.
This is ridiculously condescending... everybody is able to zoom-in with their pdf viewer.
https://www.amazon.de/KOMA-Script-Sammlung-Klassen-Paketen-L...
B) converting layout with css to Tex sounds like an absolute nightmare compared to just implementing proper word, line, paragraph, and page flowing/breaking in the browser. In fact, HTML is pretty opposed to page breaking at all, which is arguably a forte of TeX.
But when it comes to the web I feel like people spend too much time (and waste my resources) to arrive at a perfect layout. There I would prefer a simple robust layout that renders well on all devices and resolutions.
src/renderer/texpara.lisp in https://github.com/dym/closure
I write a lot of LaTeX documents and use the TeX Live distribution to typeset them, so I understand why the name TeXMe could cause some confusion.
This project only cares about protecting the LaTeX content (the MathJax supported LaTeX, not the "real" LaTeX) by hiding it from the Markdown processor, so that the Markdown processor cannot mangle the LaTeX code before it is fed to MathJax.
I could have named it MathMe, JaxMe, or something similar that would have eliminated this confusion between MathJax supported TeX/LaTeX[1] and the "real" TeX/LaTeX. However, unfortunately I did not spend sufficient time thinking about a good name for this project, thus the name TeXMe.
It’s based on LaTeXML, which does the heavy lifting of converting LaTeX to HTML. I’m more interested in the CSS applied on top of that. I want to make web output of the same calibre as the PDF output. (Or, near to it at least. :)
This to me is the unsolved problem. Interpreting a LaTeX file on the web is probably not a crazy problem -- heck, worst case scenario recompile the entire LaTeX compiler into WASM or something.
But, typesetting on the web is really hard and browser capabilities that I know of are just not up to snuff, to the point that conventional wisdom is you pretty much never justify text on the web, ever, because it won't ever look good. So generic solutions to that problem are something I'd be interested in.
The secondary problem past that is to make justified text look good while still being responsive :) But even just for fixed column widths it's something that I'd love to see.
Props, it looks like a very cool project.
I respectfully disagree. I think people use LaTeX to write scientific documents because there is no better alternatives, not because LaTeX is good.
I find the typesetting system in LaTeX is extremely verbose and over-complicated, as compared to something like HTML+CSS or Markdown.
If someone comes up with a good enough typesetting system with some inspiration from Markdown, HTML and CSS, LaTeX will lose its popularity. HTML and CSS are more expressive than LaTeX, so it should be theoretically possible.
One possible explanation for the lack of better alternatives is that the people who really need the alternatives and people who are capable of coming up with better alternatives are somewhat mutually exclusive. Former is more research-heavily and latter is more coding and user experience heavy.
pandoc -s markdown.txt -o latex-output.texBut I'm obviously just being a pedantic jerk when I say that. I get what you're getting at and you're not wrong. The real question to ask is, "is a programming language a good place to put a typesetter?"
Making things more powerful doesn't always make them better, if that extra power forces you to introduce extra complexity.
[0]: https://my-codeworks.com/blog/2015/css3-proven-to-be-turing-...
Markdown is ubiquitous because it's easy to learn and does one thing well. It's a very Unix-y idea. Markdown is a document markup language, not a Model-View-Controller framework. It's not trying to be the front end for your SPA, it's trying to allow you to type Reddit comments and blog posts.
That comes with advantages, the big one being that here's the entire documentation you need to look at if you want to start using Markdown[0] and here's one section of LaTeX's documentation[1]. I could have an entire office (programmers and non-programmers) using Markdown in about a week to a month, I couldn't do the same with LaTeX.
Of course LaTeX is more powerful that Markdown. Markdown isn't really a tool for typesetting tbh. But the point is that because extra functionality nearly always comes at the cost of extra complexity, you always need to take a step back and ask how important that extra functionality is. If someone says, "Okay, but I really need to meta-program my blog post", even at that stage I'm not sure that I push them towards LaTeX or its equivalents, because maybe it's better at that point to just jump to a full featured language like Python or Javascript.
That separation of concerns also means that I can have different people handling different things. If my office is all writing Markdown, maybe I only have one or two people who are handling the CSS for how it renders. Nobody else in the office needs to know about the CSS side of things.
[0]: https://daringfireball.net/projects/markdown/syntax
[1]: https://www.latex-project.org/help/documentation/fntguide.pd...
Markdown it's essentially ASCII art for simple documents. Nothing wrong with that, but it is not seperation of concerns. Rather, it is simplicity of some concerns.
And to make it hugely amusing, the rules in markdown for just content are not much different than TeX.
I am somewhat happy with AsciiDoctor[0], although I honestly feel like AsciiDoctor goes farther in the other direction than it needs to and ends up overcomplicating itself. I spend more time reading the AsciiDoctor documentation than I want to.
Would you be willing to expand more on what you mean about a separation of concerns vs a simplicity of concerns? I was using the term in the sense that Markdown not only doesn't embed logic, it also doesn't embed style. It's literally just content, and you use other technologies to get at the other parts.
In contrast, LaTeX embeds logic, style, and content. When I work with Markdown I don't stop thinking about style -- I just use other tools for that. That separation allows me to then give other people content access without asking them to worry about CSS.
I guess I could see that it's a flat-out dropping of meta-programming from my document generation (although Markdown is a very easy compilation target for templates). But that just kind of gets back to my original point -- do you really need to meta-program blog posts and books, given that it greatly increases the complexity of LaTeX? Would it be better to have a simpler version of LaTeX that got rid of that functionality and said, "I'm just gonna do text layout really well, and nothing else"?
I meant "simplification of concerns." Basically, if you simplify what concerns the content author can have, it is easier to encode them. In the case of markdown, you are very limited in what you can define.
And I am not at all against the ideas. I'm very partial to org-mode. For example, I wrote http://taeric.github.io/Sudoku.html using an org-mode buffer. And I've been very happy with the markup it supports. However, there is a lot of markup around some things it can do. And, I've found that the constant churn as folks improve it has been tiring. I think it has stabilized recently, which is good.
Could someone make something that was more aesthetically pleasing for the typist? Likely. Though, most of the added structures of LaTeX are often not needed for casual documents. Which is why it was oddly refreshing to read some of the core source of Knuth's books.
I know, the proper way is supposed to be to use Jekyll or Hugo or something like that, to compile to raw html on the server side but this is much simpler for the publisher. No configuration and compilation and updating on the server involved.
I do not think that the user will notice or begrudge the few milliseconds of processing in his browser. Especially if this could be converted to a WASM module. Any takers?
Impressive. A sweet spot between HTML, Markdown and LaTeX.
The one thing I would wish for is a build tool where I could insert my CSS (or a choose from available themes) and build my version of TeXMe (I have a CSS file already which I use to convert Markdown files to PDF).
I mean, yes it is not that complicated, and I might just do it manually, but for wider adoption, it would be cool to have themed versions which do not clutter the Markdown files.
I have a script which converts Markdown to HTML (using Discount; referencing my CSS file and highlight.js for syntax highlighting in code sections) and starts a headless chromium which in turn converts the rendered HTML to a PDF file.
Care to share your script?
$ sha256sum markdown2pdf.zip
97b9df1091cb29e15b699fe1b87d0de20929abfe5357d9e4db920fb24a76406b markdown2pdf.zip
It consists of multiple files (entry point is markdown2pdf.sh) and also supports a flag '-w' to watch the markdown file so you can use an editor to modify it while your PDF viewer (in my case okular) keeps updating the changes.I built it solely for my purpose, so it might not fit your needs (or taste). ;-)
Legal: I have no special requirements, so think of it as MIT license, but I included some other files (everything in the directory 'Markdown-Styling' except 'css/app.css' which is my CSS theme) which are available free of charge somewhere on the internet but might have their own licenses and terms of use.
I use Markdeep for most Markdown content, including Github, as I like being able to compose in my own editor, joe, but be able to quickly see the results before committing without having to explicitly compile anything--just tab into browser and refresh.
The "problem" with Markdeep is that it supports so many different features (including LaTeX math typesetting, but also ASCII diagram rendering) that it can be difficult stick to a common subset supported by, e.g., Github.
But in TeXMe, the Markdown + LaTeX code goes in a <textarea> element. This difference, in my opinion, leads to more robust parsing of the input. TeXMe handles some cases fine where Markdeep breaks. For example, this Markdown code breaks in Markdeep:
Here is a fenced code block that breaks in Markdeep:
```
print("unusual <string")
```
<!-- Markdeep: --><style class="fallback">body{visibility:hidden;white-space:pre;font-family:monospace}</style>
<script src="markdeep.min.js"></script>
<script src="https://casual-effects.com/markdeep/latest/markdeep.min.js"></script>
<script>window.alreadyProcessedMarkdeep||(document.body.style.visibility="visible")</script>
The `<string` in the above code is interpreted as the opening of an HTML start tag by the browser. So what looks like a fragment of Python code to a human ends up being parsed as an HTML tag by the browser that looks like: <string") ```="" <script="" src="markdeep.min.js">
Once the browser has mangled the input like this, there is no way for Markdeep to retrieve the original input entered by the user.Markdeep's answer to this problem is to use HTML entity `>` instead of `<`. But that disrupts the natural flow of taking notes in a text file where we might want to paste verbatim code.
TeXMe's answer to this problem is to put the input inside a <textarea> element. So the same example works fine in TeXMe and MdMe:
<script src="https://cdn.jsdelivr.net/npm/texme"></script><textarea>
Here is a fenced code block that breaks in Markdeep:
```
print("unusual <string")
```
The user input is within <textarea> element, so the entire input can be retrieved verbatim and rendered correctly. This is precisely the reason why I wrote TeXMe. I was writing a lot of documents related to neural networks and I needed to paste code examples without having to bother about substituting special characters with their HTML entities.The other minor difference I see is that Markdeep does not conform to CommonMark. It has several restrictions. For example, Markdeep does not support indented code blocks, two trailling spaces for hard linebreak, intra-word emphasis with asterisk, setext headings with one/two minus/equals characters as underline, and single line blockquote. TeXMe is not opinionated about CommonMark. It just hides math content, invokes commonmark.js to render CommonMark, then unhides math, and invokes MathJax to render math. MdMe is even simpler--it just invokes the commonmark.js processor on load. For example, the following code, although valid Markdown (CommonMark), does not do what we normally expect:
H1
==
Intra*word* emphasis.
> What I cannot create, I do not understand. -- Richard P. Feynman
print("hello")
Line
break
<!-- Markdeep: --><style class="fallback">body{visibility:hidden;white-space:pre;font-family:monospace}</style>
<script src="markdeep.min.js"></script>
<script src="https://casual-effects.com/markdeep/latest/markdeep.min.js"></script>
<script>window.alreadyProcessedMarkdeep||(document.body.style.visibility="visible")</script>
It works fine with TeXMe or MdMe: <script src="https://cdn.jsdelivr.net/npm/texme"></script><textarea>
H1
==
Intra*word* emphasis.
> What I cannot create, I do not understand. -- Richard P. Feynman
print("hello")
Line
break
TeXMe or MdMe is a good choice if you care about conformance to CommonMark and pasting code verbatim in your files but do not need the additional features that Markdeep provides. Markdeep has a very impressive feature set that includes task lists, definition lists, tables, diagrams, syntax highlighting, etc. This makes writing many different types of documents possible in Markdeep.TeXMe and MdMe, on the other hand, take a minimalist approach and focus only on getting the CommonMark and LaTeX (MathJax) rendering right by keeping CommonMark.js and MathJax out of each other's way.
By the way, to use MdMe instead of TeXMe, just replace "texme" in the CDN URL with "mdme".
TeXMe and MdMe (stripped down fork of TeXMe) now support Markdeep style of putting the content directly in the <body> element (i.e., without a <textarea> element) and then putting the <script> tag to load TeXMe/MdMe at the end of content.
But this method has the same caveats that Markdeep has, i.e., an HTML syntax error in the content can lead to broken rendering. Therefore, I recommend using TeXMe's original method of putting the content in textarea for more robust parsing and rendering.
For further details, please see:
* https://github.com/susam/texme#content-in-textarea (TexMe's original method of writing content in textarea that leads to robust parsing and rendering).
* https://github.com/susam/texme#content-in-body (The new method of writing content in body that makes the content look neat but has some caveats. See the "Caveats" subsection in this section for details.)
It also demonstrates documents created this way require code from two different third-party origins to work at all.
Well done!
Also, it doesn't technically need to be a preamble. Markdeep works by appending similar code to the end, which is much more innocuous when viewing the file outside a browser or when its rendered by other software (e.g. on Github).
For the kind of user input (e.g., code pasted verbatim) I was trying to handle, a preamble was technically necessary, at least a `<textarea>` start tag was necessary at the beginning of the document.
For more details about this, please see the comparison of TeXMe vs. Markdeep I have posted here: https://news.ycombinator.com/item?id=18314175
Thank you for highlighting this feature of Markdeep. I have now added this feature in TeXMe and MdMe too. We can now put the `<script>` tag to load TeXMe or MdMe at the end of content. Here is an example:
# Euler's Identity
In mathematics, **Euler's identity** is the equality
$$ e^{i \pi} + 1 = 0. $$
<script src="https://cdn.jsdelivr.net/npm/texme"></script>
However, like I have explained in another comment[1], this method of writing content requires the content to not have any HTML syntax errors, otherwise the output would be mangled. Markdeep has this limitation too. There is a 'Caveats' section[2] in the README now that discusses this in detail.The original method of writing content with a single line of HTML code at the beginning of the content does not have this limitation. But if the HTML code in the beginning feels distracting, we can now put the `<script>` tag in the end too.
That'd be a neat way to quickly create documentation websites for open source projects!
> Render Markdown Without MathJax > To render Markdown-only content without any mathematical content at all ...
<!DOCTYPE html>
<script>window.texme = { useMathJax: false, protectMath: false }</script>
<script src="https://cdn.jsdelivr.net/npm/texme"></script><textarea>
### Atomic Theory
Atomic theory is a scientific theory of the nature of matter, which
states that matter is composed of discrete units called atoms. It began
as a philosophical concept in ancient Greece and entered the scientific
mainstream in the early 19th century when discoveries in the field of
chemistry showed that matter did indeed behave as if it were made up of
atoms.
Right now, it requires a few more lines of HTML code to disable MathJax related processing. I am planning to add a simpler mechanism later, so that something like merely appending "?markme" to the TeXMe URL in the <script> tag is enough to set it to Markdown-only mode. If anyone has a better idea regarding this, please let me know or send a pull request.This is a stripped down fork of TeXMe that does not contain any code for MathJax related processing.
Here is a minimal example:
<!DOCTYPE html><script src="https://cdn.jsdelivr.net/npm/mdme@0.1.0"></script><textarea>
# Atomic Theory
**Atomic theory** is a scientific theory of the nature of matter, which
states that matter is composed of discrete units called *atoms*.
This is absolutely fresh--published only a few minutes ago. If you find any issues, please report it.I hope you like this.
If you can provide a reproducible example of the issue, it would help me to understand the issue and fix it. I suggest filing a new issue with the details.
[1]: https://opendocs.github.io/mdme/examples/e04-style-custom.ht...
I am kind of surprised that the styling does work in your minimalist html (style=none) example because there the <body> is missing too and yet the style is applied to it.
You might wish to take this pitfall into account in your exposition, probably by putting <head> and <body> sections into your "valid html5" example nonetheless.
Though I acknowledge that, strictly speaking, you had done nothing wrong in your individual examples.
The <body> element's start and end tags can indeed be omitted under certain conditions (which are being met in the "Valid HTML5" example). Here is the actual specification quoted from https://www.w3.org/TR/html52/syntax.html#optional-tags:
"A body element's start tag may be omitted if the element is empty, or if the first thing inside the body element is not a space character or a comment, except if the first thing inside the body element is a meta, link, script, style, or template element."
Then in "Example 6", it provides the following as an example of a valid HTML5 document:
<!DOCTYPE HTML>
<title>Hello</title>
<p>Welcome to this example.</p>
Now coming back to your issue, perhaps the selectors in your style do not match the actual elements that TeXMe is setting in the page? I am just guessing here because I need the actual code to be able to nail down the cause of the issue.If you can share the complete code you have either here or at https://github.com/susam/texme/issues/new, it would really help in determining the cause of the issue.
One thing that's not immediate clear from the examples: how (if at all) are you avoiding Flash-Of-Unstyled-Content on initial load? Might it be better to put the markdown in a <noscript> block rather than (or as well as) a <textarea> when editing isn't required?
If you write the document with HTML markup you then lose this flexibility.
How so? HTML can be always embedded inside markdown
If you are going to write HTML then I'd use a subset of that and convert to everything else (including markdown if desired).
The first and third tools require you to sign up. How exactly do you create a distributable Markdown file there that can be read both in editor as text and viewed as HTML in browser? The second link does not recognize Markdown.
Isn't it better to use something like my https://html-notepad.com where you can just create/edit your document in WYSIWYG and use Markdown as one of input options? (Markdown is coming there).
I think that this is the best of two worlds.
In those cases, having an online editor and an rendering engine that supports something like markdown, is beneficial.
Thank you for reporting this. I have fixed it now.