I ported LaTeX to Javascript
manuels.github.com
manuels.github.com
Yes, I used emscripten to port it to Javascript. It was not that hard. Emscripten had three bugs I had to fix (the hardest was to find that the %g format was not supported by emscripten's sscanf).
But it was compiled almost like for x86: first convert the pdftex 'web' souce code to c using web2c, then compile it to LLVM bytecode and the LLVM bytecode to JS.
Top, top work. I love it. Now to try with my monster list of packages...
I want to generate a bunch of bytes programmatically, then have the user click on a button, and allow the user to save a file containing the generated bytes. This should run entirely client-side, with no talking to the server (except for static loading of the HTML page, JS files, images, css, etc). I've been wanting to do this since I wrote my first Java applet way back in the 1990's, and haven't been able to find a way; it's a personal long-standing unscratched itch for me.
Apparently with HTML5 it's possible to do this, since this project does it! I wasn't able to find where in the code the downloading happens, and I'm not sure what this concept is called, which makes searching difficult. Thanks!
EDIT: Browsing commits instead of the source tree was fruitful [1]. What I'm looking for is called Data URI [2]. The inverse operation -- programmatically processing uploads on the client side -- can be achieved with the File API [3]. Now I need a few days to think about what startups will become possible with these capabilities!
[1] https://github.com/manuels/texlive.js/commit/b7b7eef27846473...
[2] http://en.wikipedia.org/wiki/Data_URI_scheme
[3] https://developer.mozilla.org/en-US/docs/Using_files_from_we...
Just as annoyingly, IE10 has createObjectURL too, but only allows the created URLs to be used for <img>, <audio> and <video>. To save a blob, you have to use something ms-prefixed instead, and this doesn't allow inline PDF viewing.
The projects fill somewhat different niches, however. MathJax's ideal use case is adding math support to a CMS like a blog or wiki. This project looks like it's better for offering a Web-based LaTeX-to-PDF compilation service for articles with 100% compatibility with the original implementation.
Also, this project is a great resume-builder if the author is looking for a job that involves wrangling build systems or cross-compiling to JavaScript!
The first package supported is geometry. If you want to add another, append the required files to the supported_packages array.
This is what it looks like for the geometry package: https://github.com/manuels/texlive.js/blob/master/website/ma...
I don't like TeX. I use it and i like the output, but I have never realy understood the language and therefore I don't like it. The syntax for optional arguments([]) seems very odd to me, aswell as the separation between mouth and the rest. The support for named parameters is IMO very hacky.
Wouldn't be a Tcl based macro processor with the tex-Backend nice? Or is this silly?
The "programming language" is horrible. It's a clear example of a turing tarpit. For example you don't have arrays, so you must fake them. You don't have function, so you have to return the value in a glob@l. Monkey patching is considered an art, but this makes many of the different packages slightly incompatible.
The "printing library" is amazing. If you only want to do a standard thing (and someone else had programmed it) the result is nice.
The other advantage is that everyone knows it, so if you make a package correctly it is easy to use by mathematicians that can't program in LaTeX.
(It's also no coincidence that the development of BibTeX coincides with the crack epidemic, but I digress.)
But TeX has a extremely flexible macro system, so it is possible to fake optional arguments. That's what LaTeX do. LaTeX defines a standard way to have optional parameters and make easy to define and use them. It has some strange optional parameters like in \newtheorem, because they are really fake optional parameters. The starred versions of the command also are fake, and have to be defined using a trick. But if you are lucky and don't look under the hood everything works quite well.
I don't know enough about ConTeXt, but I hope that it has a better system for optional arguments.
In some way, this is similar to what happens with some features of high level languages and assembler. For example, the exceptions in C# or Java are translated sooner or later to assembler, but assembler doesn't have exceptions.
This script reqiures a lot of computation in JS and communication between the website and the webworker (~300kb). Maybe Safari is unable to cope with this in a decent amount of time.
The problem is that all of the JS engines involved have various failure modes in which they end up way slower than the others, though... We're talking 10x-1000x slower. And the problem with performance is it only takes your script hitting one serious performance bottleneck to make the speed of the rest of it not really matter. :(
Hence my interest in testcases that point out such performance cliffs, so they can be removed.
Or even just a bug report at bugs.webkit.org
Thanks :)
1. Hitting "Compile" does whatever it should do successfully. 2. Then hitting "Open PDF" opens a new tab which stalls out and crashes
I would almost say its actually a PDF thing that is crashing Safari vs a JS thing.
At least, that's what has happened consistently on my machine with latest Safari and WebKit nightly.
Edit: However, taking the generated PDF from Chrome and dropping it into Safari loads it fine, so maybe it is a JS thing.
Edit 2: Upon further inspection it appears the new tab is being opened with just a data URL representing the entire PDF. It wouldn't surprise me if that's the problem (Safari not being able to handle huge data as a URL). I recall running into a similar crash in MobileSafari because of that a while ago.
Drag and dropping the file itself into Safari, works fine.
I'd also appreciate info on whether Latex packages will be supported in the future. At the moment this does not seem to be the case, though I can imagine this would be a challenging problem to solve for in-browser compilation of documents.
In theory one could search the LaTeX code for '\usepackage' stings and download the required files and mount them into the virtual filesystem of emscripten.
See https://github.com/manuels/texlive.js/blob/master/website/ma...
lstat(/bin) failed ...
/bin
:
No such file or directory
warning: kpathsea: configuration file texmf.cnf not found in these directories:Note that the browser might prevent the website from opening the PDF (Pop-up blocking)
Would a download button appearing when the pdf has compiled be harder to do?
Nice demo
Note that it is somewhat broken for me: http://imgur.com/twIG75m
From what I see, the main work horse is this 250k "binary": https://raw.github.com/manuels/pdftex.js/master/release/pdft...
This webworker script in combination with a .fmt (latex format file) generates the PDF. All in JS. @manuels__ this is awesome!
Yes, the main work is done in the webworker (~250kb) + the tex format file latex.fmt (~700kb)
In printing, text is usually emphasized with an
\emph{italic}
type style.
In my pdf it looks like the every other letter, starting with the first, is missing from the word italic. It looks like " t l c ". Maybe something is wrong with the fonts. I'm using the latest Chrome on OSX.I could not figure out which file is missing, yet. but as soon as I find it, I will include it and fix this.
Firefox 18.0
And yeah, expected error messages like the no-/bin should probably be filtered out. But man, great idea!
I would hardly call this porting something.
EDIT: I just remembered I ported Microsoft Word to Linux by running it inside a Windows VM.
That analogy is not correct. He isn't running existing binaries in an emulator. He cross-compiled code to a new platform, using a new compiler and toolchain, and made sure everything worked in the entirely new environment. And as he mentions in a comment, he found, reported and fixed some bugs in the compiler and toolchain while doing so.
And remember that the web platform is different than a normal desktop environment. You can't just compile code and expect it to run perfectly (it often does for small apps, but larger ones generally no), for example, the main loop has special requirements, threading as well, etc. It's the same as porting a Windows game to the iPhone, for example - there are different APIs, different expectations of how the OS works and what it allows you to do, etc. Such ports take work.
So your dismissal of his work is very unfair.