Sphinxtr: Creating a Portable PhD Thesis
jterrace.github.com
jterrace.github.com
I put my thesis together in HTML using Pandoc, customized to look a bit like a Tufte publication with margin notes. It had animations, anchor links when cross-referencing paragraphs, and in my eyes, made sense for a document like a thesis.
My committee members, on the other hand, were in consensus that a thesis should be a PDF with page numbers, not some multimedia document with hyperlinks. When I made a PDF[0] and submitted it to the university, someone even checked it to make sure the page numbering switched from Roman to Arabic at the correct point as stipulated by the submission guidelines (it didn't, hence my knowledge of their thoroughness).
Considering that universities today often don't even want a hard copy of a thesis, this is a world that is tied very strongly to the paper paradigm.
[0] Fortunately Pandoc did this quite nicely, though I did have to make a few changes to deal with citations. Incidentally, there is no good, cross-platform way to put an animation in a PDF.
I think that more generally when a document becomes sufficiently complex it is better to use a page-oriented, typesetting approach. You'll want to fiddle with where certain page breaks and figures end up, and if you do that you might as well stick with one canonical format.
You can view it as the paper paradigm. Alternatively, you can also see it as sticking to a uniform, canonical format. Getting fluid layouts right is a very difficult problem, and manually getting the layout right for a fixed format is simply a much better solution for this purpose.
Which is more than a little quaint, in 2012.
Right, and I have even seen page numbers in citations of theses. It wasn't common in my field at least, but it happens. For a journal citation, generally page numbers tell you where in the volume that article is, not the content that is being referred to. I think there is flexibility in citations—one can refer to a figure, table, section or even paragraph number with at least equal clarity as a page number. There is convention here, I agree, and my point in my original post is that the convention is followed dogmatically rather than pragmatically in academia.
Getting fluid layouts right is a very difficult problem, and manually getting the layout right for a fixed format is simply a much better solution for this purpose.
I have given this topic a lot of thought. Partly from working on an HTML document for a thesis, partly because I've worked in the area of cognitive ergonomics. At the moment our tools aren't great for this, but that's really our fault as software designers. Fluid layouts should be better than fixed ones.
Concerns about fluid layouts are new, a byproduct of current technology. Printing technology didn't afford us these problems. You can remember when it was commonplace to see websites that contained text as bitmaps, to force readers to consume content as laid out by the designer. There were sites that rigidly spec'd font sizes and browsers that allowed that. Today, the best web designers accept the fact that content will be consumed on different screens (sizes and pixel densities) by users with different preferences. They design layouts that can gracefully handle a large font size stipulated by the reader, rather than design the way that a designer would for print.
So what is the underlying issue with fluid layouts?
Well, we have text that relates to graphics, and as Tufte points out in some great books on the topic, it's particularly useful (and historically very common) to have the relevant text presented in conjunction with graphics. That's an obvious point, but with fluid layouts this can unpleasantly 'break.'[0] This is a solvable issue (even using current web technologies), in a way that can evolve beyond what paper can do for us.
I'd like to see a markup that allows me to designate which graphics are relevant to a block of text, and have those figures shown together. For example, if the display allows it, once you start scrolling past a relevant figure it could slide to the margin of the text and stay pinned alongside as you read relevant text for reference. Perhaps the reader could select other figures to pin, or unpin figures as well.
Think of the number of times you've flipped between pages of a PDF to read text and look at the figure described. Sometimes I end up opening the same document twice and putting them on my screen side-by-side. That's the kind of hack us human factors types love putting in slide shows about how we should be designing for users.
With high-resolution displays getting to mass-market price points, we are only missing the right tools to take technical documents to the next level. As I said, this could be done in HTML, CSS and javascript today, but it's not author-friendly. Even if the tools did exist to make this easy, it would frowned upon in the academic world for being different. But some fields are more progressive than others, and eventually we'll move past the page as a paradigm. I've got to applaud the OP for taking steps in that direction.
My futurist speculation: the move away from pages in the scientific community will happen at the same time that it will move away from traditional journals as an idea distribution channel. I don't think that's in the immediate future, though.
[0] In fact, this often breaks with fixed layouts too, except it's only really bad during document creation. Think figure placements in LaTeX. Fortunately the document creator has to do this work once, and then all of the readers see the exact same thing. Better tools for fluid documents would make things friendlier for the author as well as the readers.
While the idea of multi-format thesis, or at least double format - PDF and HTML, is very compelling, I doubt there could be good enough solutions for that for any thesis which contains more then just text and some inconsiderate amount of figures/tables/formulae at all.
[0] http://www.w3schools.com/cssref/pr_print_pagebi.asp [1] http://www.princexml.com/
"Just write and pay attention to content, not formatting," has led to staring at the clock, wondering how 5am came around so quickly more nights than I would care to admit.
God forbid you use Sweave.
I'm looking back 22 years and getting the twitch again. All that time I saved not playing with fonts, kerning, margins, and line height has been burned aligning figures or tables or yes, trying to get that frickin' float on a page at least vaguely proximal to its reference.
There are sphincters in your eyes.
There is a sphincter in your esophagus.
Your body is full of sphincters.
The anal sphincter is just one of many sphincters.
And now you really do know.
As an example, if people lived on Venus, they would be called "Venerials" by the proper genitive form. Alas, doctors got to it first (Venus, Roman goddess of love...) so it was changed to "Venutians". (Source: Neil DeGrasse Tyson's podcast, StarTalk radio).
Still true though.
I'm going to keep telling myself that your use of 'genitive' right next to 'Venerials' was entirely conscious and deliberate.
I'm sorry, but this cracked me up. It brought out my inner Beavis and Butthead. I apologize. But, I mean, come on. Most people do associate the word sphincter with the butthole, and surely the creator should have known that. (Unless the, umm, cheeky association is intentional?)
GitHub repository: https://github.com/jterrace/sphinxtr
Disclaimer: I'm immature...
[...]
2. (Literature) A citation from some author, or a sentence
framed for the purpose, placed at the beginning of a work
or of its separate divisions; a motto. EpigraphicI use vim, change a line, and hit ":w". My git-onNotify [0] script detects a change and issues "make show". The Makefile uses rubber or latexmk to build a pdf, then issues "gnome-open $PDF", which opens the new version in my pdf viewer. If my screen is tiled, the preview on the side just updates.
Essentially, I just save my tex file and wait for the change.
[0] https://github.com/beza1e1/dot/blob/master/bin/git-onNotify
I also have a script which periodically converts the source files from LyX to both HTML and PDF, then dumps them in the webroot of the Apache server running on my Uni computer. This folder has a .htaccess file which restricts access to my supervisors and myself using the Uni's LDAP server.
It works a treat for me.
[0] - http://www.lyx.org/
I usually use perl/sed/awk + markdown for generating html from my own made-up mini-formats. I'd love to keep the sources in org-mode instead, but I wonder whether elisp is the right language for the type of text munging I want to do.
Modifying the HTML output (adding classes etc.) was fairly OK, I haven't tried anything crazy though.
Another concern is: is it practical to include figures drawn with Tikz? I find it the easiest way to lay out many things, but it effectively means LaTeX lock-in.
[0] http://worrydream.com/LearnableProgramming/ [1] http://worrydream.com/ScientificCommunicationAsSequentialArt...
Every time someone invents a new markup format for absolutely no reason, I die a little bit inside.
Incidentially, semantics, not presention, was originally the point behind HTML, but it got warped and twisted over time into a kinda-sort presentation-oriented language.
restructuredtext, on the other hand, was originally developed to create documentation for Python programs. It might be good for that purpose (wouldn't know; haven't used it.) But it's certainly not good for typesetting mathematics, research papers, academic quotations and so forth. Hence the large amount of wheel reinvention going on here. It's a little bit like writing your research paper using JavaDoc comments. Sphincter indeed.
On the other hand, TeX was developed by Donald Knuth, a guy who spent his entire life doing research and writing papers about it. It has excellent math support, and is a true semantic language. I've written a few papers in TeX and been very happy with it.
Anyway, if RestructuredText were good at typesetting research papers, there would be no need for this project, would there?