The future of education is plain text
simplystatistics.org
simplystatistics.org
- Notepad
- Microsoft Word
- Twiki
- Various proprietary WSYWIG that compiles to HTML
- JIRA
- Raw HTML
- Markdown (several flavors)
With nearly every kind of migration, there are numerous pain points. The "raw" formats are a nightmare to edit and update, and the compiled ones require several hours of changing syntax, image locations, etc.
I've been getting so tired of having to re-do stuff on different platforms that more of my docs are starting as Plaintext and then written in pseudocode markup for areas that I know will change on every platform (e.g. generating a table of contents, image tags, etc).
Having just coded an entire website from scratch that was basically just documentation, Markdown comes remarkably close to doing what I want, except when the common format fails to meet my needs, which forces me to then have to switch to a specific flavor of Markdown in order to get something as basic as tables.
The docs of mine that seem most resilient to platform shifts (other than plaintext) are the ones that are written in or compiled to longstanding formats like LaTeX or HTML.
So perhaps my takeaway is, write in something readable that compiles to something widely available. That will provide the least headache.
RestructuredText is very powerful, and has an official specification. Basically what Markdown is missing. Yet, it is plain text and easy to parse.
foo
======================
... the easy and built-in extensibility is a huge plus though, as is Sphinx and the pretty good output writers for many formats. Sphinx btw. works with Markdown, too.Quite possibly my biggest annoyance with rST/Sphinx is the hanging toctree issue, and poor support for multi-format images/figures.
The 100 % offline search in Sphinx is somewhat of a mixed bag and doesn't always give good results, though it's fine for reference works.
And plain text is itself a nightmare to parse and do anything beyond reading it as a block.
Obviously it's not harder to parse than a PDF.
If someone told me they needed me to parse a "plain text" file, to me it sounds like they'd want me to provide some statistics on an unstructured text blob.
Free form text does have properties, such as byte length, character count, word count, line count, sentence count, paragraph count, and so on. Most of which is simply variations on whitespace delimiters.
But an individual plain text file... all on its own? Could be a shell script? Could be a newspaper clipping? Could be a dictionary file? Could be an array of XYZ vertices, edges and faces? Could be all of the above, and an email signature at the end?
Someone tells me: "here have this plain text object," and I presume it is a monolithic blob of ASCII, if I'm lucky, and maybe it's War and Peace or Moby Dick expressed in emoji if I'm not.
Well, most would agree markdown is a plain text format for example, but it's not totally unstructured.
Still it's a bitch to parse compared to something like JSON.
Imagine it wasn't decided on Markdown. E.g.: I don't use markdown when making plain-text notes. I make it all hierarchical and indent using "- ". I.e. headings are top-level, sub-headings are indented by one, and so on. Eventually, you realize that all things can be represented as just lists of nested lists. No need for some arbitrary "heading" type, just nest it all based on the concept.
It's like DocBook with Markdown like syntax
[1]: http://pandoc.org/
I just did essentially the same, in reStructuredText.
The battle is so lost that I don't even bother fighting it, but there are many nice things about reST: you have to explicitly escape html to get it passed through as html (feature/bug), tables are easy, especially the list tables, I'm used to it (feature? bug?), ...
It includes a rich text mode for easier collaboration with non-LaTeX users [2], and you can also write in Markdown if you like :) [3].
Feedback always appreciated if you do give it a try, thanks.
I really like how easy it is to experiment with the LaTeX formatting (something I wasnt super familiar with) and immediately see the output.
I had someone send me their resume template on overleaf and it was super easy to get a similar product with my personal touch.
The only feedback I would have is it was a little awkward to do folders within folders (this was a few months ago) and I had to ken the "hey put a path separator in the name" before I got it.
Good luck with your future projects!
I used to use OneNote for everything, until I moved to Macs... Not only there was no Mac version at first, OneNote STILL doesn't let you import old local notebooks into the Mac version.
With Ulysses/Quiver at least I know that what I write on the Mac will still be readable on Windows/Linux..
Python has a lovely library for it and combined with jinja you can get a hell of a long way in a couple of hundred lines of python as a static site generator.
Format is dirt simple: date alone on a line is the only special element. The main thing I really miss is the ability to scribble drawings inline.
Html or plaintext is all I trust now. A nice editor on top of it and I'm good. If you have to do larger document writing, I recommend Madcap Flare.
Here is an example of using the IPython kernel to evaluate inline Python code within an OrgMode document.[1][2]
More information on how to create multi-language notebooks with OrgMode Babel here[3]
[0] http://minimaxir.com/2017/06/r-notebooks/
[1] http://kitchingroup.cheme.cmu.edu/blog/2017/01/29/ob-ipython...
Though, I think this is a valid concern (and I hate that). It is just similar to how people used to have markup in their documents to see what they were typing, but people are lured by anything that hides this markup. See, nobody actually likes typing \bold{hello} instead of just highlighting and bolding. More, people just want the markup to be hidden.
Orgmode shines not because it hides anything. But because it made a very coherent set of macros to type the markup that I want to use. That and source blocks. (Ok, mainly source blocks.)
This is basically why I switched to Asciidoc. And if you need to collaborate with others you can point to different editors that people can use, there is even AsciidocFX which is made specifically for new users.
Example demo javascript webapp[1]
Orgmode Vim[2]
OrgMode IntelliJ [3]
OrgMode Atom [4]
OrgMode Sublime [5]
OrgMode VSCode [6]
[0] https://github.com/mooz/org-js
[1] http://mooz.github.io/org-js/
[2] https://github.com/jceb/vim-orgmode
[3] https://github.com/skuro/org4idea
[4] https://atom.io/packages/organized
For example, I once attempted to run this org-mode "notebook", ironically titled "Reproducible Research with Emacs Org-mode", and found I had to make significant cosmetic changes to get it to build: https://github.com/eliask/orgmode-iKNOW2012
Maybe things have improved since, but backwards compatibility is important for these kinds of formats.
Plain text: so that no one can own the distribution method.
Plain text: so that no one can own the creation method.
Plain text: so normal people can recover data even when partially corrupted.
Plain text: so you aren't forced to see jarring ads.
Plain text: so that there are no tracking pixels.
Plain text: because connecting information with hyperlinks doesn't require all of HTML or even computers.
Plain text: because it's good enough for metadata.
My future and knowledge is in YAML-fronted markdown and YAML metadata for binaries. Let's take back our data. Look out for Optik.io.
But as evidenced by this and other articles, as everyone is getting fed up with the current state of data, people are coming around to the idea.
You might want to be a bit more clear in your communication. When someone links a URL/domain, you'd expect that there's something there.
I mean, why not put up a splash page on your domain, telling something about your product (what is it? Something with plain text? A Markdown-based static site generator, maybe? Who knows.), or maybe a form to subscribe to a mailing list.
Did you think people would put that domain in their bookmarks and re-visit it themselves?
I want to help organize data using an open format. I think user control of presentation is crucial as AR and VR work further into daily life -- especially considering accessibility.
Splash page is now up, including the product description, and form to subscribe to a mailing list. Please check it out, I'd love to hear what you think!
How do you plan to monetize plain text?
There's nothing to melt down
This site can’t be reachedoptik.io took too long to respond. Search Google for optik io
ERR_CONNECTION_TIMED_OUT
(Just for reference, I do want to build a business on open data, but not by owning your data. I would like to help people manage their data.)
EDIT: Website is live! Try https://www.optik.io until DNS catches up, if optik.io is not responding for you.
Truly the limits are blurry. Even assembly code has an ASCII representation.
a_0+1/(a_1+1/(a_2+1/(⋱+(1/a_n ) )))
Past that into an editor that accepts the spec, and you'll get something like http://imgur.com/a/7hBwv
Disclaimer: I work @Microsoft improving on some math features, and we are the main implementors of the spec. Which makes me sad, since it is an open spec and it is really powerful!
Try to implement them though, you will run into a lot of issues.
Handbook for Spoken Mathematics (1983) http://web.efzg.hr/dok/MAT/vkojic/Larrys_speakeasy.pdf
MathPlayer https://www.dessci.com/en/products/mathplayer/
TalkMaths http://talkmaths.sourceforge.net/
Language and Mathematics: Bridging between Natural Language and Mathematical Language in Solving Problems in Mathematics http://file.scirp.org/pdf/CE20100300008_45591409.pdf
My wife is a math teacher, and the piss-poor experience of writing mathematical notation outside of MS Word keeps them stuck on Word. As much as she and I love LaTeX's math notation, she can't get her departmentmates onboard with that kind of syntax.
They need to maintain notes, guides, tests, quizzes, etc. and basically running a team OneDrive is really the only option because all the alternatives utterly fail a group of non-technical mathies.
Microsoft Office has two internal math formats, one of them is Ecma Math[0], the other the other is the "Unicode Nearly Plain-Text Encoding of Mathematics"[1], it is a Unicode standard and uses only Unicode standard characters. It has the unfortunate property of being hard to type on its own (the integral character isn't on most people's keyboards), but it is pretty easy to read as just plain text. If you copy an equation out of Word you'll get something like this: ∫(x^2/2)
[0] https://blogs.msdn.microsoft.com/murrays/2006/10/06/mathml-a...
[1] http://unicode.org/notes/tn28/UTN28-PlainTextMath-v3.pdf
Disclaimer: I work @Microsoft improving on some math features, and we are the main implementors of the spec. Which makes me sad, since it is an open spec and it is really powerful!
Maybe it could be integrated into the Unicode shaper (HarfBuzz/Uniscribe/AAT) with a language code for "maths".
Unfortunately not! I am wondering if I should make requests to the Windows team internally, the Managed wrapper class they ship has the math features compiled out. :( I'm in Office, where we use an internal UWP safe C++ version of Rich Edit. We bundle it with our AppX and load it up like any other C++ DLL.
Please post over on https://wpdev.uservoice.com/ and get others to upvote! If internal and external asks line up, it becomes much easier to argue in favor of doing a feature.
OK, I went and looked at the standard and it seems like they're thought about this: they allow input as either ∫(x²/2) or ∫(x^2/2) but they apparently output as ∫(x^2/2) because it's more general and editable, like if you wanted to change the exponent to something that didn't have a predefined Unicode superscript.
Thus assume only + - and * / are unambiguous, everything else be explicit about ordering.
EG: ∫((x^2)/2)
Mind blown. I only write a little math, and I can't imagine my first choice not being MathJax/LaTex in a plain text file.
Storytime. A couple years ago, I was in the midst of copyedits on a book with a bunch of math in it. The copyeditors were using a different version of Word.
When they sent the edits back to me, the math was gone. Completely.
Not only that, but when I tried to copy-paste the math in from a prior draft, the Word file refused to save.
Eventually, I had to reconstruct every one of the damn things, by hand.
(Happy consequence: I caught and corrected an error doing so. But still!)
That would never happen in a plain text format. And this was the experience that made me abandon word processors for good.
LaTeX:
\begin{bmatrix}
0 & 1 \\
1 & 0 \\
\end{bmatrix}
AsciiMath: [[0,1], [1,0]]
[0]: http://asciimath.org/ # this is a header
and the more elaborate this is a header
================
which conveys more "headerness" and is more readable than the first.the same could be done with something like ASCIIMath, where a plain-Unicode representation could be an intermediate form between ASCII and MathML. Why? Because keyboards don't type in Unicode, but you don't want to be storing only the final output - storing an intermediate Unicode form seems best, assuming you can keep modifyign it using ASCII and then going ASCII + Unicode => Unicode.
We also published a short post on 'the stoic resilience of the PDF within the digital ecosystem' recently [3], which seems relevant...although it's just a short background piece.
if we use a conservative definition of "computer and programming literacy", then it's not a requisite for either.
markup languages are not programming languages. most people have some familiarity with some markup language.
If you were over 60 in 1985 and are still posting to HN, that would be really cool.
I mean, it's not a joke for its intended purpose; that's fine. It typesets the #$@! out of documents. But TeX is a joke for mathematical notation in particular. Put another way: the best way to understand what an arbitrary mathematical expression in LaTeX really means, is to render it as an image and read that image. TeX can be understood best as a concise way of writing a certain class of vector images, and when you are reading TeX you are reading a computer program which generates an image.
I'm not saying it can't be used for this context, of course it can, I have used it a ton and found it quite enjoyable. The fact that it's an image-based representation makes it very easy to switch from `\int_A dx~\int_B dy~f(x,y)` to `\iint_{A\times B} dx~dy~f(x,y).` However let's not mistake the fact that TeX does not know and does not want to know how you are using the `\int` and `\iint` symbols, is 100% OK with omitting those `~` characters, and has no semantic conception of what `dx` and `dy` are. If TeX were the CAS that it doesn't claim to be, as far as it's concerned that expression could be canonically refactored to `\iint d^2fxy(x,y),` since no one wrapped the `dx` and `dy` in curly braces. TeX claims to be a typesetting and layout system, and it does that well; it's not trying to be a universal mathematical notation.
The contention of the original link posted is, all of these image-based formats like PDF and lecture videos are going away. This may or may not be true, but if it is true then TeX is not going to survive the death of images, precisely because it is a programming language for a class of images.
Maybe something else will survive. The biggest player right now is the Wolfram language, of course, but that can look terribly unwieldy too.
But the point is, mathematics written as a plaintext document needs an interpreter, unless you have been writing papers with math notation for some years. But it is still far more cumbersome (and looks ugly) than writing formulas by hand. And that's not exactly plaintext learning material anymore, then, even though the sources are text files (unlike Word documents).
My ideal workflow (the one I dream about) would be a document camera -like setup that parses math I write on a paper (or a blackboard) into LaTeX (or MathML) style format and ~immediately renders it as beautiful document (with MathJax or similar tool). Like Overleaf, but without typing. (And of course then I could open the source file in text editor to make edits.)
I'd buy mobile app any LaTeX OCR software.
Some kind of plain text math notation using numpy expressions would be cool.
If math notation is unclear to you, it's only because you spent little time learning and using it. Mathematicians care about clarity even more then you, as they actually do indeed spend large parts of their lives reading and writing math. They are very quick to adopt new notation, if it brings meaningful benefits over the old one -- for example, the commutative diagrams and category theory language is now commonplace in all fields of algebra and topology, because it is much easier to draw a diagram and claim it commutes rather than name all the maps involved and write down all the equalities.
I think markup languages like markdown which are both fairly easy to convert into other formats and deliciously human readable are the way to go.
They are OK in theory, for the reasons you state.
The problem is the gratuitous use of PDF which we have all experienced - here's a common (pathological) example:
Document author starts with plain text - no special formatting or fonts, no images, etc. Somehow their toolchain converts that into a PDF file that contains No text, but rather an image of text.
The result is a big, bloated, unnecessary use of PDF that cannot even be parsed or used with anything but a graphical PDF viewer since the text is now gone - there is nothing but a picture. Of text.
Kindles etc. are great when you're mostly reading a flow of text. For anything that benefits from design layout -- positioned graphics, sidebars, footnotes, etc. PDF on a 10" tablet is often better.
Technically, it could be anything. The layout is specified, and everything stems from this.
In addition PDF adapts, by design, a model of paper to the web. It's a "horseless carriage" file format.
Font size, page width, cut/paste and presentation in general should be the reader's choice, not the writer's. The Web manages this, sort of.
The OP is right on in this regard. Even the TeXs of this world, while better than the binary formats, have upgrade complexity.
Sometimes. There are a lot more design options available if the creator of the content maintains control over layout, fonts, etc. Sometimes this doesn't matter--if it's a block of text for example. Fairly simple layouts also render pretty well on the web.
Different content works better or worse with different approaches. One isn't intrinsically superior.
PDFs are awful.
Could have stopped right there.The spec is 1000+ pages, references other docs, has many omissions and contains much that is apocryphal, or at least wildly inaccurate
right now, today, tuesday june 13th, 2017, safari will not open a .pdf to a specific page or bookmark, neither on the desktop nor in ios.
chrome can. firefox can. but safari cannot. which means neither the iphone nor the ipad can do it.
imagine if -- on the web -- you couldn't deep-link to an anchor in the middle of a webpage, but merely to (the top of) the webpage itself.
that's not the only deficiency of the .pdf format. it's not even the most galling one. it's just the one that happens to be hamstringing a certain project of mine at the moment. and it's illustrative.
I believe that MediaWiki, AsciiDoc or LaTeX are particularily well-suited for this purpose.
MediaWiki is already widely known and widely used for knowledge accumulation, namely, in Wikipedia. The downside is, of course, that this wiki language has quite some limitations.
AsciiDoc is a well-designed language with a stable definition for years. (Compare this to Markdown: Are you using the original Markdown? Or a fork of some fork of Markdown?) Also, AsciiDoc can be converted to beautiful HTML as well as beautiful PDF. Also, the clean definition of AsciiDoc means you have no trouble with nesting. For example, in a table cell you can put everything: enumerations, code listings, and so on. You can even put a new table within a table cell if you need to.
LaTeX is the de-facto standard in mathematics and parts of computer science, and has proven to be a stable standard, too. For example, arXiv doesn't accept your generated PDF as black-box, they want your LaTeX source and generate the PDF themselves. (That way, they can, for example, automatically produce PDFs with hyperlinks from documents which originally had no hyperlinks.) The downside is, of course, that LaTeX is not as readable as a plaintext-like format.
So for any "serious" / rigorous documentation purposes, either AsciiDoc, MediaWiki or LaTeX are the way to go.
Markdown stresses ease of readability. I don't find the language low quality at all, just limited. Limited isn't necessarily bad, depending on its purpose.
> isn't GitHub the only reason why Markdown gets pushed so much?
Reddit's use seems to significantly predate Github, and I would say reaches a far greater audience. Unique users visiting in April 2017 is 1.285 billion[1]. Some fraction of that is probably unique people (given anonymous desktop and mobile usage), but given how large the number is, I imagine it's still a very large number.
> If GitHub chose AsciiDoc instead of Markdown as their base, we would be in a better position now.
I agree that Github would be in a better place now, but I don't think that would really have changed anything for anyone else. Even if you want to make a case that Github would have influenced programmers who would have then used it in other projects, I think you need to account for Stack Overflow also. I think it was arguably much more popular and used by a far wider audience than Github for a long time (and may still be, I'm not sure).
Markdown was used widely because it was simple, and users would actually bother to learn and remember the very few options they had. People are more used to it now, and at this point, sure, they might accept something more complex (especially if it built on the rules they already internalized), but I don't think we can blindly assume they would have accepted something more complex.
1: https://www.statista.com/statistics/443332/reddit-monthly-vi...
Edit: s/Unique IPs/Unique users/
Nobody forces you do use all of them. If you just want to use AsciiDoc as "Markdown with more coherent syntax", stick to a small subset that is equally trivial to learn.
And if you need more, AsciiDoc's comprehensive user guide is a huge advantage. In Markdown, you'll have to look around for all kinds of forks. And god forbid if you want to use two additional features that were not designed to work together in the first place. Compare this with AsciiDoc's clear extension system where you can hook up everything and it won't interfere with each other.
I don't see how it's unfair. I think AsciiDoc is much more complex, by just about any metric you want to use. I'm not saying it's necessarily worse by many metrics, just that by the particular metric of getting average people to use it, and not just a random subset of it, Markdown's simplicity is beneficial.
> Nobody forces you do use all of them.
No, but for random internet user there's a real trade-off between what you're trying to accomplish and what it takes to accomplish it, when you're just trying to write a simple comment. Reddit has a link that says "formatting help" below the comment box, and when you click it, it shows every option you have for special formatting with examples in a table with nine rows and two columns, including the header. They could have chosen AsciiDoc, but then they would have to make a decision about what features to elevate to the quick help and which not to, and very possibly which to disallow for their use case.
Markdown is simple for users to use, simple for user to understand, simple for developers to implement, and simple for developers to decide about. That simplicity is both why it was adopted by developers and why users bothered to learn and use it. As I alluded to before, AsciiDoc may have been a better choice in the end, but I don't think it's quite as simple as AsciiDoc does more stuff, so it's better. It's all about the trade-offs, and there's been plenty of discussion on that[1] before, from both sides.
1: Just google "worse is better".
which is why so many people had to "extend" it with different "flavors", which has now created a massive mess of inconsistencies.
sometimes worse is better. and sometimes it's just plain worse. and sometimes it's the worst kind of situation you could ever imagine.
if instead of adopting markdown, people would have extracted a small subset of asciidoc (which predated markdown) or restructured-text (which also predated markdown) to serve the brain-dead use-cases that markdown claimed, those subsets would've been just as "simple" to learn, but also leveraged more cleanly when people sought to extend the light-markup toolkit to longer-form documents.
but the blogosphere thought it was hot shit back then, and took great delight in pushing things viral. ergo markdown. so now we're stuck in a bad situation.
They are all mostly consistent with the core markdown. They are inconsistent in their extensions. Markdown itself does have problems in that there was no formal spec, but that's mostly been resolved with CommonMark[1]. They even go so far as to document the different extensions that have been developed with their different syntaxes[2]. You might be tempted to call CommonMark a replacement, but it's not, it's really just a formalization of a spec based on Markdown.pl the the test suite that resolved some ambiguities.
> if instead of adopting markdown, people would have extracted a small subset of asciidoc (which predated markdown) or restructured-text (which also predated markdown)
In that case, why not Setext, which is from 1991? I'll tell you why, because Markdown was meant to codify already in use norms, and to emphasize readability over all else:
Readability, however, is emphasized above all else. A Markdown-formatted document should be publishable as-is, as plain text, without looking like it’s been marked up with tags or formatting instructions. While Markdown’s syntax has been influenced by several existing text-to-HTML filters — including Setext, atx, Textile, reStructuredText, Grutatext, and EtText — the single biggest source of inspiration for Markdown’s syntax is the format of plain text email.
To this end, Markdown’s syntax is comprised entirely of punctuation characters, which punctuation characters have been carefully chosen so as to look like what they mean. E.g., asterisks around a word actually look like emphasis. Markdown lists look like, well, lists. Even blockquotes look like quoted passages of text, assuming you’ve ever used email. - Markdown Syntax "Daring Fireball – Markdown – Syntax. 2013-06-13.[3]
> those subsets would've been just as "simple" to learn
I think not. For some, including me, markdown was almost zero-cost. It's how I wrote email.
> but the blogosphere thought it was hot shit back then, and took great delight in pushing things viral. ergo markdown.
I think that's highly simplistic, and ignores the realities. One of which is that it was pushed on Reddit, which has become one of the largest and most used sites on the internet. I find it hard to believe the blogospere opining on it (because it's not actually used on all that many blogs) has had more sway than them on this topic.
2: https://github.com/jgm/CommonMark/wiki/Deployed-Extensions
3: http://daringfireball.net/projects/markdown/syntax#philosoph...
it's fairly easy to get the brain-dead part "right". even down to replicating gruber's original bugs and his corner-case complications.
> They are inconsistent in their extensions.
that's precisely my point. and the crux of the problem.
> Markdown was meant to codify already in use norms
markdown's markup did not differ significantly from that of asciidoc or restructured-text. all of them, including setext, leveraged existing conventions from e-mail and usenet.
> and to emphasize readability over all else
since nobody is meant to actually _read_ raw markdown, i've never understood why everyone cites that passage so religiously, other than that is part of the origin story mythology.
> I think that's highly simplistic, and ignores the realities.
due mostly to netnewswire, which installed gruber as its default mac-blogger, gruber's reach was phenomenal when blogging first went viral. if you don't understand the power of that reach at that time, it's probably because you weren't around. and that group of "cool internet kids" still flaunts itself, most notably recently in the nearly-immediate widespread uptake of json-feed.
the _only_ reason markdown was the choice of the masses was because it looked "easier" to a lazy tl;dr mentality. which is a false economy for which the light-markup revolution will have to continue to pay for years down the line.
well, that coupled with the fact that markdown has a catchy name. one cannot deny that. that helped too.
at any rate, kbenson, i'm off to a school reunion, so the last move here will be yours, if you choose to make it. we've hit the point of severely diminished returns anyway.
Because that's not a universal feel, and some people do read it. I write a subset of markdown normally in text. I use asterisks for bold, use a hash for section headings, and use unordered and ordered lists as defined. I value that I write the same thing, and sometimes it's just text and sometimes it gets prettified, and I really don't need to care the majority of the time whether it does or not, because for the most part people understand the conventions used in the plain text.
Here's the kicker, in one job I designed a system to send email to customers that took advantage of this, and if you supplied a text message to email and the markdown version was different, automatically generated a multi-part email with the plain text part being the markdown, and the HTML part being the generated output from the markdown.
> due mostly to netnewswire, which installed gruber as its default mac-blogger, gruber's reach was phenomenal when blogging first went viral. if you don't understand the power of that reach at that time, it's probably because you weren't around. and that group of "cool internet kids" still flaunts itself, most notably recently in the nearly-immediate widespread uptake of json-feed.
I think you vastly overestimate the pull Gruber had over the general people at that time. I didn't know anything about him, but it wasn't because I wasn't around, I was already working in the industry. It was because I didn't have anything to do with Apple products and didn't care. Which is the same for most people. We're talking about three years pre-iphone here. Before the unibody macbook. Apple's core product that was tapping a wider audience was the iPod. If you weren't following Apple as a customer and fan, chances are you didn't know or care who Gruber was. I certainly didn't.
But Gruber wasn't the only author. Arron Schwartz invented it with him, and Aaron Schwartz was helping out an early Reddit a year later. Again, I think you vastly overestimate Gruber's role over actual use in popular sites, such as Reddit, and later Stack Overflow.
> well, that coupled with the fact that markdown has a catchy name. one cannot deny that. that helped too.
I won't deny that at all! I think that probably has more to do with it than Gruber's advocacy as well. :)
> at any rate, kbenson, i'm off to a school reunion
Enjoy! I've got another year before I have my 20th.
> we've hit the point of severely diminished returns anyway.
Agreed. We're really just refining our prior points but not making any headway in persuading each other.
my only note now is that i was never trying to "persuade" you. or anyone else. think whatever you like, wrong or right.
Github is widely used too, and being used in code, it's likely that it will appear in more places than just online wikis.
Remember that one of the major breakthroughs of the World Wide Web was that HTML meant documents were no longer plaintext.
I like text-based formats, but I'm not convinced that "Being non-binary is a huge plus" for parsing. With binary formats you can assume that documents are generated by a tool, which is at least trying to be compliant with a spec, so barfing on noncompliance is more acceptable. With text you have to be prepared to cope with any kind of rat dance imaginable.
SGML has the SHORTREF feature which allows custom Wiki syntaxes such as markdown, but also casual math. It works by applying a context-dependent (parent element dependent) mapping of tokens (such as the `_` token for markdown emphasis) to replacement text (eg. the `<b>` start-element tag). Within the `<b>` context, the `_` token is mapped to the `</b>` end-element tag, in turn, ending the emphasis. In combination with tag omission/inference (such as in HTML) and other markup minimization and processing features, SGML is a quite powerful plain text document authoring format.
[1]: https://en.wikipedia.org/wiki/Standard_Generalized_Markup_La...
There is a naive assumption that all platforms and operating systems will treat your text (everything is either text or binary before it is parsed into something else) equally. This is false. When when this fallacy becomes self-evident many developers will refuse to modify their assumptions in the belief that consuming software will figure it out properly. Sometimes that is true and sometimes will absolutely break your code/prose/data. Clearly that assumption carries a heavy risk, but this is just data at rest.
When it comes to data moving over a wire the risk substantially increases, because all software that processes that text may make custom modifications along the way. You don't see it so much when the protocol is primitive like HTTP, pub/sub, or RSS (but it still does happen frequently). There are many distribution protocols are that less primitive and absolutely will mutilate the formatting of your documents, such as email (which is why there are email attachments).
That's not entirely true. For email, the only characters that have special meaning are carriage return, line feed, period, and the null ASCII character.
Other than that, you can transfer data via SMTP without any issues.
The worst was webmail, which is an email client embedded in a web page. The documents would have to be mutilated so that contents didn't leak outside of a bounded area on the page and visually kill any advertisements or other controls on the page.
If you embed HTML in email and then embed other grammars inside the HTML these applications will brutalize your document at every step. If you are fortunate and extremely defensive your document arrive at a first destination mostly undamaged, but after that any retransmission will thoroughly crush its soul.
My testing was limited to three commercial SMTP servers that I had credentials for. One of them was the SMTP server that I could access using my Hotmail account credentials. Other than changing the Message-Id header that I had manually set in the test message I was sending, I wasn't able to to see any other changes in the message that I had sent (a string of ASCII characters (0-255) excluding the ones I noted in my previous reply).
On the other hand, I have no idea what MAPI does with text.
After re-reading your original post, it appears that you're taking applications and the transfer protocol as a single unit rather than separating them out. If you use protocols like HTTP, IMAP, SMTP, or NNTP over telnet, you'll find that they don't typically mangle text (bytes) that you send outside of certain control characters like I mentioned above.
But you're definitely correct about the problems that applications pose in terms of preserving the text that they process.
Also put some JavaScript in there and see what it does.
I guarantee Hotmail will destroy the original document and do so in such a way that the document evolves from machine that touches the document. The document, from the perspective of SMTP 821/281/5321 (and so forth) is still just plain text.
I'm not disagreeing with you here and you're most likely correct. Exchange has historically not complied with SMTP RFCs. I'll try it out and see how it changes just to satisfy my curiosity.
I suspect that you were to send email like you specified through Postfix, Sendmail, or Exim, it wouldn't be altered before being sent to the next MTA or delivered to the user's mailbox.
We deserve better than this reductionist thinking. Constraints can breed innovation; but they can also just constrain.
That doesn't mean it has to be the distribution / consumption format.
One of the great things about something like Markdown is that it can be rendered to HTML trivially, to display video, equations, etc.
Same thing for ebooks, PDFs, whatever (thanks Pandoc!). It's also easy to translate between formats (e.g. .md, .org, .rst, etc.).
If a new format comes along that everyone wants, there's an extremely good chance that plain text can be rendered to it.
The reverse is not true.
When I first heard about HTML, described in some magazine article, it was touted as a way to give people a chance to have their own unique readers. For instance, a blind person might want to have their own HTML reader, that uses the hierarchy of header tags to help them navigate the document.
Today, my impression is that HTML and its successors are treated more like a general purpose programming language intended to drive visual-specific browsers. This is why we have to create target specific web pages (e.g., mobile and desktop), rather than let each target's browser render generic pages.
I have sent plain text files to some of my colleagues in the past (so that there is a 100% chance that they could read the file), and they were unable to open them because of this issue with choosing the default application, and asked me what app they needed to download to view the files.
I have a kid in middle school, and he has a tablet. These things are often pushed as "educational". Pop quiz: you walk by, and you need to determine, within a couple of seconds, if what he's doing is actually educational. Here's what you see:
lots of graphics, whizzing around the screen
-or-
black alphanumeric characters against a white background
Now, you don't actually know, but generally speaking, the second is a better indicator than the first.
I realized this applies to my own work as well. There are parts of my job that I consider extremely useful to the world, and parts that I really gotta wonder about.
If I'm looking at green or white alphanumeric characters against a black background (easier on the eyes), I'm probably at a UNIX prompt, writing code that is doing something very analytical, or writing SQL. If I'm looking at graphics whizzing by, I'm either trying to figure out how to get a drop down to repopulate with the right thing pre-selected in the latest javascript framework, or, worse, I'm so irritated with javascript frameworks that I've decided to browse the web.
Again, it's not a guarantee, but I'm starting to consider a very general guideline: if you are looking at symbols and alphanumeric characters, the odds that you are building something of lasting value is much higher than if you are looking at things with elaborate UI elements.
It's not 100%. My kid could be watching Citizen Kane and developing an interesting critical point of view. He could be reading 101 fart jokes. It's not a perfect match. There are worthy and unworthy things on both sides.
But as a general rule, for culture and career - if you're looking at plain text, that's a good sign.
https://www.theguardian.com/education/2017/mar/13/teachers-n...
And there are many more such studies available with a simple search.
Check this out:
https://www.amazon.com/Visual-Complex-Functions-Introduction...
I also don't want to come off as knocking video games. I had an Atari in the 80s. Super fun. I would have played that thing 8 hours a day if my parents had let me. They put a cap on it, along with a rule that I also spend some time with this odd object involving black letters printed on a series of white pages if I wanted to play the video game.
The Atari is long gone (though you could say it's as present as ever, in a greatly enhanced form). But those black letters on white pages are still pretty excellent, and are identical to how they were 30+ years ago.
[1] https://en.libre.university/subject/4kitSFzUe/Web%20Programm...
(1) When youcompose text you want to compose first style later. Wysiwyg mixes the two and you end up with crappy spelling and half arsed formatting most thetime.
Which is better?
I would say the second one is more effective.
There are a number of image formats that are commonly supported, as well as video and audio. Once you've decided to go beyond face to face speech in the same language, you have to make compromises. Just try to minimize those, and try to stay within conflicts that have compromises, like line endings.
I think starting from "authorship in plaintext," along with sharing plaintext files, would move things pretty far forward. (I guess there's some irony in there.)
The post correctly notes that R Markdown files are plain text, but the benefits of such (version control) are not discussed by the OP.
You can run two versions of a markdown file or a LaTeX source file through a diff, and see what's been changed. Try that with a PDF or Word file or what have you.
As I like to keep my files in version control, I use plain text formats as much as possible.
At this point I think that plain text files in a distributed version control system that can also import/export patches for emailing (like darcs and git) is clearly superior to cloud-based document storage. The promises of the latter have just not panned out - what we got instead is vendor lock-in and the Damocles sword of account lock-out/deletion/hacking. In many ways (UI responsiveness, control of data and privacy, service discontinuation due to vendor shutting down product/acquisition/bankruptcy, etc), the cloud apps are a big step backward from PC applications.
The average person run around with a smart phone in their pocket, a marvel of engineering and yet they still don't know how to use their computer to do trivial tasks.
It's kind of slow though in my experience. Diffing plain text is amazing. I recently used it to diff a research paper I was peer reviewing that the TeX source was provided for. I felt like a wizard.
(A fairly common occurrence when many people's distributed version control system consists of emailing around "thesis_v0.9.doc", "thesis_v1.0.doc", "thesis_final.doc", "thesis_final_v2.doc", etc. See also: http://phdcomics.com/comics/archive.php?comicid=1531 )
The current content for education is good, but is definitely bandwidth-heavy and is tough to maintain. But dropping to plain text will force us let go a few things that otherwise make learning more effective.
I think HTML (or a WYSIWYG style editors - that seem like plain text, but can be powerful with images, videos, animations when required) also does the same thing like plain text,
- it is always compatible (I give it to you that it takes effort to run hifi stuff on browsers, but still better than plain text).
- is easy to mix and match
- is easy to maintain (thanks to many editors)
- is light weight (not compared with plain text of course)
- is forward compatible (not possible when all browsers decided to not support HTML in its current state).
I appreciate the thought behind bringing this up. I think writing something in plain English, which can then be turned into some super cool learning material that runs everywhere would be awesome. It helps both in solving the issues mentioned in the article at the same time keeps learning effective.
Images, videos, and animations can all be done in markdown -- they're just references to files.
Even something as small as the the MathJax CDN getting retired means that I have old HTML notes which are now 'dead'. If I don't have the markdown/rmarkdown source for them, realistically, they're just not coming back.
That's why I try to use HTML as the universal storage. It's not friendly to edit by hand, but with a basic, limited editor it's wonderful.
With the text being in plain-text, it guarantees someone the base ability to just open up (e.g.) Notepad.exe if all else fails vs. trying to open up a Powerpoint '97 presentation with embedded RealMedia files in 2017.
What PDF offers is a consistent, space-persistent, formatted output. For reading longer works, it actually does matter to me where a passage appears on a page, or within a work. Spatial memory is important that way.
I read a lot of material, in a lot of different formats: paged text (manpages, console-mode browsers), formatted HTML, ePubs, DJVU, image-scanned books.
If I'm reading on a largish (9-10") tablet, PDF in one-page-up format is actually really good. Fills the screen, is almost always suitably readable. Scans of old books (thank you, Internet Archive and Google) in particular are a delight -- there's something about reading a century-plus-old library copy with markings and (hopefully not too much) marginalia, as well as the original typesetting and images.
The main problem I have with fluid formats ultimately is their fluidity. I realise that that's perverse, and that there are times when it's a real benefit, but again, I can't seem to get away from that spatial memory thing.
If I'm extracting content from works, I prefer source (LaTeX, Markdown, DocBook, etc.). Though that's another story.
The ability to spin out formats on demand would be an ideal. I'm looking into ways of doing that.
> If I'm extracting content from works, I prefer source (LaTeX, Markdown, DocBook, etc.). Though that's another story.
except it's not actually another story. it's just a different part of the same story. and a format (like .pdf) which only handles one part of the story well (such as reading) but falls apart on another part (like text reuse) is not -- ultimately -- a good solution.
but that doesn't mean .pdf is worthless. yes, it's worthless as an archival format, and as a distribution format. (and those two are the ones which people commonly pitch as _strengths_ of .pdf, unfortunately, which is misguided.)
but .pdf is fine as a one-off output-format, spun out in an on-demand fashion by an end-user who wants .pdf for their own personal reasons (which require no justification to us). this is what you mention at the end of your comment, and i, too, am working on that...
The idea of requesting, say, <item>.<extension>, where extension is [html,pdf,epub,djvu,txt,json,tex,md,csv,dir,...] would be interesting.
This presumes that there's a way to represent the content as, say, a directory listing, CSV, or JSON archive.
http://www.groklaw.net/article.php?story=20080328090328998
http://www.consortiuminfo.org/standardsblog/article.php?stor...
0: http://www.groklaw.net/articlebasic.php?story=20070123071154...
1: http://www.groklaw.net/articlebasic.php?story=20070123071154...
that severely limits its adoption with people that spend all their days defending the idea of 'attack surface'. so google draw uses svg, and exports svg, but won't import svg. github won't inline svgs in md, etc
i don't personally like the svg design, its got a lot of weird corners and has the usual screwy relationship with the DOM. but having a fully neutral vector format would be such a massive win, by all means svgs if you can get it more widely adopted...maybe some sort of sanitized/sandboxed svg subset.
postscript should have been that format, but they were so focussed on driving rasterization that they made any kind of other interpretation (i.e. editing) impossible.
Here are the associated repos:
https://github.com/JBorrow/latex-pandoc-preprocessor (pandoc does not handle cross-references from LaTeX -> markdown correctly at the moment. Can be installed with pip install ltmd)
https://github.com/JBorrow/lectures-in-the-middle (the website & dodgy compilation script). The docs for this are pretty rubbish at the moment but the configuration file (litm.cfg) is pretty verbose.
If you are interested in setting it up and need more info, it's probably best to chat via email. It should work just fine already, and as I said students love it, but I'm not completely happy with middleman and the huge number of dependencies.
Also: if you're interested in contributing to the new frontend then let me know, I'm looking for contributors.
Or something like DevonThink.
https://stackoverflow.com/questions/87350/what-are-good-grep...
1) vimwiki configured to use .gpg extensions, then use the gnupg vim plugin. Great wiki syntax, easy links, transparent encryption, optional markdown formatting. Downside is no embedded images.
2) Typora with encfs. Great visual interface and the developer has promised some new features (such as a file browser) to make it more useful as a note-taking app.
The main thing I don't like about it is that there's no simple way to ingest data quickly from a mobile device.
I think you should rethink this a little, constraining or adapting your medium will provide you with different results that can be better or worse, depending on your goal of writing.
Oh.
For exchange and processing data (data != information), UTF-8 text with contextual formatting like Markdown, CSV, etc is nice as it is generally tool agnostic.
But for conveying information (information != data), plaintext sucks. That's why Markdown is a copout -- the formatting still matters... and humans don't perceive markup coding as well as the rendered result. Humans do better with visual queues when interpreting written data. Formatting, typesets, bullets, tables, graphs, etc all help use process and contextualize data.
If being able to reference information or data over time, where time > 10 years, you need to think about what you're doing. (Archivists do this professionally.) PDF is the quick path to address this for format-sensitive applications, as big & important institutions like the US Courts use PDF for their documents -- it isn't going anywhere. Big datasets are more complex to deal with... you have to decide whether the raw data should be preserved vs. the processed/analyzed data, etc.
For those who want to point me at html, html cannot be rendered without a browser and created without a messing with tags, but unicode 𝐜𝐚𝐧.
Word is probably one of the most successful software products ever devised. There's a reason -- it solves a problem.
Word is succesful, but you can't highlight syntax in Word, apply styles in MD, do WYSIWYG in HTML (cause fonts).
The process took about a year. My own estimate is that we lost about 2-3 months to tool-related problems, without any benefit whatsoever. I have never in my 20 year career understood the point of word-processing tools, and never will.
And those tools make money. That is wrong.
... wait for it ...
> And those tools make money.
Is this UTF-8, latin-1, 7 bits ASCII…?
I sent someone an important text (7-bit ASCII) file once. Much later, I found that they had tried to read it using NotePad. They had spent an inordinate amount of time trying to format it into something they could read. Had they opened it in WordPad, Word, a web browser, or practically anything else it would not have been a problem, but by default text files are opened with NotePad.
Please don't get me started on tabs vs spaces.
When I think of lecture notes I think of two uses:
1. An informal reference/mnemonic for the lecturer. This use suggests a format that suits the particular lecturer (which may not necessarily be text, or even a digital format).
2. Potential answers to exam questions for situations where there are too few instructors (i.e., lecturer plus TAs) chasing too many students over too short a time for students to practice critical thinking.
Are there other uses? If not, I can imagine standardizing notes across schools could be a detriment by streamlining "plugging-and-chugging".
After all, not all of us are able to learn by just reading a bunch of text. Some of us need graphical (possibly video) or even interactive material as well. Such material is at high risk of being lost (see also: the abundance of interactive science demonstrations online that are implemented as Java or Flash applets).
Before MIME was a thing,I was possible to embed images in emails by including a block of uuencoded text in the middle of the email message (which was otherwise in plain text).
I suppose that could be one option assuming the client program can render them.
Let's agree on markdown/rich text + non-interactive media (images, sound, video)
Everything else works fine with just LF. A lot of new Windows software even seems to ship with LF format configuration files. Especially games, probably because it is so common to have a Windows game with Linux backend servers, so the developers are working with both.
So anyway, I'd go with just LF unless your software is Windows only.
Let the fractured world of code pages rest in peace. Unicode may not be perfect, but it sure beats the alternatives.
Please no. UTF-16 needs to die a painful death.
This person has obviously never been tasked with reading text files from old mainframes.
I was using Windows at the time. I finally discovered UltraEdit.
www.hellolyra.com
Lyra brings it back.
I wonder whether the indent really has any advantage over the blank line separation between paragraphs of text.
(Scroll to the bottom)
I have been meaning to look into what alternatives are out there.