It would be nice if someone wrote an extension/post-processor applying some OCR to the SVGs, so that plain text could be extracted. I tried to do something like this myself, but failed. What was most annoying was that the SVGs the app writes are actually outlines of the pen strokes, after applying "pen width". As far as I've seen, OCR systems prefer raw strokes as input, ideally with pen pressure information. I don't have enough experience in Machine Learning either to try to build a new model for this use case from scratch.
edit: A sample article draft/experiment I wrote with it:
- as a HTML+SVG: https://akavel.github.io/post/2018-05-31-stylish-elephants/2...
- as a PDF: https://akavel.github.io/post/2018-05-31-stylish-elephants/2...