2. Saves context for llms and they are often trained on markdown so work best with it.
3. For search. Can search markdown much better than html with postgres
Wrote https://markdown.download to help me with these
The templates are very basic, eg. <p> <b> <a>, etc. Users can also customise these via a WYSIWYG editor.
We then use turndown.js to convert the rendered HTML email to markdown which we then use for the text version of the email.
My use case involves scraping job boards so that I don't have to doomscroll them myself anymore, and storing them in Markdown makes them smaller while also removing a bunch of extraneous classes and structure.
Further, the side project I'm working on for managing all of this can then render them in a way that makes sense.
Not exactly a common usecase I wouldn't think but it's good to be able to do this.
I’ve sometimes been converting it back to md to include the text for each exercise alongside my solutions.
In my case I used a custom HTML to Markdown converter that was specifically built to support only what I needed in order to convert those Advent of Code exercises to markdown.
Mine was also written in Rust.
Markdown in the database is also easier to look at, reason with and takes up a lot less space. Especially if the content was pasted from say MS-Word to a Content-Editable field... omg the level of chaos there.
I swim around a lot in the "XML High Priesthood" pool, and the latest new thing is this: AI (sucking down unstructured documents) isn't capable of efficient functioning without Knowledge Graph, and donchaknow a complex XML schema and a knowledge graph are practically the same thing.
So they're glueing on some new functionality to try and get writer teams to take the plunge and - same old same old - buy multimillion dollar tools to make PDFs with. One sign of a terminal bagholder is seeing the same tech come up every few years with the latest fashionable thing stapled on its face. They went through a "blockchain" phase too, where all the individual document elements would be addressable "through the chain".
Anyway . . .
Anyway, thing is, there's a teensy shred of truth in what they're saying, but everything else about what they're suggesting would, I think, either not work at all, or make retrieval even less dependable. Also, to do what they're trying to do, you don't actually need a gigantic full on XML schema. Using Asciidoc roles consistently would get you the same benefit, and would save a hell of a lot of space in a very limited window.
That's why tools like this exist: https://jina.ai/reader/
Demo: https://r.jina.ai/https://news.ycombinator.com/item?id=40695...