I'm very excited about it.
Let's build a pipeline: all the pdfs -> markdown them all -> archive.org them all
I'm very excited about it.
Let's build a pipeline: all the pdfs -> markdown them all -> archive.org them all
I don't think that is the right approach for archiving. The preferred pipeline would be
all the pdfs -> archive them all -> markdown them
This way you can always re-run the conversion as bugs are fixed and improvements are made. Generally archivist prefer to save as close to the source material as possible, because every transformation from there can only lose data.
I opened the first example to a random chapter (1.4 Formal and natural languages); within the first three paragraphs it:
- Hallucinated spurious paragraph breaks
- Ignored all the boldfacing
- Hallucinated a blockquote into a new section
This is not a tool to produce something for humans to read.
Maybe it might be useful as part of some pipeline that needs to feed markdown into some other machine process. I would not waste my time reading the crud that came out of this thing.
It's a stunt.
FWIW PDF is actually great for distribution. It allows you to invisibly embed all the raw data used to generate the document that the end user is seeing, in whatever format you want. So if you are generating your PDFs by using PrinceXML to render HTML, you can embed the raw JSON used to generate all of the text, graphs, charts, etc. Now most people don't actually do this of course, but that isn't the fault of the spec.
With a website and available source code, any dev working on it later on can still add accessibility, tweak contrasts and fonts and add screen reader hints, etc.
It's much harder to do so for PDFs after the fact. And PDF viewer apps may or may not even support the accessibility annotations. By contrast all the major browsers and operating systems have OK support for web accessibility.
They can hold so many different types of data so that they're extremely difficult to parse.
Because of this, you can put several malicious programs into them for RCE.
That way, if someone archives many PDFs, there can be a plethora of different RCE vulnerabilities just waiting for the user to discover.
It's a wonderful dream for any malicious actor.
* /s
Even though this would only cover a small part of the needs or use cases, it will still be hugely useful if it works well.