Marker: Convert PDF to Markdown quickly with high accuracy
github.com
github.com
For anyone trying to do OCR on any pdf with math in it, definitely do try nougat. It's very easy to install (just a python package), and extracts the math, text, tables and beyond (in a .mmd file) with a single command line command. It also runs reasonably fast for personal uses - it takes about 30 seconds to convert a 6 page document using CPU only on my 4 year old i5 laptop.
I'm looking for a food OCR model to help me transcribe sections of RPG books to markdown. Ideally, I'd like the emphasis such as bold or italics to be transcribed.
The combo of text, numbers, and math symbols seems similar to technical and academic writing, but often has weird formatting, text boxes in the margins, and many diagrams.
Comparing two things doesn't inherently imply the previous thing was touted about with superlatives. It's just a way to juxtapose the new thing with something that may be familiar. As you said, nougat is easy to install/run so it makes sense they'd compare it. Would it be better if they could add more libraries in the comparison? Absolutely; that'd be helpful.
Nougat is a great model, and converts a lot of PDFs very well. I just wanted something faster, and more generalizable.
I noticed marker downloaded a PyTorch checkpoint called `nougat-0.1.0-small`, do you use nougat under the hood too or is that just a coincidence?
I want to extract financial statements from pdfs which are in tables, would Nougat be suitable for that use case?
I'm very excited about it.
Let's build a pipeline: all the pdfs -> markdown them all -> archive.org them all
Even though this would only cover a small part of the needs or use cases, it will still be hugely useful if it works well.
FWIW PDF is actually great for distribution. It allows you to invisibly embed all the raw data used to generate the document that the end user is seeing, in whatever format you want. So if you are generating your PDFs by using PrinceXML to render HTML, you can embed the raw JSON used to generate all of the text, graphs, charts, etc. Now most people don't actually do this of course, but that isn't the fault of the spec.
With a website and available source code, any dev working on it later on can still add accessibility, tweak contrasts and fonts and add screen reader hints, etc.
It's much harder to do so for PDFs after the fact. And PDF viewer apps may or may not even support the accessibility annotations. By contrast all the major browsers and operating systems have OK support for web accessibility.
They can hold so many different types of data so that they're extremely difficult to parse.
Because of this, you can put several malicious programs into them for RCE.
That way, if someone archives many PDFs, there can be a plethora of different RCE vulnerabilities just waiting for the user to discover.
It's a wonderful dream for any malicious actor.
* /s
I don't think that is the right approach for archiving. The preferred pipeline would be
all the pdfs -> archive them all -> markdown them
This way you can always re-run the conversion as bugs are fixed and improvements are made. Generally archivist prefer to save as close to the source material as possible, because every transformation from there can only lose data.
I opened the first example to a random chapter (1.4 Formal and natural languages); within the first three paragraphs it:
- Hallucinated spurious paragraph breaks
- Ignored all the boldfacing
- Hallucinated a blockquote into a new section
This is not a tool to produce something for humans to read.
Maybe it might be useful as part of some pipeline that needs to feed markdown into some other machine process. I would not waste my time reading the crud that came out of this thing.
It's a stunt.
I regularly hand transcribe RPG PDFs scans from dubious sources that have not always been run through OCR to have selectable text. If it has, it wasn't always done very well.
It's literally faster to type it all myself than fix all the errors from copy-pasting (or after using OCR to turn it into text).
Even if the file was an official PDF the formatting would often get screwed up with lots of double or triple spaces and even tabs included between words.
This would save so much time if I can get it to work. Thanks for sharing!
Heh, that was my immediate thought too. There's a ton of RPG stuff that never had any kind of physical release and is totally orphaned as IP.
https://github.com/tesseract-ocr/tesseract/releases
I suppose it depends on your use-case. For personal tasks like this it should be more than sufficient, and won't need user details/cc or whatever to use it.
In practice, I would use the Markdown output and plug it into any tool that converts that into the desired final output format.
I wonder if this could somehow be used directly by calibre. I think calibre's pdf->epub conversion isn't amazing. In particular, tables often end up broken.
I have a question regarding the output of Nougat: Where do the "hallucinations" come from (just scroll through the Nougat output of the Think Python example to see what I mean)?
Nevermind, i just read it runs it through an LLM, so hallucinations are par for the course.
What kind of PDF are you tweaking it for? How does it handle handwritten annotations?
I have a set of PDF files, and this week was thinking how I can link them to an LLM and be able to ask questions about them. So this was very timely.
I did a quick side-by-side testing against Nougat, and it clearly works better. On a handful of PDFs I tested, Marker extracted considerably more text (the text did not have any math, just academic papers), finished the job faster, and did not crash on any pdf, while Nougat took a lot longer to finish, and sometimes crashed due to out-of-memory error (could not allocate more than 7GB RAM!)
Random old issue for example: https://thetech.com/issues/33/34
>Due to the licensing of the underlying models like layoutlmv3 and nougat, this is only suitable for noncommercial usage.
Does this mean it isn't suitable if I wanted to use it in a product for sale or I cannot use it for tasks at my work? I would like to try to use this at work to convert vendor documentation to include in our internal wiki.
I've been through a number of options in the past and this is what I've settled on.
My workflow still takes manual tweaking. When I find floated figures with captions, the lines get intertwingled and need to be unintertwingled. So I'm not surprised it didn't work for you.
Good luck, report back if you find what you're looking for. I'm always on the lookout for a better way.
The other option I have started looking into is the PDFCPU library for Go. It is a bit more low-level than PDFMiner, but one gets out very well structured info, that seem it might be possible to post-process quite well, for one's particular use case and PDF layouts: https://github.com/pdfcpu/pdfcpu
I also now tried the Marker tool in the OT, and it seems to do a reasonable job. It did intermingle some columns though, at least in some tricky cases such as when there were a round shaped image in between the two columns. One note is that Marker doesn't seem to retain styling like italics though.