DeepDoctection: Document extraction and analysis using deep learning models
github.com
github.com
An example use case: given a url, figure out content type (article vs product page for example), if product page, automatically extract all product details and specs without manually mapping xpath or css lookup paths
The mozilla/readability library is a good first step though.
Edit: I'm thinking of something like LlamaIndex
Yes, I could try to figure out which bits of the page need to be sent along with the prompt, but that's hard in the general case. Squeezing a bit more out of the prompt window by stripping out unnecessary boilerplate is easy by comparison. (Multi-page articles are another headache).
There was a show HN here a couple of days ago that used chatgpt to create a scraper based on a webpage. Hook the 2 together and you are basically there!
The biggest enterprise users are doing thousand+ of pages a minute and also turn document extraction into a scaling distributed systems problem
[1]: https://www.ibm.com/cloud/blog/exploring-ibms-new-optical-ch...
If there are any other libraries folks have seen out there like this, I’d love to try them out.
I’m not sure if they’re being intentionally annoying or if someone thought this was actually helpful for the thousands of independent contractors who track their expenses down to the penny?
However, fundamentally I completely agree with you. Information we seek should not be bound to the medium it is stored on in this day an age. I wish we could get out of the containerized knowledge but it seems to me we are creating ever more virtual containers in which information is stored. I for one only get a glimpse of the vast amounts of information TikTok is making available to it's users when it is posted on one of the few websites I visit.
I guess the reason we still think of information being in books and on paper is because we are human and its hard to get rid of millennia of habits and institutions that have grown around us to accommodate for our limited ability to grasp the universe.
Since pdfs are created so many different ways, do you have some examples and links of the pdfs that are awful?
We have a file conversion API that supports DOC/DOCX/ODT/PDF/TEX to Markdown conversion in one line of cURL (or you programming language of choice).
(Disclaimer: I'm the product lead for the Zamzar API).
* AWS Textract [0]
* Microsoft Azure [1]
* Google Cloud Vision [2]
I personally use Azure, combined with OCR correction using GPT to convert a scan of my daily journal (Apple Notes creates a PDF that is nothing but a bunch of images) -> Markdown -> Extract tasks and then add them to my Reminders app using CalDav. Azure has one of the best OCR for handwritten text, but for normal document extraction (read: printed text), any service would do a reasonable job.
[0] https://docs.aws.amazon.com/prescriptive-guidance/latest/pat...
[1] https://learn.microsoft.com/en-us/azure/data-factory/solutio...
This is the prompt I bought from promptbase. You basically provide GPT with some examples on possible OCR errors, and then you give it the OCRed text and it tries to correct it
An LLM is useless (or not as useful) in OCR for forms where we are trying to extract the name "John Smith" from a name field whereas a HMM trained specifically on name fields may be able to do a better job.
LLMs may perform better as a post processing step for running text (such as book pages).
This looks interesting as well. Haven't tested yet.