Let's crash the fediverse! https://wire.wise-relations.com/
33 karma · joined January 26, 2019
Let's crash the fediverse! https://wire.wise-relations.com/
We have build a web client + REST API that allows to use the API for free for small personal projects.
https://konfuzio.com/en/ocr-api/
It supports handwriting, correction of HOCR text via the webrowser, automated language detection.
We use the text to allow large enterprises to train document categorization and data extraction AI in a low/now code UI.
Disclaimer: I'm one of the founders.
Find the source code here https://github.com/konfuzio-ai/document-ai-python-sdk/blob/b...
Many other file types are supported. Have a look at https://dev.konfuzio.com/web/api.html#supported-file-types
https://colab.research.google.com/github/konfuzio-ai/documen...
Request access to the full data set via info@konfuzio.com
Disclaimer: The first half of the article provides a general framework how to use Active Learning in production. In the section "How does Active Learning work in practice?" we visualize how we implemented active learning in production using our OCR & Document AI platform. Some people might feel the latter part promotional that we add this disclaimer.
Feel free to add comments on "How are you using active learning in production?"
Elevator Pitch: Konfuzio automates complex back office processes in large enterprises and supports them to use more data to generate insights. OCR, extract, validate & process documents.
Konfuzio enables business users to first OCR, then categorize documents and finally extract information. Optionally, data scientists can use Konfuzio as an automated text and image labeling tool with a user-friendly web interface to maintain high-quality data sets to build world-class AI, integrate their AI as containerized workloads and improve them with humans in the loop.
Konfuzio enables IT experts to use AI(s) in a non-invasive way in existing processes, especially for unstructured input data like documents or emails. Helm & Nagel GmbH hosts Konfuzio in Germany. The data processing agreement is GDPR-compliant under the Technical Organizational Measures (TOMs) specified in the Terms and Conditions, which provide redistribution rights. On-prem customers deploy Konfuzio either on a single machine or their Kubernetes Cluster via Helm.
Register for free: https://app.konfuzio.com
Test our free public Python SDK: https://pypi.org/project/konfuzio-sdk/
View the changelog of our server (proprietary software): https://deep-tech.com/changelog_app/
View the on-prem documentation of our server: https://deep-tech.com/pdf/konfuzio_on_prem.pdf
Get to know our team: https://konfuzio.com/en/team/
You can copy and paste our setup in Jekyll and Gitlab pages.
IMHO existing SSG need educated users to produce sophisticated homepages not another SSG. Still jekyll feels rather complicated to use, I described my set-up issues in the post:
https://www.helm-nagel.com/how-to-score-a-perfect-100-percen...
https://www.helm-nagel.com/how-to-score-a-perfect-100-percen...
The code of this post is available on GitLab.
https://gitlab.com/c.helm/multiple-languages-fast-jekyll-exa...
Feel free to add any feedback.
thanks for your questions! We have updated the post and included the answers to your questions.
- Of course, this step can be omitted if the documents already have text embeddings. However, it is often necessary to read tables or scanned documents, for example. In our software solution, the users can decide for any project if they want to use text embeddings, Tesseract, or a commercial OCR.
- With page segmentation or also called layout analysis, we refer to the division of a document into separate parts.
- This is done with our own trained model because we couldn’t achieve the needed outcome with off-the-shelf software like Tesseract or Abbyy FineReader.
The page segmentation API is already live. The PDF summarization API is work in progress. We just wanted to share our approach already now to incorporate any feedback! We are also working on the retraining loop to fine-tune our model on a small sample of other documents. We support this for custom NER models and document classification so far.
Best Chris
Disclaimer: I am the co-founder of Konfuzio, a start-up founded in 2016 based in Germany.