HNHacker News
TopNewBestAskShowJobs

konfuzio

33 karma · joined January 26, 2019

Konfuzio extracts, categorizes and validates information in documents. www.konfuzio.com
submissionscomments
konfuzio··on Show HN: s@: decentralized social networking over static sites
I think: A new way of old school forum would bring real UGC rel="ugc" that brings value to human and agentic readers.

Let's crash the fediverse! https://wire.wise-relations.com/

konfuzio··on Briar has been removed from Google Play
FYI: Your logo is not displayed on mobile.
konfuzio··on Show HN: My Free OCR Website – want feedback
https://centraluseuap.dev.cognitive.microsoft.com/docs/servi...

We have build a web client + REST API that allows to use the API for free for small personal projects.

https://konfuzio.com/en/ocr-api/

It supports handwriting, correction of HOCR text via the webrowser, automated language detection.

We use the text to allow large enterprises to train document categorization and data extraction AI in a low/now code UI.

Disclaimer: I'm one of the founders.

konfuzio··on Show HN: Annotate Images and Scans with BIO-Scheme Using Konfuzio SDK
Retraining NLP models, like flair, often require the data structure to be in the BIO scheme. For scanned documents or images, we convert visual annotations to the BIO scheme using OCR, we transform the bounding box to the start and end offsets of each annotation and its label. In the new release of our SDK, this conversion can be done using the method get_text_in_bio_scheme() of the Document class.

Find the source code here https://github.com/konfuzio-ai/document-ai-python-sdk/blob/b...

Many other file types are supported. Have a look at https://dev.konfuzio.com/web/api.html#supported-file-types

konfuzio··on [dead]
Find the Colab Notebook here:

https://colab.research.google.com/github/konfuzio-ai/documen...

Request access to the full data set via info@konfuzio.com

konfuzio··on Few-Shot Learning in Practice: GPT-Neo & 'HuggingFace' Accelerated Inference API
Good to see that few shot learning is now even easier using the widget. We have had good results with writing emails and reported about it in our blog: https://www.helm-nagel.com/open-sourced-gpt-Neo-writes-indiv...
konfuzio··on [dead]
Besides the resources in our article, we are aware of libraries like modAL, ALiPy, or libact as active learning frameworks. Interfaces to collect, correct and revise feedback are rare. In addition, to the general framework, on how to use Active Learning in production, we visualize how we retrieve feedback from subject matter experts/users.

Disclaimer: The first half of the article provides a general framework how to use Active Learning in production. In the section "How does Active Learning work in practice?" we visualize how we implemented active learning in production using our OCR & Document AI platform. Some people might feel the latter part promotional that we add this disclaimer.

Feel free to add comments on "How are you using active learning in production?"

konfuzio··on Show HN: Alternative to Google Document AI made in Germany
Over the past 6 years we have built an AI to process documents. We now share our current status-quo and are looking for your feedback :)

Elevator Pitch: Konfuzio automates complex back office processes in large enterprises and supports them to use more data to generate insights. OCR, extract, validate & process documents.

Konfuzio enables business users to first OCR, then categorize documents and finally extract information. Optionally, data scientists can use Konfuzio as an automated text and image labeling tool with a user-friendly web interface to maintain high-quality data sets to build world-class AI, integrate their AI as containerized workloads and improve them with humans in the loop.

Konfuzio enables IT experts to use AI(s) in a non-invasive way in existing processes, especially for unstructured input data like documents or emails. Helm & Nagel GmbH hosts Konfuzio in Germany. The data processing agreement is GDPR-compliant under the Technical Organizational Measures (TOMs) specified in the Terms and Conditions, which provide redistribution rights. On-prem customers deploy Konfuzio either on a single machine or their Kubernetes Cluster via Helm.

Register for free: https://app.konfuzio.com

Test our free public Python SDK: https://pypi.org/project/konfuzio-sdk/

View the changelog of our server (proprietary software): https://deep-tech.com/changelog_app/

View the on-prem documentation of our server: https://deep-tech.com/pdf/konfuzio_on_prem.pdf

Get to know our team: https://konfuzio.com/en/team/

konfuzio··on The computers are getting better at writing
GPT Neo works well for drafting emails. Find our approach here: https://www.helm-nagel.com/open-sourced-gpt-Neo-writes-indiv...
konfuzio··on Ask HN: Simple(st) Static Site Generator?
We watented a fast, SEO optimized, multi language page, where blog posts can be written in markdown.

You can copy and paste our setup in Jekyll and Gitlab pages.

IMHO existing SSG need educated users to produce sophisticated homepages not another SSG. Still jekyll feels rather complicated to use, I described my set-up issues in the post:

https://www.helm-nagel.com/how-to-score-a-perfect-100-percen...

konfuzio··on 100% PageSpeed Score in 4 Steps on GitLab Pages: Google SEO Experience Update
Read the blogpost:

https://www.helm-nagel.com/how-to-score-a-perfect-100-percen...

The code of this post is available on GitLab.

https://gitlab.com/c.helm/multiple-languages-fast-jekyll-exa...

Feel free to add any feedback.

konfuzio··on Automatic Text Summarization in PDF Documents with Faster R-CNN and PEGASUS
Hi Walter,

thanks for your questions! We have updated the post and included the answers to your questions.

- Of course, this step can be omitted if the documents already have text embeddings. However, it is often necessary to read tables or scanned documents, for example. In our software solution, the users can decide for any project if they want to use text embeddings, Tesseract, or a commercial OCR.

- With page segmentation or also called layout analysis, we refer to the division of a document into separate parts.

- This is done with our own trained model because we couldn’t achieve the needed outcome with off-the-shelf software like Tesseract or Abbyy FineReader.

konfuzio··on Automatic Text Summarization in PDF Documents with Faster R-CNN and PEGASUS
We collected all models in our documentation on https://deep-tech.com/training_documentation.html#text-modul... Extractive PEGASUS is not yet there. What do you exactly mean by highlighting/underline?
konfuzio··on Automatic Text Summarization in PDF Documents with Faster R-CNN and PEGASUS
Hi notafraudster, is this dataset or your approach public? Perhaps we can can collaborate to expand our approach. FYI: We detect text embeddings automatically and decide thereby if we need OCR. Thanks for the feedback!
konfuzio··on Automatic Text Summarization in PDF Documents with Faster R-CNN and PEGASUS
Hi David, thanks for reporting the link issue! We fix it. It should be https://app.konfuzio.com

The page segmentation API is already live. The PDF summarization API is work in progress. We just wanted to share our approach already now to incorporate any feedback! We are also working on the retraining loop to fine-tune our model on a small sample of other documents. We support this for custom NER models and document classification so far.

Best Chris

konfuzio··on What's so hard about PDF text extraction?
PDF are a pain, and Konfuzio is an AI software to make PDF content machine readable. On our journey to structure PDF content or even scans we have been supported by large enterprises in the banking, insurance and audit industry. We are in closed beta for data scientists, so feel free to request a demo and free user account.

http://www.konfuzio.com

Disclaimer: I am the co-founder of Konfuzio, a start-up founded in 2016 based in Germany.

konfuzio··on Show HN: Eiten – open-source tool for portfolio optimization
Does the optimization include rebalancing + trading costs?
konfuzio··on Show HN: Pdf.js Express – PDF annotation, e-signatures, and form filling
Hi mmckelvy, Co-Founder of konfuzio.com here. If you are looking for a tool to extract information from documents and download results via API or CSV ping me. You can train your custom AI Model using our PDF Annotation Tool. There is no coding required. Using your documents as training data you can go live within a few hours or use our pre-trained models. We don't have pricing plans on our homepage, as we mainly focus on large enterprises at the moment. But I would love to hear your feedback. Best Chris