OmniPage Ultimate is an entire workflow that's turn-key and ready-to-go.
OmniPage Ultimate is an entire workflow that's turn-key and ready-to-go.
Generally speaking, Kofax -- or any other off-the-shelf OCR tool -- can't process tables properly. They get confused by headings, total rows, and the like. Hence the R&D effort by that Microsoft team to develop a purpose-built AI-driven tool that can specifically identify these elements and then output the result not just as ASCII text, but as a spreadsheet.
Either way, whether you are talking about the Microsoft AI or the Kofax tool, the effort is measured in man-years, man-decades, or perhaps even man-centuries.
You can just ask ChatGPT to do similar tasks for you in seconds.
A real-world use case I had for GPT is to fix up the formatting of badly copy-pasted tables. E.g.: I had created a cloud VM and pasted the summary tab into text editor, and only then noticed that the cells ended up on individual rows, interleaved with the headings. Even if I had used a regex to undo the damage, the headings mess up the alternating row interleave. E.g.:
Overview
Hardware
SKU
A2
Zone
2
Software
Operating System
Windows
Image
Licensed
Yes
Chat GPT can "undo the damage" because it understands not only what text is likely to be a heading, property, or value... but it has also seen specifically the properties of the type of cloud VM I was working on! It understands these things to the same level as I do, and can fix up the formatting the same way I would.Sit down for a second, and ask yourself: How much time and effort would you have to invest to solve this style of problem, in general. E.g.: given some random, broken table formatting for an arbitrary but well-known subject, fix up the formatting.
Years of effort?
Decades?
Whilst to the untrained eye, sure, the NGINX config looked great, but it didn't function, and it didn't include the URL rewrites as instructed, hallucinated a bunch of rewrites that didn't exist and would serve no purpose, and despite refining the prompts over many days, and consulting with self-proclaimed "prompt engineers", it still didn't give the expected output.
It's neat. Reliable? Not in my experiences for my needs, but I am genuinely glad you're making it work for what you need, that's definitely cool, provided it doesn't hallucinate in an unnoticed capacity. It's a lot of trust to place in an LLM.
The trick is to know what they can and can't do, and use them where they're useful enough.
Someone here on HN quipped that ChatGPT is "isn't intelligent" because it couldn't come up with a revolutionary new battery chemistry.
Like... what are you expecting? A god?
LLMs are useful when you can validate the output yourself. Similarly, they're useful when the output doesn't have to be precise but the inputs are English.
For example, they're awesome "filters" for human input as measured against human metrics. "Is the following text rude? Output YES or NO only?" is very useful and works well enough, right now.
They're also useful when you need to iterate to find something, or when your search parameters are incredibly vague but have a narrow Venn diagram intersection.
I've been using ChatGPT as a replacement for the /r/tipofmytongue sub-Reddit. It knows everything well enough to be able to interactively find what I'm looking for to an extent that is super-individual-human, and in some ways beyond what even a very large collection of humans can achieve.
In my opinion, the mental model you should use when evaluating LLMs-vs-Human capability: When asked a question, the LLM has to answer without being able to iterate, back-track, or even use a scratch pad such as pen&paper or a text editor. Basically it's the same as an "oral exam", where you have to stand in the middle of a room and get grilled by a professor to determine your knowledge on a subject.
Don't compare "human with tools and unlimited time" to "LLM with no tools and seconds of time". Compare "human being interrogated in an empty room" and then it is much more clear where an LLM rates.
GPT 4 is definitely super-human in some areas, such as general knowledge and translation between languages.
No human knows as much, or can speak as many languages.
Ask yourself this: can any human, when asked to "invent something", just do it, then and there?
Can you share the prompt and your expected output? I'm interested to see where it goes wrong.
TL;DR. I start by asking it to generate sentences where subject-object inversion yields a meaningful sentence where the verb's meaning is shifted metaphorically. For instance "I smoke the cigarette vs the cigarette is smoking me". After some back and forth it comes up with:
The painter captures the landscape vs. The landscape captures the painter. The gardener shapes the garden vs. The garden shapes the gardener. The chef creates the dish vs. The dish creates the chef.
Maybe that will change your mind.
And how well does it do them? Correct me if I'm wrong, but for the time being all you know about its ability to replace that other MS tool is a "yes, yes it can" by a MS dev? Don't you want to withhold judgement until after you've spent some time actually using ChatGPT to do OCR?