The whole thing reminds me of the blockchain hype train. Still using and loving databases here for the foreseeable—still writing things by hand for the foreseeable, and loving every moment.
The whole thing reminds me of the blockchain hype train. Still using and loving databases here for the foreseeable—still writing things by hand for the foreseeable, and loving every moment.
Here on HN, a few days ago, there was a post about Microsoft publishing a GitHub repo that contained a "table recognizer AI". Basically, you feed it PDFs that contain horrible scanned images of finance records, and it spits out Excel spreadsheets. For some reason, Microsoft had just "thrown this over the fence" and released it to the public for free. This, despite man-years of effort developing the thing. It was working, and everything.
I made a comment wondering if Chat GPT 4 with the vision extension could solve the same problem. One of the devs that had worked on the aforementioned AI (for years!) mentioned that yes, yes it can.
Game over.
Those years of effort had just been replaced with a one-sentence English-language prompt that starts with "Please output a table from..."
If this doesn't blow your mind, then... I don't know how to help you understand just how much has changed, virtually overnight.
OmniPage Ultimate is an entire workflow that's turn-key and ready-to-go.
Generally speaking, Kofax -- or any other off-the-shelf OCR tool -- can't process tables properly. They get confused by headings, total rows, and the like. Hence the R&D effort by that Microsoft team to develop a purpose-built AI-driven tool that can specifically identify these elements and then output the result not just as ASCII text, but as a spreadsheet.
Either way, whether you are talking about the Microsoft AI or the Kofax tool, the effort is measured in man-years, man-decades, or perhaps even man-centuries.
You can just ask ChatGPT to do similar tasks for you in seconds.
A real-world use case I had for GPT is to fix up the formatting of badly copy-pasted tables. E.g.: I had created a cloud VM and pasted the summary tab into text editor, and only then noticed that the cells ended up on individual rows, interleaved with the headings. Even if I had used a regex to undo the damage, the headings mess up the alternating row interleave. E.g.:
Overview
Hardware
SKU
A2
Zone
2
Software
Operating System
Windows
Image
Licensed
Yes
Chat GPT can "undo the damage" because it understands not only what text is likely to be a heading, property, or value... but it has also seen specifically the properties of the type of cloud VM I was working on! It understands these things to the same level as I do, and can fix up the formatting the same way I would.Sit down for a second, and ask yourself: How much time and effort would you have to invest to solve this style of problem, in general. E.g.: given some random, broken table formatting for an arbitrary but well-known subject, fix up the formatting.
Years of effort?
Decades?
Whilst to the untrained eye, sure, the NGINX config looked great, but it didn't function, and it didn't include the URL rewrites as instructed, hallucinated a bunch of rewrites that didn't exist and would serve no purpose, and despite refining the prompts over many days, and consulting with self-proclaimed "prompt engineers", it still didn't give the expected output.
It's neat. Reliable? Not in my experiences for my needs, but I am genuinely glad you're making it work for what you need, that's definitely cool, provided it doesn't hallucinate in an unnoticed capacity. It's a lot of trust to place in an LLM.
The trick is to know what they can and can't do, and use them where they're useful enough.
Someone here on HN quipped that ChatGPT is "isn't intelligent" because it couldn't come up with a revolutionary new battery chemistry.
Like... what are you expecting? A god?
LLMs are useful when you can validate the output yourself. Similarly, they're useful when the output doesn't have to be precise but the inputs are English.
For example, they're awesome "filters" for human input as measured against human metrics. "Is the following text rude? Output YES or NO only?" is very useful and works well enough, right now.
They're also useful when you need to iterate to find something, or when your search parameters are incredibly vague but have a narrow Venn diagram intersection.
I've been using ChatGPT as a replacement for the /r/tipofmytongue sub-Reddit. It knows everything well enough to be able to interactively find what I'm looking for to an extent that is super-individual-human, and in some ways beyond what even a very large collection of humans can achieve.
In my opinion, the mental model you should use when evaluating LLMs-vs-Human capability: When asked a question, the LLM has to answer without being able to iterate, back-track, or even use a scratch pad such as pen&paper or a text editor. Basically it's the same as an "oral exam", where you have to stand in the middle of a room and get grilled by a professor to determine your knowledge on a subject.
Don't compare "human with tools and unlimited time" to "LLM with no tools and seconds of time". Compare "human being interrogated in an empty room" and then it is much more clear where an LLM rates.
GPT 4 is definitely super-human in some areas, such as general knowledge and translation between languages.
No human knows as much, or can speak as many languages.
Ask yourself this: can any human, when asked to "invent something", just do it, then and there?
Can you share the prompt and your expected output? I'm interested to see where it goes wrong.
TL;DR. I start by asking it to generate sentences where subject-object inversion yields a meaningful sentence where the verb's meaning is shifted metaphorically. For instance "I smoke the cigarette vs the cigarette is smoking me". After some back and forth it comes up with:
The painter captures the landscape vs. The landscape captures the painter. The gardener shapes the garden vs. The garden shapes the gardener. The chef creates the dish vs. The dish creates the chef.
Maybe that will change your mind.
And how well does it do them? Correct me if I'm wrong, but for the time being all you know about its ability to replace that other MS tool is a "yes, yes it can" by a MS dev? Don't you want to withhold judgement until after you've spent some time actually using ChatGPT to do OCR?
This is like saying you could do CAD via ChatGPT, sure, but it's an objectively crappy experience versus an actual piece of CAD software with an appropriate UX.
Using different pieces of software isn't a problem for most adults.
For now I'd treat it more like a sparring partner who can help you with your ideas and support you throughout your process, rather than a magical genie that can magically solve your problem for you. And in this manner I find it to be very useful indeed.