OCR is VERY good
OCR is VERY good
—-
The following is the declaration of James Lambert, a soldier of the Revolutionary War in North America.
The said James Lambert on this day personally appeared in the Probate Court of the County of Dearborn in the State of Indiana and at the November Term of said Court (1841), it being a court of record established by the laws of Indiana and made oath that:
On the 25th day of March 1842 he will be eighty-five years old; that he was born in the State of Maryland; that he is now a resident of said county and has been for the 27 years last past; that he has lived in Virginia, Maryland, Pennsylvania…
—-
Considering the people involved are experts in their field, are certainly aware of OCR capabilities, and have publicized a need thusly:
... the National Archives is looking for volunteers who can
help transcribe and organize its many handwritten records ...
Perhaps "random humans" can perform tasks which could reshape your belief:> OCR is VERY good
Asserting an ulterior motive without supporting proof is to engage in conspiracy theories.
Sometimes a cigar is just a cigar.[0]
If it's that easy, then do it and be the hero they want.
Or maybe, just maybe, "a pretty big chunk of their problem appears to be well within the state of the art" is a sweeping generalization lacking understanding of the difficulties involved.
This is a strawman[0] argument. You proclaimed:
A lot of what they want transcribed is totally
straightforward to OCR
And I replied: If it's that easy, then do it and be the hero
they want.
So do it or do not. Nowhere does my finding "something hard" have any relevance to your proclamation.(I don't think you need to Wikipedia-cite "straw man" on HN).
Awesome.
Can you guarantee its results are completely accurate every time, with every document, and need no human review?
> I'm sorry, but I declare the burden of proof here to be switched.
If you are referencing my stating:
If it's that easy, then do it and be the hero they want.
Then I don't really know how to respond. Otherwise, if you are referencing my statement:> Perhaps "random humans" can perform tasks which could reshape your belief:
>> OCR is VERY good
To which I again ask, can you guarantee the correctness of OCR results will exceed what "random humans" can generally provide? What about "non-random motivated humans"?
My point is that automated approaches to tasks such as what the National Archives have outlined here almost always require human review/approval, as accuracy is paramount.
> (I don't think you need to Wikipedia-cite "straw man" on HN).
I do so for two purposes. First, if I misuse a cited term someone here will quickly correct me. Second, there is always a probability of someone new here which is unaware of the cited term(s).
> > If it's that easy, then do it and be the hero they want.
> Then I don't really know how to respond.
If someone says a thing is easy, and you respond by demanding they do it a million times to prove that it's easy, you are the one that has screwed up the burden of proof.
What I was trying to convey is that no matter what tooling is employed for this specific problem, a human will have to review the result to ensure accuracy. Since the problem is defined as reading cursive handwriting, people who can do so intrinsically obviate the need for tooling.
Whether or not automation can produce reasonable results is moot when the requirement is to have as accurate as possible a one-time transcription of cursive handwriting documents.
That's not a claim that processing the entire archive would be trivial. And even if it was, whether that would make someone the "hero they want" is part of what's being called into question.
So your silly demand going unmet proves nothing.
Also, "give me an example please" is not a strawman!
If you actually want to prove something, you need to show at least one document in the set that a human can do but not a machine, or to really make a good point you need to show that a non-neglibile fraction fit that description.
I made demands of no one.
> Also, "give me an example please" is not a strawman!
My identification of the strawman was that it referenced "find something hard" when I had said "be the hero they want" and that what is needed in this specific problem domain may be more difficult than what a generalization addresses.
> If you actually want to prove something, you need to show at least one document in the set that a human can do but not a machine, or to really make a good point you need to show that a non-neglibile fraction fit that description.
Maybe this is the proof you demand.
LLM's are statistical prediction algorithms. As such, they are nondeterministic and, therefore, provide no guarantees as to the correctness of their output.
The National Archives have specific artifacts requiring precise textual data extraction.
Use of nondeterministic tools known to produce provably incorrect results eliminate their applicability in this workflow due to all of their output requiring human review. This is an unnecessary step and can be eliminated by the human reading the original text themself.
Does that satisfy your demand?
Whatever you want to call "If it's that easy, then do it"
> LLM's [...] Does that satisfy your demand?
That's a different argument from the one above where you were trying to contradict tptacek. And that argument is flawed itself. In particular, humans don't have guarantees either.
> provably incorrect results
This gets back to the actual request from earlier, which is showing an example where the machine performs below some human standard. Just pointing out that LLMs make mistakes is not enough proof of incorrectness in this specific use case.
I respect what you have chosen to contribute to this conversation, but neither need nor seek your approval of mine.
Experts are asking for the help of non experts.
> Anyone with an internet connection can volunteer to transcribe historical documents and help make the archives’ digital catalog more accessible
I'm largely aligned with your interpretation of "random humans", with a clarification below. The experts I was referencing are the ones you identified:
> Experts are asking for the help of non experts.
The call to action by the archivists (experts), IMHO, has the intent to engage people with interest in the topic. So not really random from a mathematical definition, but perhaps better thought of as "unknown interested parties."
Granted, this is my unsubstantiated opinion.
I do.
> OCR is VERY good
Uh, my experience is extremely different.
Denying/downvoting reality is always an option, of course.
And BugsJustFindMe can't downvote you, because it was a reply to him. So don't bite his head off over it. You got downvoted because you were a jerk, plain and simple.
Refraining from reflexively pooh-poohing AI with uninformed and/or out-of-date opinions is also an option, but not one often exercised on HN.
It gets old not being able to carry on a discussion without squinting at grayed-out text, simply because someone pointed out that humans aren't robots and should no longer have to emulate them.
It gets them wrong for me, but maybe it will get them right for you. Maybe you're better at prompting or have access to a better model or something.
Still, here's the first one, via Gemini 2.0 experimental: https://i.imgur.com/HtnwfHp.png
How does the response look? Did it correctly identify the language as Old French, at least? Even if 100% made up, which I have a feeling it is, it's a more credible (not to mention creative) attempt than most non-specialists would come up with.
o1-pro, on the other hand, completely shat the bed: https://i.imgur.com/mivdjkA.png I haven't seen it fail like that in a LONG time, so good job, I guess. :) I resubmitted it by uploading the .jpg directly, and it mumbled something about a "Problem generating the response."
Second image:
Gemini 2.0 seemed to have more trouble with this one: https://i.imgur.com/oEktMP6.png
o1-pro gave another error message, but 4o did pretty well from what I can tell (agree/disagree?): https://i.imgur.com/7iR1y7U.png I thought it was interesting that it got the date wrong, as '1682' is pretty easy to make out compared to much of the text.
In summary, I think you broke o1-pro.
Yes! But that's the easy part. :)
> I was talking about OCR'ing modern English cursive handwriting
Yeah, see, I think that's a very narrow expectation. Archive paleography is substantially broader than that. I'm not saying that the tools are useless, but they're often still not better than humans directing focused care and attention.
> o1-pro, on the other hand, completely shat the bed
The result is absolutely hilarious though! So kudos to the model for making me laugh at least.
> 4o did pretty well
It is indeed pretty good and very impressive as a technological feat. The big problems I guess are:
1) Pretty good isn't necessarily good enough.
2) If one machine gets it right and one machine gets it wrong, can a machine reconcile them? Or must we again recruit humans?
3) If a machine seems to get a lot right but also clearly makes important factual errors in ways where a human looks and says "how could you possibly get this part wrong, of all things?" (like the year), how much do we trust and rely on it?
It seems likely that a mixture-of-models approach like this will be a good thing to formalize at some level. Using appropriately-trained models to begin with seems even more important, though, and I can't agree that this type of content is relevant when discussing straightforward OCR tasks on modern languages.
1682 is a number though, language independent, and you noted it as being extremely obvious to a human, even one who can't read any of the other language. So I do think the tools are useful, but people probably still need to be there for now until better models for this are made that stop getting especially obvious parts wrong.
> The agency uses artificial intelligence and a technology known as optical character recognition to extract text from historical documents. But these methods don’t always work, and they aren’t always accurate.
They are absolutely aware of the advances in these tools, so if they say they're not completely there yet I believe them. One likely reason is that the models probably have less 1800s-era cursive in their training set than they do modern cursive.
It's likely that with more human-tagged data they could improve on the state of the art for OCR, but it's pretty arrogant to doubt the agency in charge of this sort of thing when they say the tech isn't there yet.
I don't necessarily agree with her conclusion because she wasn't participating directly in the thread and wasn't completely responsive to some of the points raised, but still, it appears that there are a few instances of difficult-to-read handwriting where OCR is still coming in second to skilled human interpretation.
If you want to provide a good faith answer at least make it English. I assume this is French but it’s obviously much harder to evaluate on both ends when you’re mixing up the language.
Which parts of "OCR" and "human" stand for "modern english"?
Are you suggesting that humans can't read or write in french? Because I can point to a lot of them who would disagree.
Are you aware of CAPTCHA[0] images?
https://github.com/noCaptchaAi/NoCaptcha-Ai-Browser-Extensio...
The original assertion was:
I would challenge you to find a picture of text
that you think a human can read and OCR cannot.
Not if many CAPTCHA image challenges could be automated. Unless the tool referenced guarantees 100% correct solutions for all manipulated text images.As long as that's the case, CAPTCHAs probably won't be considered truly obsolete.
What does your OCR say that these say? The first one isn't too hard for a human (assuming appropriate language skill). The second one is a bit more difficult.