Linear Book Scanner – Open-source automatic book scanner (2014)
linearbookscanner.org
linearbookscanner.org
Years ago I once wrote a little tool in Java called bookbuilder, where you could turn the pages manually, make a photo and then run an automatic process on all images to build a searchable pdf.
I used https://boofcv.org/, an impressive Computer Vision library in pure Java, still exists and it is pretty fast, too.
It was able to detect the page contour, deskew it, flatten the image and remove finger contours by matching the skin tone, then build a PDF with integrated invisible OCR Layer without any user interaction. I remember that I was working on line slope detection with some kind of watershed algorithm to improve the flattening part.
Fun project, I wonder if I have the source code laying around somewhere... even the download page is gone today. This was long before I went open source with all of my little side projects, because I never thought it could be interesting for someone else :-)
I just spent 30 minutes clicking through and inspecting every example.
When I saw the linear bookscanner the first time, I realized it was upside down. If it were suspended, and able to move itself, then all the issues about not knowing the mass of the book, and dealing with friction could be avoided. Counterweights could keep the force constant, and it would be moving a known mass when scanning any book.
[1] https://www.treventus.com [2] https://youtu.be/SdipuAuWsEs?si=dFWRtva5gO2oM91o
Unless you're concerned with binding the pages to form a new book. I think that would be possible with the leftovers.
Not sure how much of the design is protected and how much inspiration one can take for a non-commercial DIY project like the one presented by OP.
Regarding frequency of torn pages in the FAQ:
> Prototype 1 could scan the majority of books without damage, but may tear one or two pages in some books. Out of 50 books tested, 45% had one or two of their pages either torn or folded. This is a very early prototype and there are many areas for improvement in the design.
In my opinion, this is mostly acceptable. Especially if a future revision reduces the 45% to somewhere around the ~10-20% range. If I had the space for a device like this, I would definitely consider building one.
https://blog.archive.org/2021/02/09/meet-eliza-zhang-book-sc...
[0] https://en.m.wikipedia.org/wiki/Authors_Guild,_Inc._v._Googl....
Not if the book is irreplaceable, e.g. old and out of print. No torn pages are acceptable.
And before computers, there was no electronic copy.
With two or more photos or a stereo image (new iPhone?) one could triangulate to infer a flattened page, and produce images that look like they came from cut pages in a flatbed scanner. Now just pay someone well in Ethiopia to carefully turn pages without damage.
As any researcher can attest, our digital libraries now hold a century of scanned work of questionable quality. AI could infer scans indistinguishable from an outline font format original on an 8K monitor.
I once helped consult on the 1980's font wars, turning old formats and digital scans into Postscript and TrueType fonts. This was hard then, but will soon be understood as the "correct" way to scan text, when software catches up.
For the scientific literature, we need a ChatGPT equivalent to reconstruct LaTeX source that can reproduce each page. (We really need a successor to LaTeX that isn't such an arcane language, and can author fixed and flowable text with equal ease.)
I think the problem isn't that software can't do de-warping well, it's that by the time you set up everything for book scanning you might as well use a setup that doesn't need it.
Even for relatively high-mass objects such as books. Slow boats are slow but exceedingly efficient.
The main risk would likely be container loss off a ship. Possibly environmental damage if spending much time in warm humid climates.
Working a typical volume the letter “e” will appear hundreds of times and be identical, so there should be lots of data to help resolve ambiguities in the poorer parts of images.
Not to mention data that can be used across volumes.
That being said, modern phone cameras are going to produce "scans" above 300 DPI, and while 600 DPI or higher might be tricky they're stills possible if you take partial shots of a document, assuming you can focus that close.
What you lose in quality you make up with convenience, I suppose.
I've been part of a preservation project and scanned a LOT of magazines. Generally, it doesn't matter much if you place it on a flat bed or not. Flipping the page manually takes a while either way.
There's already a known way for scanning magazines and books very fast: cut the spine and feed the pages to an automatic scanner. This is of course not applicable to anything you'd like to keep around after scanning, because your copy is destroyed.
All in all, the best way to automate scanning without destroying the item, will have to combine a top level camera with a machine to turn pages. I believe this is what was going on in Google's massive scanning project.
Maybe using x-ray could work for "scanning" some books without having to turn the pages. But I suppose there'll be a new set of problems to solve there.
Of course, that only gets you the printed text. You might lose notes and doodles in the margins, or other physical evidence. But, it's certainly promising for works that are too delicate to physically open and inspect
I keep thinking so much collaborative potential is not being utilized. Imagine each (unique) Google Book is basically an editable wiki where people can directly correct OCR errors as they come across them (with an associated Talk page where they can give explanations, etc)
https://www.theatlantic.com/technology/archive/2017/04/the-t...
The insurmountable problem is that Google doesn't have, and can't get the rights to distribute all these books. Even regular Google employees can't see them (and I tried when in Legal there).
So this is neither a hardware problem nor a software problem. Unfortunately.
Very interesting read. Thanks for sharing it.
On the software side though, progress marches on: https://facebookresearch.github.io/nougat/ is downloadable and great.
https://news.ycombinator.com/item?id=29223815
That risk of "plausible but incorrect" has been a concern even before AI.
Especially if you project a grid over the pages using e.g. a scanning laser.
Here you can see how flattening works (this was handled by the library, we didn't need to do any custom code):
https://youtu.be/DPu0iJtK2sI?t=1542
There's also a feature where it tells you to turn the page, detects that it has been turned, takes a photo, etc. And in the background it flattens, splits into pages and OCRs the photos. With a little practice you can scan and OCR a whole book at 1-5 seconds per page.
https://youtu.be/DPu0iJtK2sI?t=1909
Then it saves the OCRed book and it can read it to you whenever you like.
Check out Nougat: OCRing scientific papers with a deep net trained end to end. It was released by Meta a few days ago.
“PDF format leads to a loss of semantic information, particularly for mathematical expressions. We propose Nougat (Neural Optical Understanding for Academic Documents), a Visual Transformer model that performs an Optical Character Recognition (OCR) task for processing scientific documents into a markup language, and demonstrate the effectiveness of our model on a new dataset of scientific documents.”
The only valid archival approach has to be taking a good photo / scan while the page is as flat as possible.
Every further processing can be done later based on that, as a separate step... as technology advances.
Also carefully FLIPPING THE PAGES is literally the whole problem. Everything else can be solved by lowering a glass pane on the book and automatically taking a photo from above.
https://www.inforum.com/newsmd/ndsu-students-book-scanner-in...
But again, for cheap mass-market works, most archivists probably won't care about destroying one copy out of a million to preserve the work. It's really only a problem for very old and very rare works
Oooh.. Second question in https://linearbookscanner.org/faq/...
> Out of 50 books tested, 45% had one or two of their pages either torn or folded
I would not use this even on my less valuable books.
They used low paid labor to flip the pages. I'll send a link later if no one beats me to it.
Back home: it's https://www.theatlantic.com/technology/archive/2017/04/the-t...
Pretty sweet 2 for 1 deal.
[0] https://www.youtube.com/watch?v=84byulcC6i4 30 seconds long
In the fullness of time, maybe I would make one of these, given that I live in an apartment and not a house with space to construct/store this scanner. I check etsy every couple of years and haven't seen someone offer a kit. I use 1dollarscan, though they've had to restrict their offering as Pearson, et al notice their existence.