DIY Book Scanner
diybookscanner.org
diybookscanner.org
"Google hereby grants to you a perpetual, worldwide, non-exclusive, no-charge, royalty-free, irrevocable (except as stated in this section) patent license to make, have made, use, offer to sell, sell, import, transfer, and otherwise run, modify and propagate this design..."But it worked, I scanned about two dozen short term library books that I needed to reference frequently during my course at a cost of about $85 per book. If I’d purchased the time limited ebooks they would have cost $125 each, and been scattered across 3 different bookstore apps.
I would scan while watching tv and could do approx 1000 pages per hour.
I also learned that I should not do carpentry and potentially saved tens of thousands by hiring a handyman or carpenter for home diy…
Still no regrets, I had a fun week of arts and crafts and got to stick it to Elsevier and other academic publishers :D
You are also making a very good point with the surprising effort and equipment this can take. I ran into trouble just fixing something on a heatsink the other day. Turns out in addition to drill, drillbit and tapping set i really need something to keep the drill straight :)
I did the same just with my bare hands and my smartphone with a rather short book and calculated it would take me ~2 weeks to create an imperfect (thumbs included) digitized copy of all my books. So that's what I did, eventually improving upon things:
- took a grill from the oven which stabilized the phone and relieved my arm
- created a couple of bash scripts to automate slicing and image compression
- run tesseract-ocr (fulltext search)
- ghostscript for making it a pdf All automated and improved over time. No big bang, just trial and error and tiny steps.
Meanwhile I have hundreds of books. They're not perfect but perfectly readable/searchable. Why I am telling you this? Keep it simple. Unless you're more interested in engineering the "machine" than the actual product it's supposed to create.
“The nature of moonlighting work is in direct conflict with the company. For example, working with or providing consultations to competitors, competing for an employer’s clients, or activities that harm the goodwill or reputation of your employer.”
https://www.aegislawfirm.com/blog/2023/01/california-moonlig...
I think it would be hard to argue that a book scanner competes with an online ebook store, since one is an archival tool and one is a commercial store. Someone could host illegal copies of copyrighted works and they would be competing with apple’s legitimate work, but those people would be the ones competing with apple, not the creator of a tool they used in the process.
DIY Book Scanner - https://news.ycombinator.com/item?id=27361815 - June 2021 (124 comments, and btw a great thread)
DIY book scanning - https://news.ycombinator.com/item?id=991897 - Dec 2009 (7 comments)
miss ya old friend
I was part of a font consulting company during the Postscript / Truetype font wars, and we reconstructed fonts from scans or earlier digital formats. Most of the work was fixing bad data. This all should be easier now; think of Peter Jackson cleaning up the "Let It Be" sound, leading to the Beatles releasing one more track.
It baffles me that book images don't get this quality of attention. As a mathematician I spend a lot of time reading old journal articles that look terrible.
Do these scanning rigs lock the mirror and shutter of the DSLR?
If not, what MTBF are they looking at, when prosumer DSLR shutter life might be around 50K actuations?
This top-search-hit other site has some "Average number of actuations after which shutter died" data is for some older models.
Consumer (lowest, 69K): https://olegkikin.com/shutterlife/canon_eos1000d.htm
Prosumer (98K): https://olegkikin.com/shutterlife/canon_eos30d.htm
That's average, so, if that data is reasonably representative of units in the the wild (I don't know), I'd think a trustworthy rating (and safe expectation) would be lower than that.
The reasons I mentioned shutter life was because I wanted to know how the scanning projects using DSLRs managed that, and also, to suggest to anyone dropping money on a DSLR for this that shutter life might be a cost consideration.
External power for the cameras is quite neat.
edit: In case anyone is curious, for battery replacements the term is (dual) "battery eliminator"
Nowadays I just use https://1dollarscan.com. Turns out to be rather expensive, but still beats all that manual work.
But you can also just set the book on a table, open it, and photograph it with your phone camera. The result is perfectly legible on your monitor.
I also wonder if using the LIDAR from iPhones for example could significantly improve the flattening process.
Sounds like a repetitive motion injury waiting to happen after you get done scanning Fountainhead.
https://www.youtube.com/watch?v=kvM-tjrS2-U
That said, a lot of automation processes in place do destroy the book by cutting the spine and scanning it.
> Better text detection by combining multiple OCR engines (EasyOCR, Tesseract, and Pororo)
FWIU there's a way to image multiple stacked sheets of ancient scrolls without unrolling them?
I may not have the time and energy to do a DIY, but have an immediate requirement to procure one. Any leads / directions / websites would be really appreciated.
Thank you!
But none of them are available or selling it currently.
Also, the forum looks dead / defunct.
No noticeable activity on any thread! Sad to see such a vibrant community gone extinct.
I know it is not an active area unlike other DIY and the requirements are also very low. But still....!
https://www.pfu-us.ricoh.com/scanners/scansnap/sv600
The software has automatic page-turn detection (so you don’t have to repeatedly press the scan button); has page-curve correction and deskew; and automatically removes fingers/thumbs from the image, in case you need to hold the pages down. Neat!
Like another commenter, I used 1dollarscan to digitize many books (to save space) but I agree that that process was more expensive than expected (and destroyed my physical books, which I have come to regret). If I had known about the scanner I just linked to, I probably would have invested in one instead.
Off-topic, but apparently Ricoh has acquired the ScanSnap brand from Fujitsu. (News to me, at least.) But unless Ricoh has changed something, in my experience it’s hard to go wrong with the ScanSnap brand for personal scanning needs.
(I have no affiliation with the companies or brands mentioned.)
maybe we'll find a method for "CT scanning" a book and using imaging techniques to reconstruct the text inside without needing to flip each page?
YES! Quite literally what's happenening here, albeit in a different context:
https://www.nytimes.com/2023/10/12/arts/design/herculaneum-s...
Ebooks are basically all epub and that's completely useless for type setting. I've had to contact authors of textbooks to try and get the latex source so I can read what the damned thing says on a screen.