Show HN: Digitizing photos of whiteboards using the command line
gist.github.com
gist.github.com
NB: This works fine on the Windows version of ImageMagick, however if (like me) you want to wrap it in a batch file you must escape the -level command by changing the "60%,91%,0.1" to "60%%,91%%,0.1" or else the batch file will misinterpret it, since %1 refers to variable names (the same as bash scripts). If you don't, you'll just get a black PNG file.
[1] : convert -fill none -draw "matte 0,0 floodfill" -type optimize -colors 64 +dither -trim -fuzz 3% +repage -strip $1 $2
Nevertheless, it's a great technique which i've just saved for such specific cases. Thanks!
[1] http://www.reddit.com/r/commandline/comments/1weqnn/cli_onel...
Anyway, I cheated by using some of these actions: http://www.fmwconcepts.com/imagemagick/
In particular the "textcleaner" script proved useful.
I'm definitely going to have a look to see if this script offers advantages over my method.
I had to sacrifice some quality in order to be able to do this with realtime video, but still, I should probably work out exactly what processing that command actually does and see if I can improve quality while still meeting the realtime requirement.
Thanks for sharing.
On a slightly related note, camscanner, an app for android (and IOS?) does something similar, but can also correct for the angle at which the photo was taken (I think it has to detect squares/rectangles in the photo). Does anyone know if this is possible using the command line too?
#!/bin/bash
convert "$1" -morphology ...
@molven, @vibragiel: Wow, I feel more than a bit ridiculous after looking into this. Turns out, those original examples I made using GIMP (same basic process, tuned the parameters to make it look nice). I'd originally put this gist together several months ago, and I hadn't done a thorough enough check before submitting to HN.
HOWEVER, I've since created new examples actually using the bash script, and the gist has been updated accordingly.
You can get these images here:
> Input1 - http://i.imgur.com/27aDJ6b.jpg
> Input2 - http://i.imgur.com/LaRWFT4.jpg
> Output1 - http://i.imgur.com/xMxM8P2.png
> Output2 - http://i.imgur.com/E3XoM3e.png
Input image - http://i.imgur.com/6o5FwxG.jpg
Output image - http://i.imgur.com/7OIOxfO.png
I tried to use vanilla Tesseract on it, but I had no luck getting anything usable out of it.
Some commercial OCR engines such as Nuance Capture SDK have built in functionality for this.
also photos of whiteboards always come out terrible, this cleans them up nicely.
Another approach I took recently was to combine this with -threshold (90 worked well in my case) and the Dilate morphology - I had a black and white set of plans printed using varying-sized dots.
Great job!
On a side note. I don't like gists. I didn't want to stray off topic so I put my complaints about gists here: https://news.ycombinator.com/item?id=7521600