This also seems like something a [convolutional] neural net should be able to learn pretty well and generating training data could be fully automated. Throw everything at it - different fonts, different sizes, bold, italic, different colors, different anti-aliasing methods, sub pixel offsets, pixelation sizes and offsets, JPEG noise, rotations, ... - and maybe also give it some language model to better reason about plausible letter combinations. I would love to see, how good this could be.