Traditional OCR pipeline would be to use some heuristics to find line of text in an image, use some other heuristics to break line of text into candidate characters. Some candidate blobs may need to be merged to make a single character, so you use a separate character classifier pre-trained on correctly segmented characters to score the candidates, and then Viterbi/A* search on those scores to find the most likely interpretation of input.
Many problems with this -- how do you tune the heuristics? How do you recover from error in an earlier stage of pipeline? How do you get character level ground truth from image/text pairs?
With enough engineering time, you can solve those problems, but it's a lot of coding and tweaking. The point of the paper is that you can skip those steps and read OCR output directly off top layers of the network.