I point you to this article: https://medium.com/@klintcho/creating-an-open-speech-recogni...
It basically describes the thing you mentioned - matching freely available audio books with the source text and using some tools to preprocess the data suitable for ASR training (alignment, splitting).