https://en.wikipedia.org/wiki/LibriVox
The recordings are public domain audio books of public domain books, so the licensing should be fine. The audio isn't annotated, but given the value involved I think it would be worth attempting to use forced alignment to annotate the recordings with their public domain source texts. Forced alignment using the sort of speech recognizer you're trying to train in the first place may be a bit "chicken and the egg", but from some experiments I've run myself existing open source speech recognizers can do it reasonably well. Humans could manually tune up the alignment to improve the quality if necessary.
As for motivating people to actually do that mundane work... well these are audio books so maybe the work isn't so mundane after all! The LibriVox recording of Tom Sawyer (read by John Greenman: https://librivox.org/tom-sawyer-by-mark-twain/) is pretty great and has been listened to by millions of people. If somebody created a "read along" web app that showed you the text of the book from Project Gutenberg getting highlighted as the audiobook from LibriVox was played, users who have an interest in reading/hearing the book could have their attention held by Mark Twain and with the right UI provide fine tuning for the forced alignment at the same time.