This is really cool. If you're looking for more datasets to train your model, here are a few relevant ones:
- https://archive.org/details/stackexchange
- http://trec.nist.gov/data/qamain.html
- http://opus.lingfil.uu.se/OpenSubtitles2016.php
- http://corpus.byu.edu/full-text/wikipedia.asp OR https://en.wikipedia.org/wiki/Wikipedia:Database_download#En...
- http://opus.lingfil.uu.se/
I'd love to see how good your model gets.