I'm no lawyer, but this repository sure appears to be relicensing the Harry Potter series under the GPL.
If all the training data is in the txt files, it is obviously trained on copyrigthed material, and immensly low amounts of text. Im impressed if the outputs even start to make sense at all.
True. The best solution for small models is to use Project Gutenberg to reduce risk: