(I put together a dataset of 190,000 books that I called books3, which llama eventually trained on. Usage rights are a big interest of mine, primarily because there’s a weird disconnect of people trying to claim copyright over models when the underlying data obviously wasn’t copyrightable.)