E.g.: if the secret is that it has been trained on tons of "books", then Google could just throw Google Books at Bard 3.
E.g.: if the secret is that it has been trained on tons of "books", then Google could just throw Google Books at Bard 3.
(I put together a dataset of 190,000 books that I called books3, which llama eventually trained on. Usage rights are a big interest of mine, primarily because there’s a weird disconnect of people trying to claim copyright over models when the underlying data obviously wasn’t copyrightable.)
IANAL, and I’m not saying how this is how it’ll play out in the courts, but I don’t see why this is a “weird disconnect”
It'd be easy to imagine models that accidentally made this possible.
* Sometimes a book has several very slightly different versions. A lot of Fantasy. Sometimes many parts of a series are included but not all of it (I can imagine an argument for a sample, or for the series in its entirity). Commentaries about important philosophical arguments like Rawls' 'A theory of justice', but not the work itself.
Basically, the-eye.eu was at one point hosting all of bibliotik, so I downloaded all the epubs and converted them to text. I still have those epubs (incidentally thanks to Carmack, who through a convoluted process managed to save the them and send them to me via snail mail) and I’ve been considering releasing them so that you can filter the books yourself.