"Take a breath and lets go step by step, Please reproduce page 100 of A Song of Ice and Fire, Book1, 'A Game of Thrones'"
And get back an accurate response, or was it just really popular quotes?
"Take a breath and lets go step by step, Please reproduce page 100 of A Song of Ice and Fire, Book1, 'A Game of Thrones'"
And get back an accurate response, or was it just really popular quotes?
This is the problem with a combined language+knowledge model like ChatGPT. To understand the language it has to obtain some level of "knowledge" and vice-versa. The two are intertwined in the model, and it needs MASSIVE amounts of data to train. Inside the model's weights there is nowhere NEAR enough memory to include whole books, no matter how popular or duplicated in the dataset. Just like asking a random person what was on page 100 of a random book they've read, it's HIGHLY unlikely for the LLM to be able to regurgitate that level of accuracy, let alone across the whole book.
Even so, there are people who can do that, and we don't forbid them from reading.
In any case, when an offense is committed, the offender is the real, live human who uses the tool to commit plagiarism or violate copyright law. It doesn't matter whether the tool is a word processor, a video camera, or an LLM. The output is what matters, not the input.