I don't think it's up to you, legally speaking: https://en.wikipedia.org/wiki/HiQ_Labs_v._LinkedIn
I mean, they could be nice and respect your robots.txt, but they certainly don't have to.
> fair use of snippets has relied on them being brief and linking to the source. Lawsuits will be immediate.
It's possible that fair use law will be expanded to cover this case, but as constructed the output of these models is generally fairly derivative of any specific original, and so probably protected under fair use. If it were spitting out exact copies of things it had read, it would probably be pretty easy to train that behavior out of it.
> I do love these imaginary scenarios where ChatGPT is going to find me the best air fryer, though. Where is that information going to come from, exactly? Barely anyone is making money writing reviews today, it's mostly farmed content. What happens when even those sites' reviews are quickly scraped and put into the next model iteration? Bing is going to have to come up with some kind of radical revenue sharing too if they want anything fresh.
I do agree with this, though. The LLMification of search is going to squeeze revenue for content creators of all kinds to literally nothing, at least if that content isn't paywalled. Which probably means that that's exactly where we're headed.