It doesn't have to be internally consistent, it just has to make them money.
It doesn't have to be internally consistent, it just has to make them money.
Not the case. That would be the case if OpenAI prevented them from using the same resources, which ChatGPT is not.
To keep that analogy going, they essentially used the ladder so hard, it broke.
It kind of is. Just think of how many services changed their data sharing policies and closed APIs due to ChatGPT training (Twitter, Stack Overflow, Reddit). Maybe the analogy is that instead of pulling the ladder, they set fire to it so it’s burning and making it harder for others to climb. Even if they didn’t set it alight on purpose, I don’t imagine they’re losing sleep over it.
Their ladder (using public data, and hiring humans to classify to taste) is still available I believe.
Not really. Once chatGPT came out, many sites changed their terms and/or significantly increased their API access costs to prevent/limit/make cost prohibitive future scraping.
I had assumed most of their web content was from Common Crawl, and the older pre-ChatGPT Common Crawl datasets used would still be available. But it looks like Twitter, for one, was not in Common Crawl.
Which is not openAI “pulling up the ladder behind them”
There is literally no way for them to avoid looking like assholes once they take enact barriers that they themselves did not have to overcome.
As far as we know.