Have you come across any specific details on their input datasets? I’ve been assuming CommonCrawl and Pile but if you have anything specific it would be very interesting what else was used.
> The news groups' concerns arose when the computational journalist Francesco Marconi posted a tweet last week saying their work was being used to train ChatGPT. Marconi said he asked the chatbot for a list of news sources it was trained on and received a response naming 20 outlets including the WSJ, New York Times, Bloomberg, Associated Press, Reuters, CNN and TechCrunch.
https://www.bloomberg.com/news/articles/2023-02-17/openai-is...
Yes, and I remember that they won that lawsuit.