All the major LLM vendors are vague on datasets because it's quite certainly copyrighted text, (Remember when Google was scanning all the world's books a decade ago?) and they don't want to give the inevitable lawsuits a jump start.
> The news groups' concerns arose when the computational journalist Francesco Marconi posted a tweet last week saying their work was being used to train ChatGPT. Marconi said he asked the chatbot for a list of news sources it was trained on and received a response naming 20 outlets including the WSJ, New York Times, Bloomberg, Associated Press, Reuters, CNN and TechCrunch.
https://www.bloomberg.com/news/articles/2023-02-17/openai-is...
Yes, and I remember that they won that lawsuit.