My understanding is that ChatGPT was trained on text from the Internet and public domain texts. There is orders of magnitude more text available to humans behind paywalls and otherwise inaccessible (currently) to these models.
Am I missing something?