However, you fail to recognize that OpenAI also created Whisper, which is a quote capable speech-to-text transcriber; and this tool easily converts audio and video (those verbal bits you mentioned) into text.
So the pool of creativity which OpenAI can train its models on is far larger than just original text; further, they demoed image modality a couple weeks back which would allow for VISUAL creative works to be parsed as well.