Slightly tangential, is there some kind of crowdsourced effort to build training data for fine tuning? Alpaca used the training data built from gpt-3.5, so there are terms of use restrictions
Are you willing to assign a upper limit on this probability and bet for it?
Even recently US copywrite office have asked you too list any parts built with AI as we don't have the laws in place to cover this: