Feds accuse China of 'systematic' distillation of U.S. AI models
cyberscoop.com
cyberscoop.com
It seems to me you can't have it both ways; either training is a violation of copyright or it isn't, and there's no consistent argument where "distillation" is a violation but other things aren't.
Bringing the prices of monthly subscriptions much closer to API rates would kill the arbitrage business and make distillation a much more expensive (and harder to hide) activity. Raise one, lower the other - whatever.
https://www.techopedia.com/trump-administration-openai-train...
Also, it takes a lot of compute. It still involves training a model, which for a very small model requires a GPU with several times more ram than the final model size, which will need to run for days to weeks. Usually you're also running the model that is being distilled from, but in this case it would require a huge database of chat logs, instead.