Training on copyrighted material does not violate copyright laws (despite what many will assume)
> Our only modification part is that, if the Software (or any derivative works thereof) is used for any of your commercial products or services that have more than 100 million monthly active users, or more than 20 million US dollars (or equivalent in other currencies) in monthly revenue, you shall prominently display "Kimi K2.5" on the user interface of such product or service.
[1] - https://huggingface.co/moonshotai/Kimi-K2.5/blob/main/LICENS...
I'm also deeply curious about this legal question.
As I see it, model weights are the result of a mechanistic and lossy translation between training data and the final output weights. There is some human creativity involved, but that creativity is found exclusively in the model's code and training data, which are independently covered by copyright. Training is like a very expensive compilation process, and we have long-established that compiled artifacts are not distinct acts of creation.
In the case of a proprietary model like Kimi, copyright might survive based on 'special sauce' training like reinforcement learning – although that competes against the argument that pretraining on copyrighted data is 'fair use' transformation. However, I can't see a good argument that a model trained on a fully public domain dataset (with a genuinely open-source architecture) could support a copyright claim.
It goes against the ML community ethos to obscure it, but is common branding practice.
[0] https://chainthink.cn/zh-CN/news/113784276696010804
[1] https://pbs.twimg.com/media/HD2Ky9jW4AAAe0Y?format=jpg&name=...
I bet Moonshot is going to make them open their wallets to avoid legal trouble.