Plenty of software on the internet on fully open license (e.g. MIT, copyleft and so on) to train on.
Also, it would be relatively easy to build synthetic datasets for training.
I, for one, don’t mind models being trained on stuff I produced and shared publicly over the last 20 years. I did it for common good, including commercial uses, and this is one of them.
Plenty of people who never produced any open source trying to argue as if if they did.