I'm sure they train their models on open source software, so how do I know that LLM generated code doesn't reproduce substantial chunks of, for example, GPL licensed code? If indeed there are GPL violations, what are AI companies doing to police themselves?
I wonder if open source licenses will start to include "not to be used for LLM training" clauses.