Calculate a rough semantic similarity score across all your snippets, and pay out a fractional reward to all originating codebases.
I think the bigger problem is that it will almost certainly lead to a proliferation of giant snippet spam repositories.
It is like the search engines vs. SEO arms race. The hope could be that such proliferation can be managed by disincentivizing such abuses. The reality might be vastly different with codes that have more regularity and better chance for AI's emulating humans than the natural language texts.
Then they need to negotiate a non-attributed contract with you before using your code to train (not sure abiut testing though).