If it won't violate IP rights, there shouldn't be a problem.
It suggests those whose code is trained upon have something to lose if the trained models are used by others.
who would want the genius of Teams, sharepoint, onedrive or powerbi in their product?
On the other hand if the repo is already public on Github then exposing it via an LLM is not introducing any new security risk.
- that it doesn't output training data verbatim
- the product is very transformative, only "learning" from training data
- There are no copyright infringements because of these two above
Well, then there's really no reason not to throw their own private code on the pile.
Specifically, I think they are less concerned with (say) specific Excel code leaking than with the knock on effects of a cheap perfect substitute.
Is there any evidence that an LLM could actually generate a perfect substitute for excel solely through prompting if only the excel source was in the training data? I hypothesize that designing a prompt for an LLM that captures all of Excel's properties would be comparable in difficulty to reimplementing the functionality without an LLM.