It shouldn't be much of a problem to ship the LLM along with attributions. (List of all sources used in the dataset - not a problem, unless they are secret, shady or illegal.) Wikipedia is one of the easy ones, since you need to attribute just 'Wikipedia' for the entire corpus instead of many individual users.
The bigger issue is when a user uses the model to generate some text. Should they attribute it when using it somewhere?
That doesn't seem very practical, since it seems that soon most of the text will be edited by LLMs and those seem to be trained on most of the web -> so pretty much everything would need to be attributed to everything. Unless someone puts a stop to this, which I find improbable.