Are they, or are they collections of probabilities? If they are probabilities, and those probabilities change from model to model, that seems like they might be copywritable.
If Google, OpenAI, Facebook, and Anthropic each train a model from scratch on an identical training corpus, they would wind up with four different models that had four differing sets of weights, because they digest and process the same input corpus differently.
That indicates to me that they are not a collection of facts.
At some point, with sufficiently many hyperparameters being chosen, that starts becoming a creative decision. If 5 parameters are available and all are left at the default, then no, that's not creative. If there are ten thousand, and all are individually tweaked to yield what the user wants, is that creative?
Not to mention all of these companies write their own algorithms to do the training which can introduce other small differences.
eg if I upload Marvels_Avengers.mkv.onnx and it reliably reproduces the original (after all, it's just a fact that the first byte of the original file is OxF0, etc)
Judge the output, not the system.
IIRC, this is wrong. Independent creation is a valid (but almost impossible to prove) defense in US copyright law.
This example is not an independent creation, but your reasoning seems wrong.