1) You can't copyright weights. A lot of people believe this. I am not sure this is true. I think it might be that there is a fair use argument that the weights are transformative, but having fair use on your infringement doesn't imply a lack of the original copyright being owned by someone. But like, this might be true, and it is not an unreasonable stance.
2) The model weights are a derived work of the training data. This feels right to me, frankly, as much as it irks a lot of people on Hacker News who are excited to use Copilot (or owns shares of Microsoft ;P). In this case, Facebook does not own the copyright as they purposefully used an open training set from third parties (including Wikipedia and OpenCrawl) as a counter to OpenAI's proprietary one.
3) The model weights are a derived work of the training code, in the same way a binary is a derived work of its source code (vs. #2 where the code is a compiler and the source code is the training set). In this case, Facebook would own the copyright... but the resulting binary program--as distributed in the machine interpretable (and executable) format of the model weights--must be GPL as Facebook amazingly used that as the license for their training code.
The interpretation of events that would allow Facebook to have some hope of arguing that the model weights are their own work--that maybe the code is more of a tool like Photoshop and they have a fair use claim on the training data--would imply something that simply is not true: that they are adding some (hopefully extensive) form of expressive input above and beyond those two inputs, in the way someone using Photoshop does when they remix someone else's art into a transformative work.
However, Facebook definitely isn't doing that: they have provided absolutely no expressive input or intent on top of those two inputs, one of which they do not own and the other of which they chose to license for us under GPL! To make this a bit clearer, separate Facebook into two parties for a moment--one which developed the tooling and the other of which ran it--to determine which of the various parties you think owns this result: if you download a program someone else wrote and click a button to run it on some data someone else owns, you simply do not own the copyright on the result. (edit: I wrote some more on this argument in the following linked comment.)
https://news.ycombinator.com/item?id=35293068
The only thing I can come up with, if I try really really hard to steelman Facebook here, is: maybe, if I were to release a binary to the world that is a compiled copy of code that I also simultaneously released under GPL, the binary might technically have been compiled from an internal pre-licensed (but identical) copy of said code; and so, while if you compile the same binary from the licensed code you get an identical output that is a derived work with GPL rights, when I do that I don't... but this feels like a perilous argument to make as you are going to have such a hard time showing that this was a reasonable way to infer your intent with the simultaneous release.
(Of course, the person who originally agreed to the terms of service attached to the download they got is a totally different matter, but that doesn't mean there are none. There are a lot of limitations to what people can extract out of you if you violate a contract. So like, by having explicitly agreed to those terms with Facebook that person should not have given the world the weights--not because they are copyrighted but merely because they were secret--and there might be some kind of ramification... but, I do not believe that would possibly apply downstream to sillysaurusx.)
(Note: I am not a lawyer. I spend a ridiculous amount of my time vs. a normal engineer working with copyright and both hearing and making arguments about copyright, including directly to the Copyright Office at the Library of Congress as part of my work on Cydia and my efforts to push back on parts of the DMCA alongside lawyers from the EFF... but like, it would be foolish to read my comment here and then embark on a project to do something that might massively infringe on someone's copyrights without running it past a real lawyer. I, certainly, have actual lawyers.)