All the methods and data are not public. We don't know what unpublished methods they're using. You can get most of the pre-training data publicly but they've probably spent a ton of money curating it and are now doing things like buying rare books. The RL training data is all (/mostly) proprietary though, and that's the real secret sauce part.