It's not open source, we have no idea about the data used to train the model, and the paper doesn't explain it all.
I would think the code to build your own is open sourced, and you can feed it any data you'd like. That's the open source part, not the part where they are running the model.
Have I misunderstood this?
I think it’s kind of an overdone complaint and I usually ignore it, and besides it looks like there’s a huggingface project ongoing where they’re trying to replicate the training process for this model anyway.