As someone who used to work in adjacent field, I have a question that I hope you can help me with. DL usually requires a large volume of data, but I don't know whether there is a huge pile of experimentally determined protein structures for training. Did DeepMind find a clever method to get around data-size issue or there is really a lot of known protein structures? I skimmed one of their earlier papers and it seems training data was in tens of thousands. I am surprised that is enough for training DL structure prediction.