As usual with ML, I now wonder how “similar” the test set is to the training set, compared to the examples that are neither in the training set, nor the test set:
TODO: 200 million - 170,000 training - 100 test ~= 199.8 million proteins
TODO: 200 million - 170,000 training - 100 test ~= 199.8 million proteins