How do you avoid accidently encoding knowledge into your algorithm? You can't test your algorithm on a different sample of naturally occurring mathematics , summer wet have only one.
Information leaks are indeed the number one challenge when evaluating this type of approach. Our evaluation procedure does not rigorously guarantee against information leaks, but it does allow for an apples-to-apples comparison with previous methods (which had the same potential issue). And showing an advantage over these methods is what we were trying to achieve.
So in summary: it's hard to prevent, but we're doing our best.