A little off-topic.
> Our fully preprocessed CodeSearchNet Corpus is available for download on Amazon S3
I am surprised that Github went with S3 for this download. Isn't there a Azure equivalent of S3 for large object storage ? This just shows the dominance of AWS.