Great summary! Reminds me a lot about Leon Bottou's work on using deep learning to learn causal invariant representations. (Video: https://www.youtube.com/watch?v=lbZNQt0Q5HA)
We can view the augmentations of the image as "interventions" forcing the model to learn an invariant representation of the image.
Although the blog post did not frame it as this type of problem (not sure if the paper did), I think it can definitely be seen as such and is really promising.