Curious; is it better to train locally on something like a 2080ti 11G or go for colab and offload checkpoints to S3?
Asking because it seems V100 performance (or the other colab paid GPU) is worth the occasional instability if you’ve set up checkpoints.