In my humble opinion, we never had failures of GPU even for large scale training. Our current training batch job is a 20GB json file which takes 6 hours just to load and has been running for more than 15 days with not a hiccup. And we are using the older Tesla T4.
GPUs have memory constraint issues but if you can plan and work around it, I havent seen it crash in real life.