You're right, but a) you have comp efficiency issues with MCMC, and b) just empirically MCMC models don't work as well as gradient descent + NN for many tasks.
We're also ignoring the benefits of a posterior distribution, which is useful for understanding the data-generating process.