Hi! Thank you for the comment, I'm the writer on the article. When we did our training we actually needed the computer for months at a time. It takes 1-2 months to tune a model and we were running exps almost 24/7. So one project put us in the break even point to build.
Was it possible to paralelize the process? Seems like training on few dozen cloud GPUs and finishing much faster would save a lot of money in terms of engineer time