What was the main motivation behind this? is GCP cheaper?
This is in contrast to the easiest way we found to deploy the same architecture on AWS using Elastic Beanstalk, which involved one really big instance (that was constantly growing as we added more models), and the costs that come with that.
The bigger issue was that he had to use bigger machine as we added more custom ML models for our customers. New architecture gives us huge $$ saving and more visibility into performance of each model.