Given the architecture demands and shortcomings noted in the article, it would likely be more efficient to cut out the overhead invoked by Cloud Run and AI Platform and just run everything in Kubernetes Engine (backed by Knative for Cloud Run-esque autoscaling). This would also solve the latency and scale-to-minimum size issues, and likely be cheaper in the long term at the cost of a bit more configuration to get it started.