What we focus on is the runtime layer underneath. You can run us behind Cloud Run or inside your existing GCP setup. The difference is at the GPU utilization level when you’re serving many models with uneven demand.
If your workload is steady and high volume on a small set of models, the standard cloud stack works well. If you’re juggling dozens of models with spiky traffic, the economics start to look very different.
As an example, we’re currently being tested inside GCP environments. Some teams are experimenting with running us behind their existing Google Cloud setup rather than replacing it. The idea isn’t to swap out Cloud Run or Vertex, but to improve the runtime efficiency underneath when serving multiple models with uneven demand.