Not to mention, if it's an ML workload, you'll also have to factor in downloading the weights and loading them into memory, which can double that time or more.
Imagine, you have a very small weak model, and you have to wait 20 seconds for your request.
For your first request, after having scaled to 0 while it wasn’t in use. For a lot of use cases, that sounds great.