I’ve noticed an increase in services like this lately. What gives? Is there some sort of ML serverless offering made available on GCP?
Also, when it comes to summarization- you don’t really need to infer each run, you can throw up a pretty simple caching system. Which means repeat requests are far cheaper and faster.
I used cloudflare workers as a proxy / caching layer with KV in front of an AWS lambda to do article extraction and SageMaker spinup (with a small cache on the AWS side too- to catch in progress jobs)