The thing which could be a great offer is ability to supply data. I’m not sure if it involves up training or something else. Imagine I have a bunch of documents and I want answer to be based on my data.
Especially for the set of problems with "Step 1: Hosting embeddings for some large corpus", it feels like there's a really useful role for offering static query/AKNN search atop popular datasets.