We wouldn’t bat an eye at using S3, EC2, and RDS as a host your own setup. The only difference here is that startups are moving faster than incumbents.
FWIW that’s one reason why Steamship (disclaimer: I’m the founder) aggregates all AI services under a single API key and interface. It’s to deal with the insane glue-code hassle of running this stuff on your own.
If you're looking to self-host chat memory rather than go all in on Supabase, there's Zep: https://github.com/getzep/zep
Full disclosure: I'm a co-author.
That said, I suspect summarization, translation, etc will take time. I’d suspect under 1 year.
Like which one?
And there are some very new 65b finetunes I have not tried.
Huggingface is full of finetunes now, and I believe a 33b model can be finetuned on a single 3090.
Llama.cpp is developing some kind of training, but I have no idea what the requirements will be.
Maybe a year before we are the level of stablediffusion?