Still relatively simple, the stack being LLM hopefully most of the actual “stack” work will be transferred inside the LLM. Example, if context size becomes unlimited, could do away with vector dbs.
> Claude offers fast inference, GPT-3.5-level accuracy, more customization options for large customers, and up to a 100k context window (though we’ve found accuracy degrades with the length of input).