GPTCache: Slash Your LLM API Costs by 10x
github.com
github.com
As most ML is inherently probabilistic, it seems reasonable to make an LLM cache both semantic and _stochastic_, i.e. you wouldn't want the same answer every time you use "pick me a color" as prompt. Injecting the original LLM (GPT, Bard, etc) response as prompt for alpaca or some other model could make this cache virtually invisible.
If you're leveraging LLMs for your projects, it's definitely worth giving GPTCache a look!
Not to mention the chatgpt code interpreter plugin allowing sandboxed python execution and many beginners starting to code with llms, nearly everything will be in python eventually.