Does anyone know how to estimate the cost of inference using your own Llama2 model? This article talks about the cost of fine tuning it, but not what to expect when running it in production for inference.
In particular, it would be great to know how the inference cost compares to gpt3.5 turbo and gpt4