What's the cheapest way to run e.g. LLaMa2-13B and have it served as an API?
I've tried Inference Endpoints and Replicate, but both would cost more than just using the OpenAI offering.
I've tried Inference Endpoints and Replicate, but both would cost more than just using the OpenAI offering.