If you want to query the Llama-2 models, you can use Anyscale Endpoints [1]. Note: I work on this :)
Llama-2-70B is $1 / million tokens, which is the most cost-efficient on the market that I'm aware of.
Llama-2-70B is $1 / million tokens, which is the most cost-efficient on the market that I'm aware of.
Edit. I'm sure it's answered on your site but sometimes it's better to include it right here! :)