Show HN: Open-source LLM provider price comparison
github.com
github.com
https://github.com/BerriAI/litellm/blob/main/model_prices_an...
Simple UI to search:
[1]: https://openrouter.ai/models/meta-llama/llama-3.1-70b-instru...
I think Together recently introduced a different price tier based on precision but otherwise it is usually dark.
Email me at ed at shadeform dot ai if we can help
Then I realized, that I changed the provider. And the new one quantized Llama 3.1 with fp8.
Then I tried Hyperbolic [2], because they offer the model in different quantizations. As result, Llama 3.1 was better than Llama 3 or at least on par.
I like your charting, many have taken this task and then lose interest.
similar other tools for inspiration https://llmprices.dev/ https://www.llmpricing.app/
What no one is doing is focusing on GPUs, what is the cost of running L3-8B on an A100 or H100 per second.
It's a lot of work, your target users is companies that use Runpod and AWS/GCP/Azure, not Fireworks and Together, they are in the game of selling tokens, you are selling the cost of running seconds on GPUs.
P.S: I am from Inferless.
What sets them apart is that they have speed and latency as well.
I wish that OpenRouter would also expose the amount of output tokens via API as this is also an important criteria.
Without any quantization our current price is 30cts ingest and 50cts output per million tokens. [1]
how about adding more models and providers?
and making a sortable table?
Agree, we will add a MUI table very soon. Also some charts.
I genuinely want someone to roast the way I did my benchmark process described there. Want something good enough yet easy to run.