The only thing I take issue with is the phrase "LiteLLM Without the Bloat." A lot of the features that have been removed (like cost tracking, streaming, caching) are... kind of the core value proposition of LiteLLM for many of their users.
The only thing I take issue with is the phrase "LiteLLM Without the Bloat." A lot of the features that have been removed (like cost tracking, streaming, caching) are... kind of the core value proposition of LiteLLM for many of their users.
The fix was deployed 2 weeks later (!!) to main (but you could downgrade of course).
Or that other time they broke model selection if you had selected "this key can used all models of their team" than the only model in the auto-selection for harnesses was an invalid "all-team-models" entry. Fixes this one in 1 week though.
all of them on the :latest docker tag btw
Imagine Sqlite adding heavy features from Postgresql, e.g. row-level security.
litellm.register("foo", CustomAIProvider)
litellm.do_whatever("foo/my-cool-model", "what is 2+2")https://www.getmaxim.ai/bifrost/resources/benchmarks
Having ran both LiteLLM and Bifrost for months, I can largely confirm the numbers from those benchmarks for myself.
That's not to say Bifrost wouldn't have been better, but the choice to use LiteLLM was arrived at after a fair bit of internal discussion (most of which predated my addition to the team), and so far we've seen nothing from LiteLLM that has been contradictory to the pros/cons they thought would be the case when LiteLLM was adopted.
Or in other words, the org will be happy indeed when they have solved so many of the rest of the problems we've had in AI uptake that the difference in latency between one AI gateway or the other becomes a problem to be solved.
THIS! And it's way cheaper than others like Kong =)