HNHacker News
TopNewBestAskShowJobs

ij23

118 karma · joined June 8, 2022

building https://github.com/BerriAI/litellm
submissionscomments
ij23··on LiteLLM Migrates to Rust
LiteLLM maintainer here. Some context on why we are doing this

Over the past year we've heard the same thing from our users and community, they want the fastest and litest AI gateway.

This change allows us to address two of the most common problems we hear from users latency spikes under load and memory leaks/OOM kills that take pods down

We believe a Rust hot path is faster and bounded in memory, so those whole classes of issues go away.

It will be a gradual, non-breaking change. The Python SDK and proxy stay exactly the same, under the hood they start calling the Rust binary through PyO3, one component at a time, each proven in production before the next. The sub-1ms figure is gateway overhead (what we add on top of the upstream call), and we're aiming for a sub-100MB binary. Happy to share benchmark methodology if folks want to poke at it.

The whole gateway will be running on Rust by December 1, 2026.

Full announcement: https://docs.litellm.ai/blog/litellm-rust-launch

ij23··on Tell HN: Litellm 1.82.7 and 1.82.8 on PyPI are compromised
Hi all, Ishaan from LiteLLM here (LiteLLM maintainer)

The compromised PyPI packages were litellm==1.82.7 and litellm==1.82.8. Those packages have now been removed from PyPI. We have confirmed that the compromise originated from the Trivy dependency used in our CI/CD security scanning workflow. All maintainer accounts have been rotated. The new maintainer accounts are @krrish-berri-2 and @ishaan-berri. Customers running the official LiteLLM Proxy Docker image were not impacted. That deployment path pins dependencies in requirements.txt and does not rely on the compromised PyPI packages. We are pausing new LiteLLM releases until we complete a broader supply-chain review and confirm the release path is safe.

From a customer exposure standpoint, the key distinction is deployment path. Customers running the standard LiteLLM Proxy Docker deployment path were not impacted by the compromised PyPI packages.

The primary risk is to any environment that installed the LiteLLM Python package directly from PyPI during the affected window, particularly versions 1.82.7 or 1.82.8. Any customer with an internal workflow that performs a direct or unpinned pip install litellm should review that path immediately.

We are actively investigating full scope and blast radius. Our immediate next steps include:

reviewing all BerriAI repositories for impact, scanning CircleCI builds to understand blast radius and mitigate it, hardening release and publishing controls, including maintainership and credential governance, and strengthening our incident communication process for enterprise customers.

We have also engaged Google’s Mandiant security team and are actively working with them on the investigation and remediation.

ij23··on Open-Swarm – use 100 LLMs on OpenAI swarm framework
OpenAI's multi-agent framework swarm only supports models from OpenAI.

OpenSwarm uses LiteLLM to add support for any LLM AnthropicAI, MistralAI, Ollama, Huggingface, GroqInc, Replicate

ij23··on Show HN: Self-Hostable Algolia DocSearch Replacement
Canary is awesome! we use Canary for our doc search at LiteLLM (you can see it here: https://docs.litellm.ai/docs/)

It's really useful to be able to specify the search space for a specific query (example: Canary allows search for the query "sagemaker" on our docs or on our github issues )

ij23··on Show HN: I built an OSS alternative to Azure OpenAI services
hi i'm the maintainer of litellm - we persist rate limits, they're written to a DB: https://docs.litellm.ai/docs/proxy/virtual_keys

- LiteLLM Proxy IS Exactly Compatible with the OpenAI SDK

ij23··on Are Open-Source Large Language Models Catching Up?
I'm the LiteLLM maintainer, can you elaborate what you're looking for us to do here?
ij23··on Llama2 on Replicate faster than ChatGPT?
Ran some testing and discovered llama2 on replicate is faster than chatgpt!

Code - https://github.com/BerriAI/litellm/blob/main/cookbook/Evalua...

Are others seeing similar results?

ij23··on Show HN: liteLLM Proxy Server: 50+ LLM Models, Error Handling, Caching
What local/in-K8-cluster models servers would you recommend adding ?

Should we add support for llama.cpp and vllm.ai in the proxy server ? Or should we assume you can host them on your own infra and the proxy server requests your hosted model ?

ij23··on Show HN: liteLLM Proxy Server: 50+ LLM Models, Error Handling, Caching
Yes, you use your own API keys. You can set them as env variables. Either set them as os.environ['OPENAI_API_KEY'] or set them in .env files: https://litellm.readthedocs.io/en/latest/supported/
ij23··on Show HN: Litellm – Simple library to standardize OpenAI, Cohere, Azure LLM I/O
thanks for sharing, while your library looks really powerful my goal with Litellm is simplicity
ij23··on Show HN: Litellm – Simple library to standardize OpenAI, Cohere, Azure LLM I/O
good points, probably going to add streaming output, function calling support. As for retries tenacity does a great job already
ij23··on Show HN: Litellm – Simple library to standardize OpenAI, Cohere, Azure LLM I/O
azure models have custom names - eg I call mine 'chat-gpt-test1', I require some flag to know if it's an azure model
ij23··on Show HN: Litellm – Simple library to standardize OpenAI, Cohere, Azure LLM I/O
Thank you !
ij23··on Show HN: Litellm – Simple library to standardize OpenAI, Cohere, Azure LLM I/O
Thanks!