11 karma · joined June 18, 2020
Curious what that means? Let's chat: <hn username> at gmail
Work with https://portkey.ai/ (Observability, evaluation, and AI gateway for foundation models)
On Portkey, since we decouple prompt templates from your code - you can continue iterating on the prompts on Portkey and just reference the prompt ID in code. Any change you make to Portkey prompts automatically get reflected in the app because the prompt ID keeps pointing to the latest / published version.
does that make sense?
Sidenote: Arch is def interesting!
A typical user we've seen at Portkey is a mid or a large size eng org where a central "Gen AI team" has now come up. This Gen AI team builds services that the rest of the company uses to build whatever AI features or products they want.
To build such a service, they need traditional API Gateway features like rate limiting, access rules, and also AI-specific features like universal API to multiple LLM providers, universal routing, central guardrails, AI-native observability + central dashboard for other stakeholders, and more.
It can absolutely be a plugin on top of existing Gateways.. like we've explored putting Portkey on Kong, but the need for a dedicated AI Gateway still remains, that can do all of these things I described in an easier way.
Probably, solutions like Langchain/Llamaindex etc. also fit in somewhere here, but a dedicated service for "ops" related issues for LLM APIs is something that we're seeing orgs adopt as a good practice.
The caching part isn't open source yet, but part of our internal workers. Would be very cool to open source it!
Quick notes on the analysis: - This is based on data from 100+ organizations globally, doing million+ requests a day via Portkey.ai. - I've randomly sampled 10,000 requests for both GPT 3.5 & 4 each day, for the past 3 months. - Variants like -0613 & -0314 are grouped under GPT3.5 & GPT4 for clarity. - I've plotted 'Latency PER Token' rather than just latency - For each day, I've counted percentile values for Latency/Token across different percentiles like 50, 75, 90, 99 to account for variance in prompt length and complexities and separate out anomalies
Happy to share more or answer questions, if any.
Would still be cool to learn technical details.
I'd say they are executing well.
Some more info here: https://docs.portkey.ai/overview/introduction
Makes sense on the deployment thing. I like what Vellum is doing. Helicone and Portkey also let you do deployments of prompt templates through APIs.