HNHacker News
TopNewBestAskShowJobs

retrovrv

11 karma · joined June 18, 2020

Student of artful reality.

Curious what that means? Let's chat: <hn username> at gmail

Work with https://portkey.ai/ (Observability, evaluation, and AI gateway for foundation models)

submissionscomments
retrovrv··on API, Claude.ai, and Console services impacted [resolved]
Your best bet is having an account on AWS Bedrock & Vertex AI so you're able to route your request to the same model (such as claude-sonnet-4) but on a different provider.
retrovrv··on Show HN: Any-LLM – Lightweight router to access any LLM Provider
there's portkey that we've been working on: https://github.com/portkey-AI/gateway
retrovrv··on Show HN: Any-LLM – Lightweight router to access any LLM Provider
we essentially built the gateway as a service rather than an SDK: https://github.com/portkey-AI/gateway
retrovrv··on Is outbound going to die?
just wanted to say, fantastic note. thanks for sharing!
retrovrv··on Show HN: Prompt Engineering Studio – Toolkit for deploying AI prompts at scale
thank you! we don't have a strong Evals module within our Prompt Studio at the moment. So there's no straightforward way to do that. however, we do have one of the better modules for applying guardrails on live AI traffic and setting up routing based on guardrail verdicts: https://portkey.ai/features/guardrails
retrovrv··on Show HN: Prompt Engineering Studio – Toolkit for deploying AI prompts at scale
lol. i get you - i think the 3rd point you shared - it's not exactly about CI issue - CI is already as fast as it can be. It's just that a prompt is a critical part of your AI app and any change to it needs to go through a few hoops before it's available to all users.

On Portkey, since we decouple prompt templates from your code - you can continue iterating on the prompts on Portkey and just reference the prompt ID in code. Any change you make to Portkey prompts automatically get reflected in the app because the prompt ID keeps pointing to the latest / published version.

does that make sense?

retrovrv··on Garak, LLM Vulnerability Scanner
This is actually pretty cool. Would love to try it!
retrovrv··on Show HN: Arch – an intelligent prompt gateway built on Envoy
I'm affiliated with Portkey, so can answer who would need such a proxy/gateway:

Sidenote: Arch is def interesting!

A typical user we've seen at Portkey is a mid or a large size eng org where a central "Gen AI team" has now come up. This Gen AI team builds services that the rest of the company uses to build whatever AI features or products they want.

To build such a service, they need traditional API Gateway features like rate limiting, access rules, and also AI-specific features like universal API to multiple LLM providers, universal routing, central guardrails, AI-native observability + central dashboard for other stakeholders, and more.

It can absolutely be a plugin on top of existing Gateways.. like we've explored putting Portkey on Kong, but the need for a dedicated AI Gateway still remains, that can do all of these things I described in an easier way.

Probably, solutions like Langchain/Llamaindex etc. also fit in somewhere here, but a dedicated service for "ops" related issues for LLM APIs is something that we're seeing orgs adopt as a good practice.

retrovrv··on Ask HN: Devs using LLMs, how are you keeping costs low for LLM calls locally?
came across this guide earlier - valuable insights. thanks for sharing!
retrovrv··on Ask HN: Devs using LLMs, how are you keeping costs low for LLM calls locally?
there's an open source ai gateway - https://github.com/Portkey-AI/gateway
retrovrv··on Show HN: Ragas – Open-source library for evaluating RAG pipelines
Phenomenal to see how Ragas has progressed. Congratulations on the launch
retrovrv··on Open-source tools for LLM monitoring and observability
https://portkey.ai/ - is also a nice addition to the list
retrovrv··on Show HN: A lightweight AI gateway to 100+ models, in TS
Thank you! We have built out the cache system -- we do both simple caching (matching the request strings 100%) and also do semantic caching (returning a cache hit for semantically similar requests). More here - https://portkey.ai/docs/product/ai-gateway-streamline-llm-in...

The caching part isn't open source yet, but part of our internal workers. Would be very cool to open source it!

retrovrv··on Show HN: A lightweight AI gateway to 100+ models, in TS
Pretty excited to announce this! While there are some popular and awesome AI gateways out there, like litellm, bricksai - none are written in TS, and for the TS ecosystem. Looking forward to the community's feedback
retrovrv··on OpenAI DevDay's Implications for LLM Apps in Prod
I really liked request_format param to enforce JSON outputs, and seed param to enforce deterministic outputs - both I've already started to use.
retrovrv··on GPT-4 is Getting Faster
Thanks for sharing my blog here!

Quick notes on the analysis: - This is based on data from 100+ organizations globally, doing million+ requests a day via Portkey.ai. - I've randomly sampled 10,000 requests for both GPT 3.5 & 4 each day, for the past 3 months. - Variants like -0613 & -0314 are grouped under GPT3.5 & GPT4 for clarity. - I've plotted 'Latency PER Token' rather than just latency - For each day, I've counted percentile values for Latency/Token across different percentiles like 50, 75, 90, 99 to account for variance in prompt length and complexities and separate out anomalies

Happy to share more or answer questions, if any.

retrovrv··on Show HN: Magentic – Use LLMs as simple Python functions
Super cool! Looks quite intuitive, especially for function calls.
retrovrv··on How are generative AI companies monitoring their systems in production?
There are quite a few LLM monitoring tools in the market. But for monitoring (or evaluating) RAG systems, I found Ragas to be the most helpful: https://blog.langchain.dev/evaluating-rag-pipelines-with-rag...
retrovrv··on Show HN: AI Grants and Credits Finder
There are some programs that match founders, bring in VCs and overall facilitate "Networking" - those are tagged as such. Hope that helps!
retrovrv··on Ask HN: Where to Host Llama 2?
Looks very interesting
retrovrv··on Ask HN: Is there a market for premium Figma plugins?
These guys built a few Figma plugins, made money from that, and also got acquired by Figma: https://diagram.com/
retrovrv··on Ask HN: What ChatGPT plugins do you use daily, if any?
Code interpreter, though not exactly a plugin, has become my default mode of interacting with ChatGPT.
retrovrv··on Show HN: Pykoi – a Python library for LLM data collection and fine tuning
A lot of thought has been put into this. Quite like the dashboard design, and the simplicity of using this. Congrats on the launch!
retrovrv··on Show HN: Q&A with AI Trained on Bankruptcy Law
My first question as well, especially considering how fast it is. But understand if you can't reveal.

Would still be cool to learn technical details.

retrovrv··on Ask HN: Where do people sell apps these day?
To add to what others have pointed out, if you have a good following on any social platform and can push your product there, putting it behind something like Gumroad is also an option.
retrovrv··on Tell HN: 6 months after the NYT acquisition, Wordle has started showing ads
I am actually quite intrigued with NYTimes's Games play. I tend to play mostly their games whenever I'm on mobile. And surprisingly, their India games subscription pricing is also not very prohibitive.

I'd say they are executing well.

retrovrv··on Ask HN: How are you improving your use of LLMs in production?
Not yet, but it's on the roadmap to publish. Broadly, we use clickhouse, streams, mongo, cloudflare workers as the mainstay of our monitoring infra.

Some more info here: https://docs.portkey.ai/overview/introduction

retrovrv··on Ask HN: How are you managing LLM APIs in production?
Langsmith would be useful in that case to see the traces for your chain. (Of course given that you are using Langchain for chaining the prompts)

Makes sense on the deployment thing. I like what Vellum is doing. Helicone and Portkey also let you do deployments of prompt templates through APIs.

retrovrv··on Ask HN: How are you improving your use of LLMs in production?
helping build https://portkey.ai/ where we try to take care of storing all prompts, responses, erros, costs, latencies etc and make it available to debug.
retrovrv··on Ask HN: How are you managing LLM APIs in production?
Langsmith is broadly for tracing the chains - are you looking for prompt deployment solutions?
Page 1 of 2Next →