HNHacker News
TopNewBestAskShowJobs

kaushik92

30 karma · joined June 29, 2022

YC Badge: 0xbe04d0b21b34584b5a1f732aa7f209a533c0b915
submissionscomments
kaushik92··on Show HN: FiddleCube – Generate Q&A to test your LLM
I have seen this, but I find it fairly hard to use.

Our goal is to focus on datasets and make it very easy to create and manage data.

In our next release, we will be launching a way to do this using a UI.

kaushik92··on Show HN: FiddleCube – Generate Q&A to test your LLM
We identified and solved for 2 key problems with generating data using GPT: 1. Duplicate/similar data points - we solve this by adding deduplication to our pipeline. 2. Incorrect question-answers - we check for correctness and context relevance. Filter out incorrect rows of data.

Apart from this, we generate a diverse set of questions including complex reasoning and chain of thought.

We also generate domain specific unsafe questions - questions that violate TnC of the particular LLM to test the model guardrails.

kaushik92··on Show HN: FiddleCube – Generate Q&A to test your LLM
This is great feedback! Will work on adding this soon.
kaushik92··on Show HN: FiddleCube – Generate Q&A to test your LLM
We do care a lot about data privacy. While the data is sent to an API, we do not store anything on our servers or use our user's data in any way.

We are working on getting SOC2 certified. In the meantime, we sign a legally binding agreement with our users who have data privacy needs/concerns.

kaushik92··on Show HN: FiddleCube – Generate Q&A to test your LLM
Ragas is an eval tool which needs ground truths and queries for evaluation. FiddleCube generates the queries and the ground truth needed for eval in Ragas, LangSmith or an eval tool of choice.

We incorporate user prompts to generate the outputs and provide diagnostics and feedback for improvement, rather than eval metrics. So you can plug your low scored queries provided by Ragas, your prompt and context. FiddleCube can provide the root cause and the ideal response.

This is an alternative to manual auditing and testing, where an auditor works on curating the ideal dataset.

kaushik92··on Show HN: A lightweight AI gateway to 100+ models, in TS
Congrats on the launch! Looks amazing
kaushik92··on [dead]
Migration projects are painful and no developer enjoys them. Good to see a great solution to automate this!
kaushik92··on Show HN: Litellm – Simple library to standardize OpenAI, Cohere, Azure LLM I/O
This is amazing. Really needed something like this to standardize all my different AI APIs!

On a side note - I love how quickly your team is shipping! Do keep it going!

kaushik92··on Launch HN: PeerDB (YC S23) – Fast, Native ETL/ELT for Postgres
Moving data in and out of Postgres in a fast and reliable way is exactly what my startup needs. I am looking forward to trying PeerDB!