HNHacker News
TopNewBestAskShowJobs

pchunduri6

52 karma · joined April 30, 2023

Opinions my own
submissionscomments
pchunduri6··on Launch HN: Talc AI (YC S23) – Test Sets for AI
I just tried the demo, and it looks great! Congrats on the launch!

I have a couple of questions:

1) How often do you find that the LLM fails to generate the correct question-answer pairs? The biggest challenge I'm facing with LLM-based evaluation is the variability in LLM performance. I've found that the same prompt results in different LLM responses over multiple runs. Do you have any insights on this issue and how to address it?

2) Sometimes, the domain expert generating the test set might not be well-equipped to grade the answers. Consider a customer-facing chatbot application. The RAG app might be focused on very specific user information that might be hard to verify or attest by the test set creator. Do you think there are ways to make this grading process easier?

pchunduri6··on Good LLM Validation Is Just Good Validation
This was exactly my experience trying to build a RAG pipeline that answers complex questions by generating sub-questions [1]. The hidden prompts in the complex libraries are quite difficult to get to and then start tweaking, whereas writing my own pipeline boiled down the code to 4 LLM calls. +1 for using the instructor library for quick schema definitions!

[1] https://github.com/pchunduri6/rag-demystified

pchunduri6··on Show HN: Demystifying Advanced RAG Pipelines
1) While building this system, I found that the LLM can sometimes generate unpredictable responses. For example, the LLM sometimes chooses to summarize the document even for a simple retrieval question. When using expensive LLM models, this mistake could result in 10x higher cost. In your case, the LLM could generate sub-tasks that incur significant operating overheads. Just curious if you're currently facing such issues and if you have plans to mitigate them.

2) The restart idea is neat! I often faced this scenario where only few sub-questions have some issues that need to be fixed. Tweaking them without re-running the whole pipeline seems like a useful feature in this case.

pchunduri6··on Show HN: Demystifying Advanced RAG Pipelines
This is really useful! Using LLM-assisted evaluation seems like the way to go for evaluating RAG applications. One issue I've faced while evaluating responses using GPT-4 is that the evaluation cost can go out of hand rather quickly. Do you have any measures in place or ideas on how to handle this?
pchunduri6··on Show HN: Demystifying Advanced RAG Pipelines
Yes, this is an excellent RAG use-case! The vector index that I use in the repository uses EvaDB [1] to retrieve the top-K matches to the user queries from the available data sources. So, you can manually inspect the best matches to your query from the research article and verify the correctness of the LLM responses.

[1] https://github.com/georgia-tech-db/evadb

pchunduri6··on Show HN: Demystifying Advanced RAG Pipelines
Thanks for the kind words! +1 for chainlit. I love their documentation. Do you have any specific use-cases in mind that would benefit from such pipelines?
pchunduri6··on Show HN: Demystifying Advanced RAG Pipelines
Thanks for the kind words and the great questions!

-- LlamaIndex has some excellent abstractions. In fact, I started off this project with LlamaIndex using their sub-question query engine. However, I found that the abstractions often obfuscate the prompt templates and the pipeline itself from the user. I found that writing my own pipeline was easier than trying to figure out how to engineer the prompts that LlamaIndex was using.

-- It is possible to non-serially issue independent sub-question queries (e.g., using async io). LlamaIndex does something similar. However, I would be extra careful while issuing parallel sub-queries due to the brittle nature of the system.

-- Cool project! I like the fact that the agent decision-making is clearly shown in the UI. A few questions: 1) How do you handle LLM output inconsistencies? 2) Can the user change the prompts for tasks or sub-tasks if the output is not satisfactory? Overall, a great idea and this sub-question query engine might simplify some of the abstractions here.