HNHacker News
TopNewBestAskShowJobs

sourabh03agr

75 karma · joined September 13, 2020

submissionscomments
sourabh03agr··on Tell HN: ChatGPT has become untrustworthy
I have experienced such issues with GPT-4 returning the wrong syntax, especially when asking questions about complex operations in Polars.

We are actually running performance monitoring for GPT-3.5/4 and Claude-2 to systematically track such regressions: https://demo.uptrain.ai/benchmark/. We would be curious to know your thoughts and feedback.

sourabh03agr··on Tracking Prompt Drift for GPT-4 and Claude
Hello HN! I am happy to share the Monitoring reports we have been running for the past few months to identify regression in popular LLMs like GPT-4-turbo, Claude-2, etc.

There have been numerous informal observations about prompt drifts in Large Language Models (LLMs), with the most notable case being GPT-4 showing signs of laziness, especially for coding tasks by the end of the previous year. Discussions on Twitter also hint at a decline in Claude Sonnet’s effectiveness over the past few days. Given the closed-source nature of these models, it's impossible to know what happens behind the scenes, and most often, these drifts go unnoticed until they get flagged by the community.

In November, OpenAI introduced a feature known as model seeding, which aims to ensure consistency in responses given a fixed seed (assuming the temperature is set to 0 and the fingerprint remains unchanged). However, in reality, the model repeatedly displays significant variation and unpredictability over time, even when identical or nearly identical prompts are used for the same seed.

Today, a structured approach to tracking these shifts in the model’s performance is lacking. So we decided to undertake this as a community initiative to monitor prompt drift and identify any regressions systematically.

Our methodology:

1. We've compiled a dataset consisting of 25 samples, each having a question and a context containing information to answer that question. The dataset covers a broad range of topics, including finance, technology, health, sports, and academic areas such as calculus and geology.

2. Each day, we generate responses for three different models: GPT-4-turbo, GPT-3.5-turbo, and Claude-2.1, and run evaluate them on three distinct criteria (with GPT-3.5 serving as our evaluator):

a) Response Conciseness: We examine the extent to which the response includes unnecessary information that does not contribute to answering the question.

b) Response Completeness: We assess whether the model successfully addresses every aspect of the question.

c) Factual Accuracy: We verify the correctness of the model's response against the provided context.

Some of the challenges faced:

1. Choice of evaluator model: We have experimented with different evaluators (e.g., GPT-4-turbo, GPT-3.5, Claude-1.2, etc.). A notable observation was that all these evaluators align well for objective evals (e.g., Factual accuracy) but differ in absolute scores for subjective evals (e.g., Response completeness and conciseness). However, as we are tracking regression over time (and are not concerned with absolute scores), we decided to use GPT-3.5 due to its low cost and high stability.

2. Stability of eval scores: Employing LLMs as evaluators presents its own set of challenges, chiefly the variability in scores across multiple assessments. Tracking performance by using LLM as a judge is an inherently difficult task, as the evaluation score fluctuates over multiple runs. We have made many pipeline improvements to reduce the standard deviation of scores over the same data point across multiple runs to below 2%.

Looking ahead, we plan to enlarge our benchmarking dataset as well as include additional models (ex: Claude 3). We would love to hear your feedback.

sourabh03agr··on Launch HN: Greptile (YC W24) - RAG on codebases that actually works
Congrats on the launch! Do you need Github permissions to answer questions on open-source repos as well?
sourabh03agr··on Show HN: An Experiment with One-Feature Tool Made $7164/Mo
Great read! On being visible on social media efforts, how do you measure the ROI of your efforts, like how much time do you spend vs how much revenue does it bring?
sourabh03agr··on Show HN: Wiwo – Find events happening around you. Google Maps, but for events.
Nice idea! From the perspective of someone who is in need of such a solution, I have explored similar apps before, and one of the biggest issues was that I don't look for events regularly (frequency is once a week or two). As a result, I tend to forget the name of the platform when I actually need it.

I don't have a lot of suggestions on how to solve this but if your platform can be consistently visible to your users, that would help you a lot with adoption.

sourabh03agr··on How does one detect hallucinations?
That's fair, but a lot of use cases require strict information retrieval and don't want the LLM to get creative. I am of the opinion that having an LLM which is always factually correct, is an almost impossible task and we would always need monitoring to catch and fix cases
sourabh03agr··on GeneGPT, a tool-augmented LLM for bioinformatics
Nice work! I see you have an evaluation module - what all are you evaluating for? Primarily Question-answer accuracy via Exact Match?
sourabh03agr··on The Japanese philosophy for a no-waste world
Guess Indian is far better than US and Japan when it comes to producing plastic waste. A lot of Indian states have plastic bans, drastically reducing the amount of plastic wasted via packaging or grocery shopping
sourabh03agr··on Show HN: Zenfetch – Turn your saved browsing content into an AI second brain
Love the product, a happy user here!
sourabh03agr··on Integrating Spade: Synthesizing Assertions for LLMs into My OSS Project
The authors claim that it works well for a wide variety of cases. They have defined a novel categorisation of these evaluations which help guide LLMs to generate relevant assertions
sourabh03agr··on Integrating Spade: Synthesizing Assertions for LLMs into My OSS Project
Operationalizing large language models (LLMs) is challenging, mainly due to their unpredictable behaviors and potentially catastrophic failures. A few of the essential requirements for productionizing LLM applications are evaluating them for a variety of checks as well as monitoring their performance in production.

This process carries two major friction points:

1. Identifying which evaluations to run is tricky: LLM Prompts are growing in complexity. Prompts nowadays have 50+ instructions with many repetitions. SPADE, developed by researchers at UC Berkeley, HKUST, LangChain, and Columbia University, provides a neat framework to identify the critical instructions against which we should validate LLM responses.

It has four key steps: a. Candidate generation: Using an LLM to generate an over-complete list of evaluations based on prompt diffs

b. Filtering redundant evals: Running the evals on sample data to filter out cases where the false failure rates exceed a certain threshold or the evaluation is trivial and always passes.

c. Subsumes checks: Check for subsumes, i.e. if two or more evals effectively do the same check. They check this by prompting LLM to construct a case where function one can return True and function two will return False. If such a case can't be built, functions one and 2 are identical, and one can be dropped.

d. Using an integer programming optimizer to find the optimal evaluation set with maximum coverage and respect failure, accuracy, and subsumption constraints

Their results are impressive. You can look at the SPADE paper for more details: https://arxiv.org/pdf/2401.03038.pdf

2. Running these evaluations reliably is tricky: Recently, using LLMs as evaluators has emerged as a promising alternative to human evaluations and has proven quite effective in improving the accuracy of LLM applications. However, difficulties still exist when running these evals reliably, i.e. high correlation with human judgments and stability across multiple runs. UpTrain is an open-source framework for evaluating LLM applications that provide high-quality scores. It allows one to define custom evaluations via GuidelineAdherence check, where one can determine any custom guideline in plain English and check if the LLM follows it. Additionally, it provides an easy interface to run these evaluations on production responses with a single API call. This allows one to systematically leverage frameworks like UpTrain to check for wrong LLM outputs.

I am one of the maintainers of UpTrain, and we recently integrated the SPADE framework into our open-source repo (https://github.com/uptrain-ai/uptrain/). The idea is simple:

1. Provide your prompt template

2. We use the SPADE framework to identify which evaluations to run

3. We configure UpTrain to run these evaluations on any provided data or monitor these scores in production. All done seamlessly

I would love for you to check it out and provide feedback.

Link for the integration tutorial: https://github.com/uptrain-ai/uptrain/blob/main/examples/int...

sourabh03agr··on Does having talent mean you should work hard to do justice to it?
Yes, but if one has a particular talent, is it justified for us (society as a whole) to exert pressure on that individual to utilize their talent to cause betterment of the society?

Should results for the whole society matter more than results for that specific individual?

sourabh03agr··on Does having talent mean you should work hard to do justice to it?
Well, that's one way to look at it. I guess if the talents are not (currently) socially acceptable, it seems to be perfectly fine (or even at times desirable) to not utilize them properly
sourabh03agr··on [dead]
Evaluating LLM applications is tricky! As Generative models don't have a unique Ground Truth to do a word-to-word comparison against, one typically calculate metrics like factual accuracy, context quality, tonality, etc. to evaluate them.

The exact metrics one should check vary with the application. For marketing, we look at creativity, brand tone, and brevity. Customer support chatbots are checked for hallucinations, politeness, and completeness. Code generators are assessed for correct syntax and code quality.

We have came up with a novel grouping of these evaluation metrics into four key categories to help developers navigate which metrics make most sense for their applications: 1. Evaluating Task Understanding and Context Awareness 2. Evaluating Language Quality 3. Evaluating Reasoning Capabilities 4. Custom Evaluations

Would love to hear community's feedback!

sourabh03agr··on Never say no, but rarely say yes (2011)
The title is catchy!
sourabh03agr··on How soon till we find alien life?
Yes, I am particularly excited about Habitable Worlds Observatory (HWO)
sourabh03agr··on How soon till we find alien life?
I highly doubt that :P. It doesn't make sense right -

1. The nearest habitable planet is atleast 100s of light years away. No way they can travel this far.

2. Even if an alien reaches Earth, don't see any reason on why they would want to transform and live amongst us.

sourabh03agr··on Ask HN: Cool applications of GPT-4V multimodal Generative models?
May not be a popular opinion but I wasn't much impressed by the speech capabilities. It is equivalent to having a ASR + LLM + TTS pipeline or am I missing something?
sourabh03agr··on Ask HN: Cool applications of GPT-4V multimodal Generative models?
Few which come to my mind:

1. Analyzing X-ray scans and answering questions on that 2. App designs to code generation, unit test generation 3. Allowing customer support chatbots to interact with images, especially useful for return orders

sourabh03agr··on Ask HN: What movie about entrepreneurship would you recommend?
The founder and Air are must watch. Also, Guru (Bollywood) is a nice one - story of Dhirubhai Ambani, the man who founded Reliance industries
sourabh03agr··on Show HN: Generative Fill with AI and 3D
The demo looks amazing! Congrats for your first show HN. Quick question on the technical side, do you generate the (added) objects in 3D directly or generate them in 2D and deproject it to 3D? If former, which foundation model are you using?
sourabh03agr··on Show HN: SkillPool – Accelerate your resume screening process by 10x
This looks quite interesting! Do you also consider the technical expertise of the candidates (based on their past experience) while screening the resumes?
sourabh03agr··on Show HN: Platform to practice job interviews using AI
Looks interesting! Do you plan to add technical questions in your AI-powered interview offering? As a founder who spends considerable time hiring, I typically try to judge the technical capabilities of the candidate during the first interview.
sourabh03agr··on Show HN: Kusho – Copilot for API Testing
Looks interesting!
sourabh03agr··on Outperforming larger language models with less training data and smaller models
Interesting! Do you think RLHF would be a necessity for smaller models to perform as par as state-of-the-art LLMs? In my view, instruction tuning will resolve any isssues related to output structure, tonality or the domain understanding but will it be enough to improve the reasoning capabilities of the smaller model?
sourabh03agr··on How are generative AI companies monitoring their systems in production?
You can add UpTrain to the list. We are building an open-source LLM evaluation tool with pre-built evaluations such as factual accuracy, retrieval quality, response completeness, tonality, etc. as well as easily extendable framework which allows LLM developers to define their custom evaluations by chaining individual operators. Checkout our demo here: https://demo.uptrain.ai/evals_demo/
sourabh03agr··on Launch HN: Loops (YC W22) – Email for SaaS Companies
Congrats on the launch! QQ - do you track if the sent email is getting flagged as spam?
sourabh03agr··on Show HN: UpTrain – A Practical Approach to Finetuning LLMs for Custom Use-Cases
Few details on our approach:

Step 1 - Use the visualisation functionalities (ex: UMAP for BERT embeddings) to qualitatively assess the distribution shifts happening.

From here, we saw that Dialogsum & Samsum datasets indeed belong to two different clusters and we can expect performance degradation due to data drift.

Step 2 - Use statistical techniques to identify clusters near few low-performing samples (samples selected by us) and find data-points belonging to them.

Interestingly, this gave us a nice collection of low-performing datapoints (accuracy ~ 40% lower than that on whole dataset). Upon manually inspecting, we saw some interesting behaviours around model failures which we will use to generate retraining datasets

Step 3 - Use UpTrain's Custom Signal interface to define rules for collecting edge-cases. We defined two rules:

1. We saw model outputs incomplete summaries when input text length is too long. Hence, we defined a rule on number of words in the input conversation.

2. In many cases, we saw model selects one or two sentences from the conversation as the summary. This works well generally but fails miserably when the conversation is all about negating those sentences. We defined a rule for the same.

Step 4 - We also wanted to check if we can detect a shift in the vocabulary between the two datasets and defined a custom monitor for the same. (Interestingly we saw higher occurrences of words related to Asia in the dialogsum dataset). Similar monitors can be designed to identify newer topics, sentiments, tone of voice, etc.

Would love you to play around with the tool and provide your feedback!

sourabh03agr··on Launch HN: UpTrain (YC W23) – Open-source performance monitoring for ML models
Thanks! Reaching out to Stefan
sourabh03agr··on Launch HN: UpTrain (YC W23) – Open-source performance monitoring for ML models
Can you describe your use case?

We also faced the same problem with other tools and hence building UpTrain with customisation at the core of it. Would be interesting to see if your use case fits

Page 1 of 2Next →