We are actually running performance monitoring for GPT-3.5/4 and Claude-2 to systematically track such regressions: https://demo.uptrain.ai/benchmark/. We would be curious to know your thoughts and feedback.
75 karma · joined September 13, 2020
We are actually running performance monitoring for GPT-3.5/4 and Claude-2 to systematically track such regressions: https://demo.uptrain.ai/benchmark/. We would be curious to know your thoughts and feedback.
There have been numerous informal observations about prompt drifts in Large Language Models (LLMs), with the most notable case being GPT-4 showing signs of laziness, especially for coding tasks by the end of the previous year. Discussions on Twitter also hint at a decline in Claude Sonnet’s effectiveness over the past few days. Given the closed-source nature of these models, it's impossible to know what happens behind the scenes, and most often, these drifts go unnoticed until they get flagged by the community.
In November, OpenAI introduced a feature known as model seeding, which aims to ensure consistency in responses given a fixed seed (assuming the temperature is set to 0 and the fingerprint remains unchanged). However, in reality, the model repeatedly displays significant variation and unpredictability over time, even when identical or nearly identical prompts are used for the same seed.
Today, a structured approach to tracking these shifts in the model’s performance is lacking. So we decided to undertake this as a community initiative to monitor prompt drift and identify any regressions systematically.
Our methodology:
1. We've compiled a dataset consisting of 25 samples, each having a question and a context containing information to answer that question. The dataset covers a broad range of topics, including finance, technology, health, sports, and academic areas such as calculus and geology.
2. Each day, we generate responses for three different models: GPT-4-turbo, GPT-3.5-turbo, and Claude-2.1, and run evaluate them on three distinct criteria (with GPT-3.5 serving as our evaluator):
a) Response Conciseness: We examine the extent to which the response includes unnecessary information that does not contribute to answering the question.
b) Response Completeness: We assess whether the model successfully addresses every aspect of the question.
c) Factual Accuracy: We verify the correctness of the model's response against the provided context.
Some of the challenges faced:
1. Choice of evaluator model: We have experimented with different evaluators (e.g., GPT-4-turbo, GPT-3.5, Claude-1.2, etc.). A notable observation was that all these evaluators align well for objective evals (e.g., Factual accuracy) but differ in absolute scores for subjective evals (e.g., Response completeness and conciseness). However, as we are tracking regression over time (and are not concerned with absolute scores), we decided to use GPT-3.5 due to its low cost and high stability.
2. Stability of eval scores: Employing LLMs as evaluators presents its own set of challenges, chiefly the variability in scores across multiple assessments. Tracking performance by using LLM as a judge is an inherently difficult task, as the evaluation score fluctuates over multiple runs. We have made many pipeline improvements to reduce the standard deviation of scores over the same data point across multiple runs to below 2%.
Looking ahead, we plan to enlarge our benchmarking dataset as well as include additional models (ex: Claude 3). We would love to hear your feedback.
I don't have a lot of suggestions on how to solve this but if your platform can be consistently visible to your users, that would help you a lot with adoption.
This process carries two major friction points:
1. Identifying which evaluations to run is tricky: LLM Prompts are growing in complexity. Prompts nowadays have 50+ instructions with many repetitions. SPADE, developed by researchers at UC Berkeley, HKUST, LangChain, and Columbia University, provides a neat framework to identify the critical instructions against which we should validate LLM responses.
It has four key steps: a. Candidate generation: Using an LLM to generate an over-complete list of evaluations based on prompt diffs
b. Filtering redundant evals: Running the evals on sample data to filter out cases where the false failure rates exceed a certain threshold or the evaluation is trivial and always passes.
c. Subsumes checks: Check for subsumes, i.e. if two or more evals effectively do the same check. They check this by prompting LLM to construct a case where function one can return True and function two will return False. If such a case can't be built, functions one and 2 are identical, and one can be dropped.
d. Using an integer programming optimizer to find the optimal evaluation set with maximum coverage and respect failure, accuracy, and subsumption constraints
Their results are impressive. You can look at the SPADE paper for more details: https://arxiv.org/pdf/2401.03038.pdf
2. Running these evaluations reliably is tricky: Recently, using LLMs as evaluators has emerged as a promising alternative to human evaluations and has proven quite effective in improving the accuracy of LLM applications. However, difficulties still exist when running these evals reliably, i.e. high correlation with human judgments and stability across multiple runs. UpTrain is an open-source framework for evaluating LLM applications that provide high-quality scores. It allows one to define custom evaluations via GuidelineAdherence check, where one can determine any custom guideline in plain English and check if the LLM follows it. Additionally, it provides an easy interface to run these evaluations on production responses with a single API call. This allows one to systematically leverage frameworks like UpTrain to check for wrong LLM outputs.
I am one of the maintainers of UpTrain, and we recently integrated the SPADE framework into our open-source repo (https://github.com/uptrain-ai/uptrain/). The idea is simple:
1. Provide your prompt template
2. We use the SPADE framework to identify which evaluations to run
3. We configure UpTrain to run these evaluations on any provided data or monitor these scores in production. All done seamlessly
I would love for you to check it out and provide feedback.
Link for the integration tutorial: https://github.com/uptrain-ai/uptrain/blob/main/examples/int...
Should results for the whole society matter more than results for that specific individual?
The exact metrics one should check vary with the application. For marketing, we look at creativity, brand tone, and brevity. Customer support chatbots are checked for hallucinations, politeness, and completeness. Code generators are assessed for correct syntax and code quality.
We have came up with a novel grouping of these evaluation metrics into four key categories to help developers navigate which metrics make most sense for their applications: 1. Evaluating Task Understanding and Context Awareness 2. Evaluating Language Quality 3. Evaluating Reasoning Capabilities 4. Custom Evaluations
Would love to hear community's feedback!
1. The nearest habitable planet is atleast 100s of light years away. No way they can travel this far.
2. Even if an alien reaches Earth, don't see any reason on why they would want to transform and live amongst us.
1. Analyzing X-ray scans and answering questions on that 2. App designs to code generation, unit test generation 3. Allowing customer support chatbots to interact with images, especially useful for return orders
Step 1 - Use the visualisation functionalities (ex: UMAP for BERT embeddings) to qualitatively assess the distribution shifts happening.
From here, we saw that Dialogsum & Samsum datasets indeed belong to two different clusters and we can expect performance degradation due to data drift.
Step 2 - Use statistical techniques to identify clusters near few low-performing samples (samples selected by us) and find data-points belonging to them.
Interestingly, this gave us a nice collection of low-performing datapoints (accuracy ~ 40% lower than that on whole dataset). Upon manually inspecting, we saw some interesting behaviours around model failures which we will use to generate retraining datasets
Step 3 - Use UpTrain's Custom Signal interface to define rules for collecting edge-cases. We defined two rules:
1. We saw model outputs incomplete summaries when input text length is too long. Hence, we defined a rule on number of words in the input conversation.
2. In many cases, we saw model selects one or two sentences from the conversation as the summary. This works well generally but fails miserably when the conversation is all about negating those sentences. We defined a rule for the same.
Step 4 - We also wanted to check if we can detect a shift in the vocabulary between the two datasets and defined a custom monitor for the same. (Interestingly we saw higher occurrences of words related to Asia in the dialogsum dataset). Similar monitors can be designed to identify newer topics, sentiments, tone of voice, etc.
Would love you to play around with the tool and provide your feedback!
We also faced the same problem with other tools and hence building UpTrain with customisation at the core of it. Would be interesting to see if your use case fits