Evaluating LLM applications is tricky! As Generative models don't have a unique Ground Truth to do a word-to-word comparison against, one typically calculate metrics like factual accuracy, context quality, tonality, etc. to evaluate them.
The exact metrics one should check vary with the application. For marketing, we look at creativity, brand tone, and brevity. Customer support chatbots are checked for hallucinations, politeness, and completeness. Code generators are assessed for correct syntax and code quality.
We have came up with a novel grouping of these evaluation metrics into four key categories to help developers navigate which metrics make most sense for their applications: 1. Evaluating Task Understanding and Context Awareness 2. Evaluating Language Quality 3. Evaluating Reasoning Capabilities 4. Custom Evaluations
Would love to hear community's feedback!