The exact metrics one should check vary with the application. For marketing, we look at creativity, brand tone, and brevity. Customer support chatbots are checked for hallucinations, politeness, and completeness. Code generators are assessed for correct syntax and code quality.
We have came up with a novel grouping of these evaluation metrics into four key categories to help developers navigate which metrics make most sense for their applications: 1. Evaluating Task Understanding and Context Awareness 2. Evaluating Language Quality 3. Evaluating Reasoning Capabilities 4. Custom Evaluations
Would love to hear community's feedback!