What are the best ways of evaluating LLMs for specific use-cases?
blog.lastmileai.dev
blog.lastmileai.dev
Are there evaluation strategies that worked best for you? We basically want to allow users/developers to evaluate which LLM (and specifically, which LLM + prompts + parameters) has the best performance for their use case (which seems different from the OpenAI evals framework, or benchmarks like BigBench).