Show HN: I implemented evals metrics for LLMs that runs locally on your machine
github.com
github.com
If you want to evaluate a fine-tuned model, we have integrations with LM Harness and Stanford HELM coming out. If you want to evaluate a RAG application, we have 7+ metrics available for that.
You can also create your custom metrics using our interface!