Show HN: DeepEval – Evaluation and Unit Testing for LLMs
github.com
github.com
As more companies integrate LLM/RAG applications into their operations, ensuring the effectiveness, reliability and safety of these models is hard.
About DeepEval We started with consulting on a few RAG projects and quickly realised how many issues came up when we iterated on our prompts, chunking methodologies, added function calls, added guardrails, etc. Very quickly we realised this had downstream effects that caused unexpected problems and results.
DeepEval, inspired by Pytest, aims to make iterating on these RAG and agent applications as easy as possible by building evaluation into part of their CI/CD workflow. The goal is to make deployment of LLMs as straightforward as getting all tests to pass.
Some features of DeepEval include: - Opinionated tests for answer relevancy, factual consistency, toxicness, bias. - Web UI to view tests, implementations, comparisons. - Opinionated flow for synthetic dataset creation.
We are currently in a fully operational beta release and would love your feedback and suggestions to continue improving DeepEval.
We are happy to answer any questions you have!
Here are the problems with what LangSmith offers: 1. Focusing on observability brings insights to engineering teams in terms of cost, latency, but these conclusions aren't really actionable
2. Not all teams use LangChain, plus a lot of teams who originally used LangChain to prototype is moving away as they productionize and look to improve what they've built
3. Not every LLM implementation involves a chain of thought, lLamaIndex is a good example. Apart from vendor locking, there are other use cases such as evaluating fine tuned models that are overlooked in a closed source software like LangSmith
4. Using LLMs to evaluate themselves are fine, but a lot more can be done. We offer other metrics (metrics that can be quantified from 0 - 1) such as factual consistency trained on NLI models that are much more deterministic. You can also use our package to pick and choose metrics that you care about (bias, relevancy, toxicity, etc)
Hope that answers your questions.