Does DeepEval allow you to set up custom metrics without an LLM-as-a-judge base?
If I want my result to be a JSON output, and I want to weight the keys based on some specific importance weighting, can I write a Python function/class to calculate and average those weighted scores as a metric for DeepEval?
I do have some annoyances with DSPy, but I think their approach to defining evals is decent.