I tend to find classic NLP metric more predictable and stable than "LLM as a judge" metrics so I'd try to see if you rely on them more.
We've written a couple of blog posts about some of them: https://www.traceloop.com/blog
We've written a couple of blog posts about some of them: https://www.traceloop.com/blog