Just to add to Nir's answer here:
Let's say your application takes several steps to build up a prompt dynamically, such as a RAG pipeline. You'll end up with a different prompt for potentially each user, depending on the application.
The result is you've likely increased the accuracy of the LLM, but at the expense of understanding the whole system's behavior by introducing more steps upstream of the LLM call. Those steps could be super simple, or they could be (like in our case) dozens of steps that could all potentially fail or have a bug or whatever.
And so how do you wrangle all of this in context? You need something like OpenLLMetry that treats a request to an LLM as one of several components that make up a request and/or user experience. Otherwise you're just throwing stuff at the wall, guessing at what could improve stuff (or guessing at what could make an eval score better).