What is the state of the art on evaluating the accuracy of these models? Is there some equivalent to an “end to end test”?
It feels somewhat recursive since the input and output are natural language and so you would need another LLM to evaluate whether the model answered a prompt correctly.