Tests are called "evals" (evaluations) in the AI product development world. Basically you let humans review LLM output or feed it to another LLM with instructions how to evaluate it.
https://www.lennysnewsletter.com/p/beyond-vibe-checks-a-pms-...
https://www.lennysnewsletter.com/p/beyond-vibe-checks-a-pms-...