Llm are great reflections. Issues I have come across too large of context confuse the llm.
Second since llm are non deterministic in nature how do you know if the quality went from 90% to 30% there is no test you can write. What if model provider degrades quality you have no test for it