If there is one problem I have to pick to to trace in LLMs, I would pick hallucination. More tracing of "how much" or "why" model hallucinated can lead to correct this problem. Given the explanation in this post about hallucination, I think degree of hallucination can be given as part of response to the user?
I am facing this in RAG use case quite - How do I know model is giving right answer or Hallucinating from my RAG sources?