LLM unreliability is no impediment in the application at which they most excel. Bullsh*t generation. There they sail way over the bar for e.g. most customer support.
You claon accuracy. LLMs results typically have dire repeatability i.e. same input gives different output. Hence anyone relying on tests for accuracy is kidding themselves.