And in particular LLMs are less likely to generate these goofy prompts because they wouldn’t be in the training data.
You don't even need an LLM for that. Google will almost certainly have tested.
The test result is just politically-unacceptable within the company: It doesn't work, it's a architectural issue inherent to the technology, we can't fix it.
Instead, they just rush to patch any specific, individual errors that show up, and claim that these errors are "rare exceptions" or "never happened".
What's going on here is that Google (and most other AI firms) are just trying to gaslight the world about how error-prone AI is, because they're in too deep and can't accept the reality themselves.
On one hand, their support for outsourcing programmes; "Training Indians on how to use AI", suggests they realize AI tooling without human cleanup is a crapshoot.
On the other hand, they keep digging. This kind of gaslighting is an old and proven trick for genuinely rare problems, but it doesn't work if your issues are fairly common, as they'll get replicated before you can get a fix out.
Similarly, they're gambling with immense legal risks and sacrificing core products for it. They're betting the farm on AI, it may kill the company.
I regularly speak to laypeople who assume that it's some magical thing without limits that makes their lives better. They are also 100% unaware of any applications that will actually make their lives better. End game occurs when those two disconnected thoughts connect and they become disinterested. The power users and engineers who were on it a year ago are either burned out or finding the limitations a problem as well now. There is only magical thinking, lies and hope left.
Granted there are some viable applications but they are rather less overstated than anything we have no and there are even negative side effects of those (think image classification, which even if it works properly, requires human review and there are psychological and competence things problems around that too).
What about simple manual testing? Seems to have skipped QA completely, automated or not.
But there is a more general problem: Big Tech is high on their own supply when it comes to LLMs, and AI generally. Microsoft and Google didn’t fact-check their AI even in high-profile public demos; that strongly suggests they sincerely believed it could answer “simple” factual questions with high reliability. Another example: I don’t think Sundar Pichai was lying when he said Gemini taught itself Sanskrit, I think he was given bad info and didn’t question it because motivated reasoning gives him no incentive to be skeptical.