> So Google hasn't used an LLM to generate and test weird queries ?
What about simple manual testing? Seems to have skipped QA completely, automated or not.
What about simple manual testing? Seems to have skipped QA completely, automated or not.
But there is a more general problem: Big Tech is high on their own supply when it comes to LLMs, and AI generally. Microsoft and Google didn’t fact-check their AI even in high-profile public demos; that strongly suggests they sincerely believed it could answer “simple” factual questions with high reliability. Another example: I don’t think Sundar Pichai was lying when he said Gemini taught itself Sanskrit, I think he was given bad info and didn’t question it because motivated reasoning gives him no incentive to be skeptical.