I'm assuming that Argon has at least a June 2026 date, but man, the 3 series models were a mess with newer information.
> The knowledge cutoff date for Gemini 3.8 Flash is March 2026
https://deepmind.google/models/model-cards/gemini-3-8-flash/
The "some domains" are very narrow. They likely just RL'ed popular queries.
Of course, it's called "flash," and that implies its purpose. I have little use for speed and a LOT of use for accuracy, so I'm hopeful 4.0 is much better. I saw a benchmark earlier today showing that it is much less prone to hallucinations. Let's see.
- AI is just a tool, like excel; it does what the human operating it tells it to
- next token prediction cannot be true understanding
- models can have no desires and goals, don't anthropomorphize it
However, "hallucination" is very much not one of them
It didn't work. Gemini: "Oh yeah, that obviously cannot work, it's not possible to do it through OBD2" (paraphrasing)
It was quite funny to me, but a bit less so to my colleague.
The bigger problem is that people expect LLMs to know what facts are. That assumption is even baked into the term "hallucination." Someone who hallucinates is expected to otherwise have a grounding in objective reality, to "not" hallucinate, and to be able to recognize reality from fantasy. We wouldn't allow a person who "hallucinates" as much as an LLM anywhere near the roles we give to LLMs. But everything an LLM does is as much a "hallucination" as anything else, it's just stochastically generating grammar. Some grammar just happens to be useful because of the quality of its training data, which was probably created by humans who do possess interiority and awareness of fact.
And it isn't "broken" either. Broken assumes that the correct mode of operation is to act as a source of truth or fact generation. When LLMs "apologize" for bad results, for instance they aren't actually apologizing. Try getting it to apologize for returning the correct data. It probably will. There is no cognition happening. It doesn't know either way. It isn't a calculator crunching numbers or a computer doing data analysis. It's just pattern matching.
"Hallucination" is no less correct than "confabulation" which also presupposes intent and contextual awareness. Unfortunately the way LLMs operate is so unintuitive (as opposed to the intuitive nature of the interface) that the only language we have to describe it is the language of human behavior, with all of the biases and false assumptions that brings.