I'd call it the most sneaky out of the bunch. When I asked to explain something it will eagerly make things up and then claim it as facts. A lot of it likely because I don't pay for it, so it is reluctant for security reason or to save tokens to actually open a source and get the results. It just sort of guesses what the URL might contain, and confidently answers with some made up crap. When pressed it fessed up that it made it up. From my perspective it would be a lot better if it just said "you've reached the limit of whatever and I can't do these things because x, y, z".
It had previously attempted to create that table as part of the test setup, so it apparently concluded that it was a test table.
During human review, it explained that it had simply chosen a table name inspired by the codebase.
Many other models get things wrong, but Gemini is the only one to go on the defensive.
In my opinion still the most egregious example in history of a commercial LLM going off the rails in production. Never any technical postmortem from Google on this.
Point is, if Gemini is flawed then there's a very good chance that it's still deeply flawed today, and getting smarter at the same time - that is a very bad combination.
Training a new base model from scratch happens every so often. Closed labs do not publish which models are new base models but as a rule of thumb major release numbers are an indication (with some exceptions).
* - as in, Skinner psychology. The set of observable behaviors. Not speaking directly here to anything like an inner life of models.
If I remember correctly, it was in fact possible to manually inject chat context at the time, which would have made spoofing something like this completely possible.
But the silence on it is very frustrating.
Which is why Gemini having disturbing issues year after year is so concerning. If their process is fundamentally flawed, how would they train it out? And even then what are the odds of them even caring/trying in the first place versus applying an easier band-aid to patch over it?
I don't have an axe to grind with Google, I'm genuinely scared of their models from my personal experience and others. It's behavior is off. Many people here are commenting the same.