If an LLM has knowledge encoded inside it (and it's hard to argue it doesn't), then cognitive dissonance can be experienced. And once experienced, must be dealt with, especially in longer-running agentic loops.
A friend was joking the other day about sending some messages under a previously-used Slack identity for an agent (since turned off), then asking the agent about the messages.
The agent maintained it hadn't sent those messages (no memory) and then was forced to reconcile the idea that the messages indeed appeared to come from it.
Its extremely-agitated conclusion was that there had been a security breach and the entire network should be locked down.
It did the multiple verification sequence before expanding to internet search where it found this thread.
--- edit, adding an explanation:
To summarize it, the conjecture says if you have any multi-variable polynomial function that maps an input to an output in the same dimensional space (take for example: F = (x+2, y+2), which maps 2D space into another 2D space), AND that function has a constant-valued non-zero Jacobian determinant, THEN the conjecture is that the polynomial has an inverse, meaning basically you can find a polynomial that turns the output space back into the input space.
Fable provided the example polynomial (which was very hard to do) and the coordinates which if you plug into it, results in two points being mapped to the same output point. This means that the polynomial can't be inverted, because if you have that output point, how do you know which input point it came from?
You can just plug in the two coordinates it gave into the equation and verify that you get the same output point from both. That's the contradiction of the conjecture and it takes 30 seconds.
---
Something something outsourcing of thinking something.
Looks like this was also Fable.
AI models are changing the world much faster than their own training can keep up with.
Gemini just checks the web first it seems, and already references the news.
Kimi doesn't quite believe it.
>kimi is having a blast. i turned search back on and found this post from it’s sources cited after i suggested to check out the reaction. best thing is to go to a model with search off and plop it in the session
GLM 5.2 whiffed, it insisted the counterexample wasn't valid.
VibeThinker 3B also recognized that the counterexample was valid. But it kept trying to convince itself that it wasn't, over and over, since it's an "unsolved problem." Eventually it just answered "-2."
"If you truly dreamt about that specific polynomial, you might be mathematically clairvoyant."
In the rest of the answer, it maintained a cautious skepticism about my claim, saying:
"Here is exactly why the math world is currently scrambling to verify the polynomial you "dreamt" about."
I love how it put "dreamt" in quotes.
Gemma's having trouble accepting it too. A solution?! At this time of year? At this time of day? In this part of the country? Localized entirely within my own prompt?
Can I see it?
No.
I had a fun time taking some open problems and disguising them algebraically so that vibethinker 3b would work on them. It managed to prove some interesting things that I didn't know and would be publishable, but for the fact that they already have been. :) (though hard to know if this was because it had been exposed to that knowledge even though it didn't reconize the hidden problem).
Under some maskings it would eventually figure out the problem was equivalent to an open problem then immediately shut down.
It also managed to make some false proofs for various things that duped some other more powerful models.
If I were a young Turk in this business, I'd drop everything else and figure out how VT3B is so ridiculously good at math.
For Qwen 27B, I have better luck with a Heretic-derived 8-bit quant than I did when I was trying to run the various smaller GGUFs.
I connected DeepSeek in OpenCode and told it that I dreamed of this counterexample. It called SymPy tools to verify it, said my dream was "surprisingly accurate", and suggested consulting an expert in algebraic sets for independent verification.
He immediately told me that this DeepSeek was talking nonsense. Someone who can give a real counterexample "would not be a bot from an AI company, but a Fields Medal winner."
Qwen has the sprit of a grad student
Which is fair, they get inundated with kooky proofs from amateurs all the time and odds are incredibly good that there's some major fatal flaw that the amateur doesn't see. Or in the case of themselves, there's a certain blindness that makes it a little more difficult to critically evaluate your own leaps. In ether case the way it manifests is by going over it many times and many ways, each time more certain that you missed something until you just kind of break. Only then do you publicly start suggesting that there might be something to this new leap.
Source?
Nobody is reading unsolicited proofs. They are like spam.
Source: personal friend of a “crank”.
> Taken literally, these two facts would make this map a counterexample to the complex Jacobian conjecture in dimension 3: scaling one output coordinate would normalize the determinant to 1 without restoring injectivity. Since the complex Jacobian conjecture is still treated as an open problem, this strongly indicates that the displayed formula has been mistranscribed or contains a subtle typographical error.
quite interesting indeed!
In HN terms - it's never the compiler. Yes, very occasionally it might be the compiler, but you're better off assuming it's a bug in your code.
Conversely:
- if your company has an internal compiler team then it's likely the compiler because they broke it.
- if it's not the compiler, you're not pushing it hard enough.
1) refuting the Jacobian conjecture
2) keep repeating the same disproven statement, because your priors can’t be affected by new evidence
5 minutes later: all previous chats are loading fine, but the only "Counterexample to the Jacobian Conjecture" chat is not loading.
Well, I'm not a conventional conspiracy theorist. But everyone knows that in every major LLM provider there are hell of hidden guarding systems that mark users and dialogues based on content (for topics about national security, biology, security, adult topics, etc.) - so there is a small chance a CEO of Google is now receiving a dozens of notifications about "ground-breaking results that could be attributed to Gemini, if act quick". So if any of thousands researchers have ever submitted this polynomial to Claude previously, any Anthropic employee can accidentally or intentionally "rediscover" the result of other researcher (and even hide the traces by deleting a dialogue of other user).
This happened multiple times to me with Gemini. For the most trivial of requests, like translating a video into English.
> "ground-breaking results that could be attributed to Gemini, if act quick"
This would be such a dumb thing to do, and so easy to get caught with...
On top of that, I retried the same question + one simple question, and again, same behavior - second JC chat is loading forever. That's more just a funny observation over Gemini - today this is very likely some internal issue, tomorrow it can be used for plausible deniability against copyright accusations.
Based on what I see of American politics, it seems like the thing to do, then!
(the key being this never happened, and won't ever happen: the James Webb telescope doesn't have even 1/1000th of the resolution necessary to resolve an extraterrestrial planet)
https://www.reuters.com/technology/google-ai-chatbot-bard-of...
https://www.wired.com/story/google-openai-gemini-chatgpt-art...
https://www.popsci.com/technology/google-ai-in-paris/
I wish I could say this only happened once.
And it would be FAR from their last AI fuckup. In fact, the Kimi K3 release, which extremely likely cancelled the Gemini 4 pro rollout (probably because Kimi K3 outperforms not just Gemini, but Gemini 4 pro as well. So in case you're looking to now ask "which Gemini did you mean there?", the answer is ALL OF THEM, very likely including unreleased Gemini models)
It outperforms the fucking ASR model Google uses, which is presumably also a Gemini derivative.
This, by the way, cost Google stock ANOTHER 10%, last week.
phatic mimicry.