Flash 3.5 fails exactly like in your sample: https://gemini.google.com/share/97521a8752d9
but Flash 3.1 Lite initially fails, but then corrects itself: https://gemini.google.com/share/dc0889ec85ba
Are you using the flash models? Reasoning models or extended thinking will change the result.
GPT 5.5. Instant shows the same error. If the given prompt isn't working, you can also try "300+140=460 is this correct?". I suspect that leading with the equation may be part of the issue, but haven't tested much.
This is also an probably part of extended prompt that disallowed coding, Gemini always does calculation with a little python snippet because it is deterministic and accurate.
Why would you use an LLM for this? My comment was about the jagged nature of intelligence, so the prompt provides an example of that.
You can see the entire conversation in the shared link. There was no pre-prompt. Even after pushing it to write python, it hallucinated the same output. It later told me that it doesn't have access to a sandbox through the web UI, but it could execute code in a sandbox if invoked via API.