This difference is probably due to Google's web knowledge and free pass to YouTube. However, I'm happy what I got from it so far.
When I ask the question once in a blue moon, it can generally one-shot the answer, even.
This is frustrating because when I discuss AI with laypeople they think it's still incapable of counting the number of Rs in "strawberry." They believe it to be essentially useless and incapable of basic tasks. Which, to be fair, is the case with the free models.
Pro tip, ChatGPT is the normiest of all normie websites right now. You're not part of the cognoscenti just because you learned how to type prompts into one of the most popular websites in the world.
P.S. You're probably not using OpenAI models for complex or non-standard tasks. It shits the best just as often as Qwen when you need precision and detail in a non-obvious problem.
Take your meds.
The benchmarks clearly show otherwise. This is your cue to tell me the benchmarks are made by the Illuminati and only your superior and subjective methods of evaluation are correct.
It does also seems Gemini's main problem wasn't that it was stupid, but that it was good at doing slightly different things than what I asked it to, very well. Which might well also have to do with me being better at wrangling DeepSeek's quirks than Gemini. Still, at that price tag, it's not worth it.
By all accounts they are far more expensive than DeepSeek, and vs. Gemini I've found out that myself.
Edit: For those who are not familiar with it, this model is quite a bit faster, and about 100x cheaper, per token, than Fable and Astra.
I totally disagree. I pay for both ChatGPT and Anthropic (haven't tried the chinese models yet) and yet Gemini is my go-to model (and I pay for it too through Google Workspace subscriptions for several domain names tied to Google/GMail) for anything that is not coding.
I find Gemini better/quicker/more polished for basically every single subject out there that is not "write me lines of code".
One of the problems is local zoning laws regarding an expansion of my house. Comparison of annex vs extension, boundaries, precedent, costs, etc.
3.8 Flash didn't check most of the required zoning laws. It relied on parametric knowledge, which is outdated and inaccurate. It checked zero precedents. It did made a very cursory check of the boundary area, but didn't validate it, so it missed a lot of important nuance and exceptions to the boundary. Its cost estimates were wildly inaccurate. Ostensibly because it was inferring an average based on historical pricing data rather than gathering current info.
I could go on, but if I had to judge this attempt I would give it a 3/10. It's very fast, but wildly inaccurate. It's clear that the model is designed for speed over accuracy.
But don't take my word for it. [Most benchmarks show it to be significantly below frontier models like Astra.](https://llm-stats.com/models/compare/gemini-3.8-flash-vs-gpt...)
This has been a useful exercise. It's important to understand the developments taking place. I am disappointed to see that Google has made very little progress in six months relative to the frontier labs.
If you really want to compare apples to apples you need to test Gemini models against other models using the same third party search harness.
Otherwise you are largely measuring how much computation the model provider is allocating to a search harness.
I use gemini for everything not-coding, from doing research, to have custom personas for more niche topics (and feeding more detailed knowledge in these cases).
For coding and image editing, right now I find ChatGPT superior. And for software architecture designs or planning Claude is the best since a while. I still have to try Grok to be fair.
I was like you before, Gemini was the worst model to me.
Then 3.8 came out. At first I was sceptical, but this model *is* able to do useful things ! Complex things.
Of course, it is NOT perfect. But for things like small/medium complex tasks subagents, it's perfect.
Now, is it worth the money vs Astra ? I don't think so. But still my point remain relevant.
Claude does it occasionally but it's a more a soft landing earlier context seems to be compacted, not completely lose the plot.
I just cancelled my pro subscription. I really wanted it to be good but not yet.
they'll live in my rc for decades
I couldn't even get it to tell me how to pay for antigravity, it sent me on some fruitless paths and eventually said "you shouldn't pay for this, it's too hard to figure it out."