For small models (which are probably distilled from their big ones) you can serve them economically all the time and not hemorrhage money.
For smaller models, you're competing with DeepSeek V4 Flash. (Which I think is a 284B A13B?) Subjectively, this feels about as smart as Sonnet 4.5, give or take. And it costs $0.09/$0.18 on Open Router, compared to $1/$5 for the latest Claude Haiku. See https://openrouter.ai/deepseek/deepseek-v4-flash#providers The developer antirez of Redis fame uses this as a local coding model.
DeepSeek did some extremely clever research on hybrid attention to get the prices that low, reducing per-user context cache sizes dramatically.
So, no, when it comes to low-price models, the US models probably can't sustain their current margins there, either.
Antigravity NEEDED to be game-changing. Without the stream of data that Claude, Codex, and Cursor enjoy there is little chance of getting an effective reinforcement learning loop. For the first time in its history, GOOG is at a meaningful data disadvantage, and apparently a cultural one as well.
Google literally giving everyone + student 18 month free subscription, those are source of cheap gemini + sonet,opus model that people selling/use with rotator proxy with thousands of account
they didn't lack the data
I don’t think this is necessarily true, did we all forget how much Google cared about alignment that their AI wasn’t able to render a white polar bear?
https://www.searchenginejournal.com/pichai-says-google-is-a-... (link to the actual podcast interview source within, this has a summary)
My view then was they are optimising the models for inference ability on their own hardware AND use cases, which is often speed and time to first token.
They've somehow seemed to end up with terrible compute shortages, which again is surprising given how good Google is at infra deployments AND have their own hardware. From rumors out there they are turning down enterprise deals for Gemini because they don't have the compute.
The problem is they're falling further and further behind on frontier class on coding especially, and since I wrote that article it's got even worse with open weights models undercutting them on price AND intelligence.
So my guess is that Google will continue having compute shortages until the Gemini enshittification starts.
Let's assume Google serves AI overviews on every SERP (they don't) and don't cache them (they do, afiak).
And let's assume that each AI overview is 2000 tokens (blended input/output), that's 500T tokens a month.
It's rumoured that anthropic is serving somewhere close to 10Q tokens a month.
Now it may be that AI overviews uses vastly more tokens than that per search, but I doubt it based on speed to render the overview.
My very rough napkin math on this is that maybe AI overviews is consuming 100T tokens/month max (after adjusting for caching and SERPs that don't have them), which would be 1% of Anthropic token volume.
"10 Quadrillion tokens a month means: 333 Trillion tokens per day and 3.85 Billion tokens generated/processed every single second, 24/7."
"At an incredibly cheap, subsidized infrastructure cost of $1 per million tokens, serving 10 Quadrillion tokens would cost Anthropic $10 Billion per month ($120 Billion a year) just in inference compute."
It also had this to say about how google's AI overview works: "Google doesn't just feed the LLM your 5-word search query. The system scrapes the top 10–20 web results, feeds thousands of words (tens of thousands of tokens of context) into the model, processes it, and then outputs the result."
Oh, and it does all of that in less than two seconds. Honestly, whatever Google is doing with its infrastructure is so far ahead of everyone else, I can't believe you fell for such an obvious lie.
There are also extremely obvious holes in your comment:
>Let's assume Google serves AI overviews on every SERP (they don't) and don't cache them (they do, afiak).
Try it out for yourself. Add a few random letters or punctuation. They cache nothing.
They definitely cache results - I've searched and re-searched an identical query back to back a few times and seen identical results from overview. They are definitely throwing a stupid amount of compute towards these ai results nobody is paying for - changing punctuation and stuff does get you a different response - but they're not doing no caching.
Certainly what they're doing with their infrastructure is impressive but it's not super meaningful at the end of the day for a for profit company to be really impressively good at burning tens of billion dollars on a service nobody pays for while the same tech from their competitors is quickly becoming one of the largest spend categories for many software engineering teams
Speed as a differentiator has always been Google's thing. They (used to?) show the microseconds it took to query & rank web-scale search results. Chrome, notoriously, focused on speed at the expense of resource use. The very many efforts to efficiently speed up Android & its runtime since its inception, and so on...
> their big model underperforms chatgpt 5.6
Possible but TFA claims:
We have started our most ambitious pre-training run yet, for Gemini 4 ...It would be a shame if they cannot beat Kimi K3 or Qwen3.8 Max, both of which are claimed to be Fable-like. If that is true, it will be [or would be] the first time a major American lab falls behind a Chinese competitor.
China can keep up because it's cheaper to run a frontier lab there. They also have more researchers and a stronger cultural inclination for this sort of thing. And I guess the business case in China doesn't have to work as well as it does in the US.
Not sure if this is what you meant, but their training runs are significantly cheaper. This was one of the big shockers from the Deepseek R1 paper. US foreign policy has helped to ensure that the Chinese are compute constrained, so they literally cannot buy the most expensive and powerful training rigs.
This has led to a steady drumbeat of innovations which are not revolutionary on their own but stack together to make things much more efficient.
like it literally pennies
When it is cheaper, and the "lower quality" model is adequate for the task at hand.
Plenty of problems have a low(er) skill/intelligence floor, anyone who uses the dual-mode agent paradigm (plan, then act) figures out the second phase can be completed by a less capable model. Even when disregarding costs - speed is important here because the agent can rapidly iterate without human supervision, based on compiler errors, lint and test failures
Paywalled article, but the headline is basically all you need: https://www.bloomberg.com/news/articles/2026-07-16/google-ge...