https://developers.googleblog.com/en/gemini-2-5-thinking-mod...
https://developers.googleblog.com/en/gemini-2-5-thinking-mod...
How cute they are with their phrasing:
> $2.50 / 1M output tokens (*down from $3.50 output)
Which should be "up from $0.60 (non-thinking)/down from $3.50 (thinking)"
How did I miss this?
In addition, it's also relevant because for the last 3 months people have built things on top of this.
Doesn’t mean you can’t do it, but people won’t be happy.
They told everyone to build their stuff on top of it, and then jacked up the price by 4x. Just pointing to some fine print doesn't change that.
More examples here: https://killedbygoogle.com/
What I have in mind is to start the voice response with a non-thinking model, say a sentence or two in a fraction of a second. That will take the voice model a few seconds to read out. In that time, you use a thinking model to start working on the next part of the response?
In a sense, very similar to how everyone knows to stall in an interview by starting with 'this is a very good question...', and using that time to think some more.
They are obviously excited about their price increase