It's not a bad model by any means, but I just don't know what situation I'd reach for 3.7 Flash first for. Google really needs a differentiator, especially given how hard it is to get an API key from them. They can't be high friction and non-pareto.
It's not a bad model by any means, but I just don't know what situation I'd reach for 3.7 Flash first for. Google really needs a differentiator, especially given how hard it is to get an API key from them. They can't be high friction and non-pareto.
Disclaimer: I work in Google so it might be that this link is not publicly well known
At a high level though, as a rule of thumb Google assumes that they're serving companies at Google scale first, and at a human scale second. For other companies it's the opposite. Generally what that means is the first experience you get with a Google product will route you through 8 different dashboards to set up ACLs before you've hired your 2nd employee.
Also:
> Google assumes that they're serving companies at Google scale first
So much this. I'm currently grandfathered in until the end of the year on Google's Search API, but the $35,000 they want to continue usage of my < 1000 personal searches per month, not going to happen. It has honestly been easier to use Anthropic to help me build my own search index & crawling infrastructure than deal with Google.
> It has honestly been easier to use Anthropic to help me build my own search index & crawling infrastructure than deal with Google.
That's great to hear! FWIW I agree with you that it's harder for independent professionals to get started on Google dev services than others, but the Search API is not developer service and was never intended to be.
Brave, Mojeek & Marginalia and EU Search Perspective have certainly been much friendlier to deal with.
I think that's actually a very interesting insight that would be helpful for PMs on GCloud to take note of. As a single founder, setting up Google Cloud, it's like they start out by assuming you're bigco, forcing (I assume most) of their users into a arduous process of removing components they don't need.
Google AI Studio is one of Google's solutions to this problem, but in typical Google fashion, it's bolted-on without any clear connection in the ecosystem. If you're also using GCloud, it's hard to remember it's even there.
OpenAI's platform, by contrast, is streamlined, easy to use. With Google, I feel like I need to wade through the documentation first before even using the darn thing.
My impression of anything enterprise (big or small) related to Google is that they just don't care / value solving it.
Which is sad, because "How does an enterprise customer pay for X?" is a non-trivial and incredibly important UX problem.
They have great technical solutions, but these are hamstrung by a frankly amateur understanding of how companies (startups to Fortune 500s) need to sign up, pay for, and track things.
From an outside perspective, one of the biggest gaps seems to be that internal Google product teams don't have to dogfood the full GCP et al. project/org experience. They get prebuilt billing structures (or just get to avoid them with internal cross-billing).
---
If I could waive a magic wand, Google would appoint an "Enterprise Czar", reporting directly to Pichai, who is a non-technical, retired founder / CEO.
That person would have one job: try to sign up, run (their team), and budget strategic Google initiatives (like AI) as a blind external party.
They would then deliver continuous reports to Pichai about how hard / easy this is.
Because potential Google customers don't give a shit if it's this internal team or that internal team's responsibility for integrating New Product X into GCP billing.
They care that the experience is terrible, filled with friction, and often flat out doesn't work.
I wanted a typical dev/qa/prod with medium specced boxes.
I was denied for quota, with an esoteric process for review.
I'd just made a case for deploying to GCP over AWD so got a bit of egg on my face. Went over and had it done on AWS in a few minutes.
A couple days later, the Google product team contacted me. I told them what happened.
It got escalated, and I ended up on a call with like 5 or 6 people from Google, some very senior. I told them what happened.
They made very concerned sounding noises and told me how this was a product failure on their part, how they'd get it corrected, etc... and they'd fixed my account so I could now make the machines. Of course, I was already deployed to AWS at that point.
That company grew and ended up with a pretty big cloud spend eventually. Google totally missed it.
I was at a new startup a few years later and decided to deploy to GCP.
Denied for quota.
The fact that there's two ways to get keys is also very confusing.
Even if you have GCP projects, I'd still recommend the AI studio UI - easier to figure out. Also, you can easily see the free tier in AI studio and just use your API key from there.
Because whatever internal team owns AI Studio fought for approvals to do so and GCP didn't?
But ironically my experience is that Codex/Claude navigate GCP better than Gemini.
Isn’t this the insecure thing that gives all your Google API keys access to Gemini, even those that were intended to be semi public (eg maps API keys embedded in websites or apps)
“Failed to create project, The request is suspicious. Please try again” or “ You do not have permission to create a key in this project”. You can then navigate multiple screens in GCP to make it work but it’s a hassle compared to any other provider (OAI/Ant/OpenRouter or any of the Chinese labs).
It was literally two clicks, and didn't even leave the page: the dialog asked to create a project and type in a name, I did that, clicked submit and then it was selected as the default project. One more click and I had the free tier API key.
Not saying you didn't have that experience at the time, but personally I have had zero issues with AI Studio and consider it the most dead simple/fastest dev dashboard to get started compared to the others like OpenAI/Anthropic (thanks to Google's free tier that lets you skip billing setup annoyances just to play around with Gemini).
Despite what HN threads (that are also frequently confused and talking about GCP instead) portray as universal/widespread issues or the process being complex and time consuming somehow.
Using Google products in general is an effing nightmare as soon as you have to give them money.
The one thing you want in a business is to remove friction when people want to give you money, a concept Google has never been able to understand.
Spending money via Google Pay on Android is extremely easy, Google does know how to accept customer's money (in the consumer space)
Eventually I gave up and run a few hundred million tokens (edit a few billion) through openrouter.ai using Gemini Flash 1.5 to Flash 2.5
Every since price increases on Flash 3.0 I've stopped using Gemini, too expensive for basic classification, sentiment detection, ocr etc.
As other posters said Google assumes you are some bigcorp trying to use their products. The Vertex versus AI studio confusion did not help.
I do pay for OpenAI, Anthropic, and ElevenLabs keys.
I know this is probably a pretty small edge case, but it is a bit frustrating. Any other provider lets you sign up with an email and give them a payment processor/card, but because google wants me to only use their unified workspace for signing up, I'm completely locked out now.
Sorry but it's not worth waking up with a 100k$ bill, fix your platform first.
Biggest traffic day of the decade and our site was down because of google.
"Experimental: The feature is experimental and limited in scope. You are subject to overages for around a 10 minute latency period."
Why in 2026 can't Google do a database lookup in real time? This is so ridiculous.
The primary reason for me has been that Google autonomously decides to downgrade usage tiers and then upgrade them again - and does this incorrectly.
Over the last week itself, in the span of two days, our account for first downgraded and then upgraded. This is despite matching the criteria to remain at the tier we operate at throughout.
Google Support (when you finally get to a human) has accepted that these are potentially bugs, but the first time it happened, we were rate limited so severely for ~4 hours that I find it really difficult to continue trusting Google.
By the way - I love Gemini's personality . it's phenomenal to work with
The reason i want my own harness - is the custom tools that i provide vs the low tier tools that come with the custom harnesses.
You'll would really benefit , if we could use Gemini in our own harness and not be forced to use it via agy . i've tried using gemini to circumvent - but not been successful.
If you see this - please reply here
Now following up with Google support team without luck to find the logs. Prompts send to the model and the responses including the thinking was available in the ai studio. But it’s unclear where to find the same in console.
To make matters worse there is vertex api and rebranded to Gemini something and making it very confusing.
Google being Google, their models tend to be better at finding, organizing and presenting information, from my experience.
I’d be very willing to try it out as an API, but it’s far too complicated to set up payment, and I don’t want to risk taking a wrong step and being locked out of other Google services. So Anthropic and Mistral get my money instead.
A single full wafer likely can run qwen3.6-27b alone. But won't be enough to run bigger models, which are pretty much all popular models.
1. HFT doing ass-simple arbitrage where only latency matters 2. More sophisticated slower trading taking in deeper signals
Those are two points along a continuum. If you are reacting to an earnings announcement by having an LLM read the earnings release and listen to the call, getting the results a few seconds earlier lets you get your trade in a few seconds earlier. Just because "not HFT" doesn't mean "completely latency insensitive".
I'm not sure what event-based traders are doing now, but back in the day NLP sentiment analysis was all the rage, so I'm assuming they've now incorporated LLMs too.
Not sure how much the benchmarks can be trusted though: https://www.reddit.com/r/GoogleGeminiAI/comments/1vbq5vf/com...
Maybe things there have improved some, but when I was looking it was a huge runaround.
Of course, we just used OpenRouter for testing and never touched a Gemini model anymore.
You reach for it every time you do a Google search
[my self-important Kagi shtick awakens, pokes at it's restraints]
What does this mean? Anybody can get an API key