Gemini API File Search is now multimodal
blog.google
blog.google
Everyone thought Google was pulling ahead with Gemini 3. For a minute there they had the best language model, image model, AND video model in the world. But it's like they decided to pull over for a nap while OpenAI and Anthropic flew by.
How, exactly, are they currently conquering the enterprise world with their models? What do you think Anthropic is doing?
Their latest proper model is a year old, they have no moat, no enterprise commitment.
Your comment would make sense if they would have actual success in the enterprise market and would have actual products in that area, but they don’t.
They had a brief sprint, caught up, and then dropped the ball again.
Their only current moat is their TPUs, and the fact that
1. The whole (successful) LLM world is screaming for capacity
2. They have excess capacity to rent out, just like Grok
Tells everything.
What's a "proper model?" Gemini 3.1 Pro was released 3 months ago. Gemini Robotics 1.6 was released a month ago. And Google is vertically integrated, they aren't just selling tokens, they are selling Taxi rides with Waymo. AI is a lot more than LLMs and Google is doing a lot more than LLMs.
I didn't say they were conquering the enterprise world. I said they are better positioned for the work that will be profitable in the future. Winning will mean being "good enough" for things like routine interactions with customers at the lowest cost to the business, and having customers fine tune your models using your hardware.
> What do you think Anthropic is doing?
Aside from being arrogant jerks that don't care about pissing off their customers, they're positioning themselves as the highest price provider for the highest end work. There will be a market for that, and maybe Anthropic will survive, but Google looks to me like they have a shot at being the profitable AI company.
I've mostly been doing reverse engineering with Codex, mostly related to games, but not once has the "training data cut-off date" been in the way, the most useful part comes from handing it a binary/directory and letting it prod it until it finds the answer you're looking for, I don't even have web search enabled and sometimes it might take 30-40 minutes for it to find the answer, but I never saw it be unable to find the answer because it's training data was a couple of years old.
It just feels like many google products really, they are capable of really amazing things, it's just that nobody there seem to care. I would guess they are likely optimizing more for internal use than their vast userbase.
So I put together a plan for refactoring it, step by step, with tests, etc. After literally 8 solid days of fighting with Gemini 3 Pro, I still couldn't pull it off.
I gave GPT 5.5 a chance with the same prompt, plans, and repo. I'm not sure how long it took, but when I checked in on it a few hours later it was done. All tests passed, everything exactly how I'd asked, and better (it made some improvements).
Surely she has seen Gemini in Google search but even her use of that is plummeting.
Google has so much revenue that they’ll be around for a long time. But I feel they are fumbling the opportunity with AI. Even in corporate, where we have Gemini. The conversation is fully around Claude. No one talks about Gemini.
Reports of the death of Google Search have been greatly exaggerated.
If you believe all the reports on HN about everyone's non-technical wives and grandmas, you'd have a hard time explaining the all-time highs in global usage and revenue from Google Search.
I agree with you that Claude 4.7 Opus is better than Gemini 3.1 Pro, but it's also a lot more expensive.
For my applications, I can't find better price-performance than Gemini 3.0 Flash. And it hasn't even been upgraded to 3.1 yet.
I suspect Google's target is price-performance and not just raw performance, which is how they can serve LLM responses at Google Search scale and still set an all-time record for quarterly earnings of any public company ever.
Frontier model capabilities leapfrog each other every few months, and Google I/O is in ten days, so I expect the leaderboard will change again soon.
OpenAI and Anthropic are going to get crushed long-term, and their investors are going to take a horrendous haircut.
On the other hand, Google and Microsoft already have the users (and lock-in). They just need to funnel them into Gemini and CoPilot.
Everything you said is wrong.
- DeepSeek is on V4 now, R1 is ancient history
- The models are open source open weights, which mean you can inspect what the models do and you you can choose US or EU infra providers
- DeepSeek literally has a Claude-compatible endpoint
Please don't comment and confuse other users on topics you know nothing about. Study. Then speak.
- Sure, you can use US infra providers. Together.ai is a good US provider but then it's 15X more expensive than DeepSeek's Chinese-subsidized pricing. It's really not that attractive at that price point. Anthropic and OpenAI are focused on larger models, but Grok 4.3[1] is smarter and significantly faster + cheaper than DS4[2] and by a wide margin.
- DeepSeek has a Claude-compatible messages API, but that's trivial. Anthropic has a massive API platform with things like Sessions, Files, and Agents[3]. None of those are available on DeepSeek.
1. https://artificialanalysis.ai/models/grok-4-3
- V4 will definitely move markets, especially as Claude and OpenAI keep jacking up the prices more and more. But inertia exists. Give it time.
- Most US infra providers are ~5x more expensive than Chinese infra, not 15x. But yes you are right. It does erode the cost advantage significantly. Big asterisk is that V4 seems to have solid cache hit percentage, often in the high 90s.
- Grok (and Llama) always underperform relative to their benchmark and ranking results. Don't ask my why, but it's a persistent pattern me and colleagues have noticed. I'll give them another try though, more competition is better.
- DeepSeek themselves have specifically said they prefer developing high performance models that you plug into other tooling, including Claude's. Regardless, I think it's unrealistic to expect DeepSeek to offer 1:1 suite compatibility with Claude or OpenAI. You wouldn't expect that from OpenAI <> Claude either.
Truly the main moat that OAI/Anthropic have is being 6 months months ahead of the competition in performance, which might be indefinite if the competition is just distilling their models (China) or takes many months between releases (Google).
Once you look passed the frontier of performance, it's just a race to the bottom on inference costs because there's at least 5 companies with equivalent open models at that level.
1. use Google App Scripts to duplicate all the convos saved in your google drive to another folder, and append each ".txt". Mine runs every day.
2. enable API in google cloud console, install google-api-python-client on desktop,then created a workflow in Alfred to search from desktop.
Now I can search my AI Studio convos + content just as fast as any other file on my desktop : )
You’d think this would be fairly obvious for Google to do, but it’s probably an organizational problem rather than a technical one.
For all intents and purposes Google Gemini is a totally separate company from Google search.
Teams will cross collaborate, but they have to be for specific projects with specific people.
Any app with this behind the scenes is a non-starter for me.
And anyone think that all those folks ditching Win11 will be going for or recommending any app built on this?
How much would you pay to have this yours forever, running locally, GDPR and HIPaa compliant, without the headache of privacy or subscriptions.
That´s what we offer with HugstonOne and we did it before Google. Multimodal, Lighting fast RAG, terabytes not kilobytes only :)
All you need is a 32gb ram laptop and HugstonOne, not a rocket science.