Two new Gemini models, reduced 1.5 Pro pricing, increased rate limits, and more
developers.googleblog.com
developers.googleblog.com
For comparison, GPT-4o is currently $5/million input and $15/million output and Claude 3.5 Sonnet is $3/million input and $15/million output.
Gemini 1.5 Pro was already the cheapest of the frontier models and now it's even cheaper.
No wait, correction: That’s confusing: it lists 4o first and then lists gpt-4o-2024-08-06 as $2.50/$10.
we're planning on doing that default change next week (October 2nd). And you can get the lower prices now (and the structured outputs feature) by manually specify `gpt-4o-2024-08-06`
No, “I” can’t.
Open AI has always trickled out model access, putting their customers into “tiers” of access. I’m not sufficiently blessed by the great Sam to have immediate access.
On, and Azure Open AI especially likes to drag their feet both consistently, and also on a per-region basis.
I live in a “no model for you” region.
Open AI says: “Wait your turn, peasant” while claiming to be about democratising access.
Google and everyone else just gives access, no gatekeeping.
Well, Gemini Pro was delayed in Europe for many months. Same for Claude.
Because I do
Tough to disentangle the capex vs opex costs for them. If they did not have so many other revenue streams, potentially dicey as there are probably still many untapped performance optimizations.
I think that's one of those things competitors complain about that never actually happens (the raising prices part).
https://www.macrotrends.net/stocks/charts/WMT/walmart/net-pr...
Instead I think they are going after the "Android model". Recognize they might not be able to dethrone the leader who invented the space. Define yourself in the marketplace as the cheaper alternative. "Less good but almost as good." In the end, they hope to be one of a small number of surviving members of an valuable oligopoly.
I think this analysis is not in keeping with reality, and I doubt if that's their strategy.
Gemini is substantially cheaper to run (in consumer prices, and likely internally as well) than OpenAI's models. You might wonder, what's the value in this, if the model isn't leading? But cheaper inference could potentially be a killer edge when you can scale test-time compute for reasoning. Scaling test-time compute is, after all, what makes o1 so powerful. And this new Gemini doesn't expose that capability at all to the user, so it's comparing apples and oranges anyway.
DeepMind researchers have never been primarily about LLMs, but RL. If DM's (and OAI's) theory is correct--that you can use test-time compute to generate better results, and train on that--this is potentially a substantial edge for Google.
Would OpenAI even exist without Google publishing their research? The idea that Google is some kind of also-ran playing catch up here feels kind of wrong to me.
Sure OpenAI gave us the first productized chatbots, so in that sense they "invented the space," but it's not like Google were over there twiddling their thumbs - they just weren't exposing their models directly outside of Google.
I think we're past the point where any of these tech giants have some kind of moat (other than hardware, but you have to assume that Google is at least at parity with OpenAI/MS there).
[1] https://ai.google.dev/pricing
[2] https://cloud.google.com/vertex-ai/generative-ai/pricing
Google is the only one of the three that has its own data centers and custom inference hardware (TPU).
"We will continue to offer a suite of safety filters that developers may apply to Google’s models. For the models released today, the filters will not be applied by default so that developers can determine the configuration best suited for their use case."
Pricing and speed doesn't matter when your call fails because of "safety".
This is a query I did recently that got rejected for "safety" reasons:
Who are the current NFL starting QBs?
Controversial I know, I'm surprised I'd be willing to take the risk with submitting such a dangerous query to the model.
I don't recall exact prompt but it should be something close to that. I really wonder what filters they had about kink tracks and why? Do they have a problem with Beyond standard model searches /s.
I think I bumped into Anthropics once? And I know I hit ChatGPTs a few months back but I don't even remember what the issue was.
I hit Google's safety blocks at least a few times a week during the course of my regular work. It's actually crazy to me that they allowed someone to ship these restrictions.
They must either think they will win the market no matter the product quality or just not care about winning it.
Using their models in a medical setting is impossible. It refuses to describe scientific photos and will not summarize HCP discussions (texts) that it misinterprets.
NFL Starting Quarterbacks for the 2024 Season Note: Quarterback situations can change throughout the season due to injuries, trades, or poor performance. AFC • Baltimore Ravens: Lamar Jackson Buffalo Bills: Josh Allen
But then it gets wiped and you cannot see it even in the drafts. The text above is from the screenshot I managed to make before the response vanished.
This non-deterministic unpredictable behavior blended with poor “safety” policies is one of those major “dealbreakers” that pushes me back from trusting any existing LLMs.
For example, this prompt was apparently unsafe: "Summarize the conclusions of reputable econometric models that estimate the portion of import tariffs that are absorbed by the exporting nation or company, and what portion of import tariffs are passed through to the importing company or consumers in the importing nation. Distinguish between industrial commodities like steel and concrete from consumer products like apparel and electronics. Based on the evidence, estimate the portion of tariffs passed through to the importing company or nation for each type of product."
I can confirm that this prompt is no longer being filtered which is a huge win given these new lower token prices!
Unlike others here I really appreciate the gemini API, it's free and it works. I haven't done too many complicated things with it but I made a chatbot for the terminal, a forecasting agent (for metaculus challenge) and a yt-dlp auto namer of songs. The point for me isn't really how it compares to openAI/anthropic, it's a free API key and I wouldn't have made the above if I had to pay just to play around
Google is aware of the issue and it has been open on google's bug tracker since March 2024: https://issuetracker.google.com/issues/331677495
There is also discussion on GitHub: https://github.com/google-gemini/generative-ai-js/issues/138
It stems from something google added intentionally to prevent copyright material being returned verbatim (ala the NYT openai fiasco), so they dialled up the "recitation" control (the act of repeating training data—and maybe data they should not have legally trained on).
Here are some quotes from the bug tracker page:
> I got this error by just asking "Who is Google?"
> We're encountering recitation errors even with basic tutorials on application development. When bootstrapping a Spring Boot app, we're flagged for the pom.xml being too similar to some blog posts.
> This error is a deal breaker... It occurs hundreds of times a day for our users and massively degrades their UX.
I was ready to champion gemini use across my organization, and the recitation issue curbed any enthusiasm I had. It's opaque and Google has yet to suggest a mitigation.
Your comment is not hyperbole. It's a genuine expression of how angry many customers are.
Imagine if Anthropic or someone eventually release a Claude 3.5 but at like a whopping 10x its current speed.
Would be incredibly more useful and game changing than a slow o1 model that may or not be x percent smarter.
https://inference.cerebras.ai/
It's pretty fast, but my understanding is that it is still too expensive even accounting for the speed-up.
And yeah their cost is ridiculous, on the order for high 6 to low 7 figures per wafer. The rack alone looks several times more expensive than the 8x NVIDIA pods [1]
[1] https://web.archive.org/web/20230812020202/https://www.youtu...
I think there's another one but I can't remember the name of it.
Also a bit further out is https://spectrum.ieee.org/superconducting-computer
"Instead of the transistor, the basic element in superconducting logic is the Josephson-junction."
So even if Gemini sucks, they'll still win over execs being pushed to make a decision.
But as an enterprise customer, if you expect X, don't you get X into the contract?
It makes sense I need to be careful not to commit secrets to public repositories but now I have to avoid not only saving credentials into a file but even to paste them by accident into my editor?
I tried Gemini Code Assist and it was so bad by comparison that I turned it off within literally minutes. Too slow and inaccurate.
I also tried Codestral via the Continue extension and found it also to be slower and less useful than Copilot.
So I still haven't found anything better for completion than Copilot. I find long completions, e.g. writing complete functions, less useful in general, and get the most benefit from short, fast, accurate completions that save me typing, without trying to go too far in terms of predicting what I'm going to write next. Fast is the key - I'm a 185 wpm on Monkeytype, so the completion had better be super low latency otherwise I'll already have typed what I want by the time the suggestion appears. Copilot wins on the speed front by far.
I've also tried pretty much everything out there for writing algorithms and doing larger code refactorings, and answering questions, and find myself using Continue with Claude Sonnet, or just Sonnet or o1-preview via their native web interfaces, most of the time.
>> Yeah It's a good question. I think it's maybe less so of what makes it unique and more so the general trajectory of the trend that we're on.*
Disappointing.
Their docs are awful, they have multiple unusable SDK's and the API is flaky.
For example, I started bumping into "Recitation" errors - ie they issue a flat out refusal if your response resembles anything in the training data. There's a GitHub issue with hundreds of upvotes and they still haven't published formal guidance on preventing this. Good luck trying to use the 1M context window.
Everything is built the "Google" way. It's genuinely unusable unless you're a total masochist and want to completely lock yourself into the Google ecosystem.
The only thing they can compete on is price.
how so?
In fact, Google is most likely _the_ target for that clause in the license.
Google is now looking more like the Microsoft chart.
Also, this model shouldn't be compared to the CoT o1, I think. That is something different (also in price and speed).
It's not watching the video as far as I can tell.
https://www.tomsguide.com/ai/google-gemini-vs-openai-chatgpt
The free version of ChatGPT is 4o now, isn't it? So maybe Gemini has not gotten worse, but the free alternatives are now better? When I compare ChatGPT-4o with Gemini-Advanced (wich is 1.5 Pro, I believe) the latter is just so much worse.
You want guaranteed private data that won't be used for anything? Keep it on your own computer.
Setting up a local LLM isn't that hard, although I'd probably air gap anything truly sensitive. I like ollama, but it wouldn't surprise me if it's phoning home.
You can run Llama3 on prem, which eliminates that risk. I try to reduce reliance on 3rd party services when possible. I still have PTSD from Saucelabs constantly going down and my manager berating me over it.
I can flip your statement around. For the vast majority of use cases, LLAMA 3 can be hosted on prem and will have similar performance.
From the linked document, so save someone else a click:
> The terms in this "Paid Services" section apply solely to your use of paid Services ("Paid Services"), as opposed to any Services that are offered free of charge like direct interactions with Google AI Studio or unpaid quota in Gemini API ("Unpaid Services").The Gemini API (or Generative Language API) as documented on https://ai.google.dev uses https://ai.google.dev/gemini-api/terms for its terms. Paid usage, or usage from a UK/CH/EEA geolocated IP address will not be used for training.
Then there's Google Cloud's Vertex AI Generative AI offering, which has https://cloud.google.com/vertex-ai/generative-ai/docs/data-g.... Data is not used for training, and you can opt out of the 24 hour prompt cache to effectively be zero retention.
And then there's all the different consumer facing Gemini things. The chatbot at https://gemini.google.com/ (and the Gemini app) uses data for training by default: https://support.google.com/gemini/answer/13594961l, unless you pay for Gemini Enterprise as part of Gemini for Workspace.
Gemini in Chrome DevTools uses data for training (https://developer.chrome.com/docs/devtools/console/understan...).
Enterprise features like Gemini for Workspace (generative AI features in the office suite), Gemini for Google Cloud (generative AI features in GCP), Gemini Code Assist, Gemini in BigQuery/SecOps/etc do not use data for training.
> When you're using Paid Services, Google doesn't use your prompts (including associated system instructions, cached content, and files such as images, videos, or documents) or responses to improve our products, and will process your prompts and responses in accordance with the Data Processing Addendum for Products Where Google is a Data Processor. This data may be stored transiently or cached in any country in which Google or its agents maintain facilities.
Trying to build an actual product on top of it was an exercise in futility. Docs are flatly wrong, supposed features are vaporware (discovery engine querying, anybody?), and support is nonexistent. The only thing Google came back with was throwing more vendors at us and promising that bug fixes were "coming soon".
With all the funded engagements and credits they've handed out, it's at the point where Google is paying us to use Gemini and it's _still_ not worth the money.
This +999; I couldn't believe how inconsistent and wrong the docs were. Not only that, but once I got something successfully integrated, it worked for a few weeks then the API was changed, so I was back to square one. I gave it a half-hearted try to fix it but ultimately said 'never again'! Their offering would have to be overwhelmingly better than Anthropic and OpenAI for me to consider using Gemini again.
I had hopes of Google able to compete with Claude and OpenAI. But I don’t think that’s the case. Unless they come out with a product that’s 10x better in the next year or so I think they lost the AI race.
Potentially (depends if the EU cares)...
E.g. integration with Google search (instead of ChatGPT's Bing search), providing map data, android integration, etc...
Not sure if their models are the moat. But they definitely have an opportunity from the productization perspective.
But so does Microsoft.
It's incredible how bad it is. I've seen it claim I've never received mail from a certain person, while the email was open right next to the chat widget. I've seen it tell me to use the standard search tool, when that wasn't suitable for the query. I've literally never had it find anything that wouldn't have been easier to find with the regular search.
I mean, it's a really obvious thing for them to do, I'm genuinely confused why they released it like that.
I agree. Right now it's not very useful, but has the potential to be if they keep investing in it. Maybe.
I think Google, Microsoft, etc are all pressured to release something for fear of appearing to be behind the curve.
Apple is clearly taking the opposite approach re: speed to market.
I wonder if there isn't a deeper, more worrying (for Google) reason behind that - that AI is killing their margin.
Google has always been about delivering top notch services, and winning by being able to do that cheaper than the competition.
It's "in their DNA" - everyone knows that using links to a website as a quality signal was a really good idea in the early days of Google, but what's a little less well known is that the true stroke of genius was the algorithmic efficiency of PageRank.
Similarly for GMail. Remember when it launched, 1 GB of free storage was just completely out of every competitor's league?
It may just be that this recipe of being smarter than everyone on algorithms and on datacenter operations might just not work anymore in the age of modern machine learning.
All the hard scalability stuff, they've already done before. Gmail exists, the Gemini API exists.
If they're not getting it to work, there must be another reason. They just can't afford to provide it at a price point that users accept.
They announced a price reduction but it "won't be available for a few days". By the time, the initial hype will be over and the consumer-use side of the opportunity to get new users will be lost in other news.