Google is silently winning the AI race.
Google is silently winning the AI race.
What I always wonder in these kinds of cases is: What makes you confident the AI actually did a good job since presumably you haven't looked at the thousands of client data yourself?
For all you know it made up 50% of the result.
It's the same problem factories have: they produce a lot of parts, and it's very expensive to put a full operator or more on a machine to do 100% part inspection. And the machines aren't perfect, so we can't just trust that they work.
So starting in the 1920s Walter Shewhart and Edward Deming came up with Statistical Process Control. We accept the quality of the product produced based on the variance we see of samples, and how they measure against upper and lower control limits.
Based on that, we can estimate a "good parts rate" (which later got used in ideas like Six Sigma to describe the probability of bad parts being passed).
The software industry was built on determinism, but now software engineers will need to learn the statistical methods created by engineers who have forever lived in the stochastic world of making physical products.
The main issue I had was it was surprisingly hard to get the model to consistently strip commas from dollar values, which broke the csv output I asked for. I gave up on prompt engineering it to perfection, and just looped around it with a regex check.
Otherwise, accuracy was extremely good and it surfaced a few errors in my spreadsheets over the years.
Everyone has a story of a csv formatting nightmare
It wasn't a one shot deal at all. I found the ambiguous modalities in the data and hand corrected examples to include in the prompt. After about 10 corrections and some exposition about the cases it seemed to misundestand, it got really good. Edit: not too different from a feedback loop with an intern ;)
They also remember what they did - if you spot one misunderstanding, there’s a chance they’ll be able to check all similar scenarios.
Comparing the mechanics of an LLM to human intelligence shows deep misunderstanding of one, the other, or both - if done in good faith of course.
BTW, not sure if you have experiences of delegating some works to human interns or new grads and being rewarded by disastrous results? I've done that multiple times and don't trust anyone too much. This is why we typically develop review processes, guardrails etc etc.
Oh yes I have ;)
Which is why I always explain the why behind the task.
30$ to get an view into data that would take at least x many hours of someone's time is actually super cheap, specially if the decision of that result is then to invest or not invest the x many hours to confirm it.
Can you share more about your strategy for "massive refactoring" with Gemini?
Like the steps in general for processing your codebase, and even your main goals for the refactoring.
For the bulk data processing I just used the python API and Jupyter notebooks to build things out, since it was a one-time effort.
I've spent a lot of time on prompts and tool-calls to get Flash models to reason and execute well. When I give the same context to stronger models like 4o or Gemini 2.5 Pro, it's able to get to the same answers in less steps but at higher token cost.
Which is to be expected: more guardrails for smaller, weaker models. But then it's a tradeoff; no easy way to pick which models to use.
Instead of SQL optimization, it's now model optimization.
accuracy | input price | output price
Gemini Flash 2.0 Lite: 67% | $0.075 | $0.30
Gemini Flash 2.0: 93% | $0.10 | $0.40
GPT-4.1-mini: 93% | $0.40 | $1.60
GPT-4.1-nano: 43% | $0.10 | $0.40
excited to to try out 2.5 flash
Historically speaking, if you had a 15% word error rate in speech recognition, it would generally be considered useful. 7% would be performing well, and <5% would be near the top of the market.
Typically, your error rate just needs to be below the usefulness threshold and in many cases the cost of errors is pretty small.
That's very different from just knowing the aggregate error rate.
Humans make a ton of errors as well. I didn't even notice how many I was making here until I started counting it. AI is super useful to just write get a first draft out, not for the final work.
It’s not surprising. What was surprising honestly was how they were caught off guard by OpenAI. It feels like in 2022 just about all the big players had a GPT-3 level system in the works internally, but SamA and co. knew they had a winning hand at the time, and just showed their cards first.
Gemini is what sucks from a marketing perspective. Generic-ass name.
Ask anyone outside the tech bubble about "Gemini" though. You'll get astrology.
I still think they'd have taken off more if they'd given it a catchy name from the start and made the interface a bit more consumer friendly.
The other AIs I have shown the same diagram to, have all struggled to make sense of it.
Yep, I agree! This convinced me: https://news.ycombinator.com/item?id=43661235
Not only in benchmarks[0], but in my own production usage.
They're getting their ass kicked in court though, which might be making them much less aggressive than they would be otherwise, or at least quieter about it.
Never count out the possibility of a dark horse competitor ripping the sod right out from under
Still impressive but would really put a cap on expectations for them.
It’s not clear to me what either the “race” or “winning” is.
I use ChatGPT for 99% of my personal and professional use. I’ve just gotten used to the interface and quirks. It’s a good consumer product that I like to pay $20/month for and use. My work doesn’t require much in the way of monthly tokens but I just pay for the OpenAI API and use that.
Is that winning? Becoming the de facto “AI” tool for consumers?
Or is the race to become what’s used by developers inside of apps and software?
The race isn’t to have the best model (I don’t think) because it seems like the 3rd best model is very very good for many people’s uses.
That is what we keep hearing here...The last Gemini I cancelled the account, and can't help notice the new one they are offering for free...
It seems shockingly good and I've watched it get much better up to 2.5 Pro.
As a consumer, I also really miss the Advanced voice mode of ChatGPT, which is the most transformative tech in my daily life. It's the only frontier model with true audio-to-audio.
Its more so that almost every company is running a classifier on their web chat's output.
It isn't actually the model refusing, but rather if the classifier hits a threshold, it'll swap the model's out with "Sorry, let's talk about something else."
This is most apparent with DeepSeek. If you use their web chat with V3 and then jailbreak it, you'll get uncensored output but it is then swapped with "Let's talk about something else" halfway through the output. And if you ask the model, it has no idea its previous output got swapped and you can even ask it build on its previous answer. But if you use the API, you can push it pretty far with a simple jailbreak.
These classifiers are virtually always ran on a separate track, meaning you cannot jailbreak them.
If you use an API, you only have to deal with the inherent training data bias, neutering by tuning and neutering by pre-prompt. The last two are, depending on the model, fairly trivial to overcome.
I still think the first big AI company that has the guts to say "our LLM is like a pen and brush, what you write or draw with it is on you" and publishes a completely unneutered model will be the one to take a huge slice of marketshare. If I had to bet on anyone doing that, it would be xAI with Grok. And by not neutering it, the model will perform better in SFW tasks too.
You can turn off those, Google lets you decide how much it censors you can completely turn it off.
It has separate sliders for sexually explicit, hate, dangerous and harassment. It is by far the best at this, since sometimes you want those refusals/filters.
And it replied "sure, here is a picture of a photo editing prompt:"
https://g.co/gemini/share/5e298e7d7613
It's like "baby's first AI". The only good thing about it is that it's free.
In my experience, anyone that describes LLMs using terms of actual human intelligence is bound to struggle using the tool.
Sometimes I wonder if these people enjoy feeling "smarter" when the LLM fails to give them what they want.
Learning how to "speak llm" will give you great results. There's loads of online resources that will teach you. Think of it like learning a new API.
WTF?