I wasn't going to bother trying those because I was pretty sure it wouldn't get any of them, but decided to give it an easy one (#4) and was impressed at the CoT.
Meanwhile, Google's newest 2.0 Flash model went 0 for 7.
1: https://metro.co.uk/2024/12/11/gchq-christmas-puzzle-2024-re...
Given the right prompt, though, I'm sure it could handle the 'find the corresponding letter from the landmarks to form an anagram' part. That's easier than most of the other problems.
You're saying the ultimate answer isn't 'PROTECTING THE UNITED KINGDOM'?
What do you mean? Is o1 not a single model?
There will probably be a 2.0 pro (which will be 4o/sonnet class) and maybe an ultra (o1(?)/Opus).
If they want a rematch, they'll need to bring their 'A' game next time, because o1-pro is crazy good.
However their on device TPUs lag behind the competition and Google still seem to struggle to move significant parts of Gemini to run on device as a result.
Of course, Gemini is provided as a subscription service as well so perhaps they’re not incentivized to move things locally.
I am curious if they’ll introduce something like Apple’s private cloud compute.
we need to separate inference and training - the real winners are those who have the training compute. you can always have other companies help with inference
I’m not saying on device will ever truly compete at quality, but I believe it’ll be good enough that most people don’t care to pay for cloud services.
inference basically does not matter, it is a commodity
training doesn’t matter if inference costs are high and people don’t pay for them
Training is amortized over each inference, so the cost of inference also needs to include the cost of training to break even unless made up elsewhere
Stack enough GPUs and any of them can run o1. Building a chip to infer LLMs is much easier than building a training chip.
Just because one cost dwarfs another does not mean that this is where the most marginal value from developing a better chip will be, especially if other people are just doing it for you. Google gets a good model, inference providers will be begging to be able to run it on their platform, or to just sell google their chips - and as I said, inference chips are much easier.
I don't know where did you get that price from but 1x RTX 3090 is $1,900. 16x is ~$30,000.
> The parts are expensive
Now that we invested ~$30k in GPUs, we only need to find a motherboard that can accommodate 16x pcie4 x16 GPUs, right? And we also need a CPU that can drive that many pcie4 x16 lanes?
Well, none of them exist, not even in the server parts sector let alone client commodity hardware. In any case, you'd need two CPUs so even with this imaginary motherboard we are already entering the server rack design space. And that costs 100's of thousands of $$$.
> but there is a competitive market in processors that can do LLM inference
Nothing but the smallest and smallish models. If that existed then why would you set yourself out building a 16x RTX 3090 machine?
Sorry, but you're just spitting out non-sense.
The second Apple comes out with strong on-device AI - and it very much looks like they will - Google will have to respond on Android. They can't just sit and pray that e.g. Samsung makes a competitive chip for this purpose.
Economically this fits the cloud much better.
If anything, I think the upcoming iOS AI update will bring them to a similar level as android/google.
While there is a chance that Apple might come out with a very sophisticate on-device model. The problem here is that they would only be able to compete with other on-device models. The magnitude of compute needed to keep pace with SOA models is not achievable on a single device. It will take many generations of Apple silicon in order to compete with the compute of existing datacenters.
Google also already has competitive silicon in this space with the Tensor series processors, which are being fabbed at Samsung plants today. There is no sitting and praying necessary on their part as they already compete.
Apple is a very distant competitor in the space of AI, and I see no reason to assume this will change, they are uniquely disadvantaged by several of the choices they made on their way to mobile supremacy. The only thing they currently have going for them is the development of their own ARM silicon which may give them the ability to compete with Google's TPU chips, but there is far more needed to be competitive here than the ability to avoid the Nvidia tax.
they’re a little bit less of a nobody than they used to be, but they’re basically a nobody when it comes to frontier research/scaling. and the best model matters way more than on-device which can always just be distilled later and find some random startup/chipco to do inference
Is it really that hard to imagine people have different viewpoints, and decisions than yourself without being painted as vapid, airheads?
The level of optimism for Apple AI capabilities on here is wrong. I can imagine people having wrong viewpoints, but it is wrong.
That may not be as big a disadvantage as you think.
Anthropic claim that they did not use any data from their users when they trained Claude 3.5 Sonnet.
About 7 years ago I trained GAN models to generate synthetic data, and it worked so well. The state of the art has increased a lot in 7 years, so Apple will be fine.
At best Synthetic data is a "slow follow" for training a model due to the need for human review, but a competitive model, it does not make.
Besides, did Anthropic and e.g. Mistral inherently have such troves of data to train on that Apple doesn't? For the last 6 months, Anthropic has had the SOTA model for the average production usecase.
> Google also already has competitive silicon in this space with the Tensor series processors, which are being fabbed at Samsung plants today. There is no sitting and praying necessary on their part as they already compete.
Intel had a much bigger advantage with x86, and look where we are now. I find it hard to believe that creating a good AI chip isn't a much smaller challenge than it was to do Apple Silicon. The upcoming SE uses their in-house 5G modem, another huge hardware achievement that no one else has been able to do.
With that in mind, how can you bet against Apple when it comes to designing chips at this point? It's not like Amazon et al aren't producing their own AI chips too. Let alone all of the startups like Cerebras. That indicates the moat and barriers are likely much lower than Apple Slicion or the 5G modem.
If I'm talking nonsense, do correct me.
I’m in the camp that this is the right call for consumers, instead of trying to compete on the large model side. They’ve yet to deliver on their full promise, but if they can, it’s the place where I think more of the industry will go (for consumers)
And regarding Google’s mobile tensor chips, they are infamously behind all other players in the market space for the same generation of processor. They don’t share the same advantages they do in the server space.
their models aren’t even that good. sorry apple fanboys but the talent isn’t there
Apple just isn’t very capable in this space, not sure what’s so hard to accept
I agree that the in-device inference market is not important yet.
inference hardware is a commodity in a way that training is not
And sure, poor reception will be an issue, but most people would still absolutely take a helpful remote assistant over a dumb local assistant.
And you don't exactly see people complaining that they can't run Google/YouTube/etc locally.
Most people are unlikely to buy the device for the AI features alone. It’s a value add to the device they’d buy anyway.
So you need the paid for option to be significantly better than the free one that comes with the device.
Your second sentence assumes the local one is dumb. What happens when local ones get better? Again how much better is the cloud one to compete on cost?
To your last sentence, it assumes data fetching from the cloud. Which is valid but a lot of data is local too. Are people really going to pay for what Google search is giving them for free?
Plus a lot of the "agentic" stuff is interaction with the outside world, connectivity is a must regardless.
Plus once you start with on device features you start limiting your development speed and flexibility.
That is currently Apple’s path with Apple Intelligence for example.
It has no world model. It doesn't know truth any more than it knows bullshit just a statistical relationship between words.
As the global human population increasingly urbanizes, it’ll become increasingly easy to blanket it with cell towers. Poor(er) regions of the world will increase reception more slowly, but they’re also more likely to have devices that don’t support on-device models.
Also, Gemini Flash is basically positioned as a free model, (nearly) free API, free in GUI, free in Search Results, Free in a variety of Google products, etc. No one will be paying for it.
Flash is free for api use at a low rate limit. Gemini as a whole is not free to Android users (free right now with subscription costs beyond a time period for advanced features) and isn’t free to Google without some monetary incentive. Hence why I also originally ask about private cloud compute alternatives with Google.
I see poor reception in both areas and only one has WiFi.
Pretty sure that's not doing any fancy on-device models!
That said, there was a popup today saying that assistant is now using Gemini, so I just enabled it to try. Could well have changed in the last week.
They've ceded the fast mover advantage, but with a massive installed base of Android devices, a team of experts who basically created the entire field, a huge hardware presence (that THEY own), massive legal expertise, existing content deals, and a suite of vertically integrated services, I feel like the game is theirs to lose at this point.
The only caution is regulation / anti-trust action, but with a Trump administration that seems far less likely.