Google Gemma 4 Runs Natively on iPhone with Full Offline AI Inference
gizmoweek.com
gizmoweek.com
The pattern "It's not mere X — it's Y", occurs like 4 times in the text :v
I guess I found the millennial. I haven't seen that in so long!
https://old.reddit.com/r/ChatGPT/comments/13mft8s/apparently...
LLM output doesn't have the variety of human output, since they operate in fixed fashion - statistical inference followed by formulaic sampling.
Additionally, the statistics used by LLMs are going be be similar across different LLMs since at scale its just "the statistics of the internet".
Human output has much more variety, partly because we're individuals with our own reading/writing histories (which we're drawing upon when writing), and partly because we're not so formulaic in the way we generate. Individuals have their own writing styles and vocabulary, and one can identify specific authors to a reasonable degree of accuracy based on this.
It's a bit like detecting cheating in a chess tournament. If an unusually high percentage of a player's moves are optimal computer moves, then there is a high likelihood that they were computer generated. Computers and humans don't pick moves in the same way, and humans don't have the computational power to always find "optimal" moves.
Similarly with the "AI detectors" used to detect if kids are using AI to write their homework essays, or to detect if blog posts are AI generated ... if an unusually high percentage of words are predictable by what came before (the way LLMs work), and if those statistics match that of an LLM, then there is an extremely high chance that it was written by an LLM.
Can you ever be 100% sure? Maybe not, but in reality human written text is never going to have such statistical regularity, and such an LLM statistical signature, that an AI detector gives it more than a 10-20% confidence of being AI, so when the detector says it's 80%+ confident something was AI generated, that effectively means 100%. There is of course also content that is part human part AI (human used LLM to fix up their writing), which may score somewhere in the middle.
This is the wrong thing to look at; your chess analogy is much stronger, the detection method similar (if you can figure out a prompt that generates something close to the content, it almost certainly isn't human origin).
But to why the thing I'm quoting doesn't work: If you took, say, web comic author Darren Gav Bleuel, put him in a sci-fi mass duplication incident make 950 million of him, and had them all talking and writing all over the internet, people would very quickly learn to recognise the style, which would have very little variety because they'd all be forks of the same person.
Indeed, LLMs are very good at presenting other styles than their defaults, better at this than most humans, and what gives away LLMs is that (1) very few people bother to ask them to act other than their defaults, and (2) all the different models, being trained in similar ways on similar data with similar architectures, are inherently similar to each other.
If you don't believe me, try it for yourself. Ask an AI to generate some text and give it to the AI detector below (paste your text, then click on scan). Now ask the AI to generate in a different style and see if it causes the detector to fail.
LLM is indeed just computer function that does stats. And our brains are just electro-chemistry that does stats. This is why stylometric analysis of human writing is a thing.
My previous experience with things such as you have linked to, is they used to be quite poor. I assume they're better since then, but then again so are the models.
Yes, but "better" means different things for each of these.
Detectors are trying to get better at distinguishing human from LLM-generated text.
LLMs are being improved to generate more useful (and benchmark maxxing) outputs, not to attempt to avoid detection.
LLMs are in fact explicitly trained to be as predictable as possible. The training goal is to minimize continuation prediction errors, which means they are in effect being trained to generate output where each word can be predicted by what came before it (which we can contrast to a human who tries to spice it up and keep it interesting by not being too predictable!).
RL post-training, which is especially used for computer code and math, is going to change this word-by-word predictability (detectability) a bit since the focus is now on a longer term goal rather than next word, but to some extent you could also view it as just steering/narrowing the output of the model towards that goal, not totally overriding the next-word statistics.
I don't know if there are AI detectors specifically trained to detect AI code rather than prose, but I'd expect that is more difficult to do, both because of the RL factor, and because computer code is so predictable in the first place - adhering to rigid syntax etc.
> Can you ever be 100% sure? Maybe not
The commenter I was replying to claimed exactly this. Their AI detector showed that the text was "100%" AI generated.
Compare to flipping a coin, counting heads vs tails, and trying to assess if it's a fair or biased coin. After 1000 flips if it's not close to 50/50 you would rightfully be suspect, and if it was 10/90 you should be almost certain it's biased. But you can never be 100% sure.
At this point relying on their judgement is beyond folly.
Sorry for making you snort and shake your head in amusement :D
My favorite: couldn't even prove the author is a real person. They all found no record!
The problem with the article is the complete lack of details. No benchmarks on the iPhone capable models. No details, whatsoever.
Human or LLM - the article is a whole lot of nothing.
"It's not just X – it's Y." Slop. "You're absolutely right!" Slop. "And this is key –" Slop. "This is a nuanced topic." Slop.
If this is the level of care that goes into news articles, then we're doomed. What will ultimately happen is that AI summarizes AI articles, which got summarized from another AI article, which got summarized from another AI article, .. and after enough rewriting all facts will be gone from articles. I don't care to read this slop, and I'm shocked people are so readily accepting this new state of affairs.
Running background processes might motivate the use of NPU more but don't exactly feel like a pressing need. Actively listen to you 24/7 and analyze the data isn't a usecase I'm eager to explore given the lack of control we have of our own devices.
The AI Edge Gallery app on Android (which is the officially recommended way to try out Gemma on phones) uses the GPU (lacks NPU support) even on first party Pixel phones. So it's less of "they didn't want to interface with Apple's proprietary tensor blocks" and more of that they just didn't give a f in general. A truly baffling decision.
Maybe not strictly impossible, but ANE was designed with an earlier, pre-LLM style of ML. Running LLMs on ANE (e.g. via Core ML) possible in theory, but the substantial model conversion and custom hardware tuning required makes for a high hurdle IRL. The LLM ecosystem standardized around CPU/GPU execution, and to date at least seems unwilling to devote resources to ANE. Even Apple's MLX framework has no ANE support. There are models ANE runs well, but LLMs do not seem to be among them.
[0] https://maderix.substack.com/p/inside-the-m4-apple-neural-en... [1] https://developer.apple.com/documentation/coreml
[1] https://news.ycombinator.com/item?id=43881692 [2] https://machinelearning.apple.com/research/neural-engine-tra...
It is an interesting area to explore, and yes,this is a tech demo. There is a long way to go to production-ready, but I am more optimistic now than a few months back (with Flash-MoE, DFlash, and some tricks I have).
> A new report says that Apple will replace Core ML with a modernized Core AI framework at WWDC, helping developers better leverage modern AI capabilities with their apps in iOS 27.
https://9to5mac.com/2026/03/01/apple-replacing-core-ml-with-...
https://github.com/blixt/pucky
It writes a single TypeScript file (I tried multiple files but embedded Gemma 4 is just not smart enough) and compiles the code with oxc.
You need to build it yourself in Xcode because this probably wouldn't survive the App Store review process. Once you run it, there are two starting points included (React Native and Three.js), the UX is a bit obscure but edge-swipe left/right to switch between views.
I think react native can be switched with swift
But it's more likely it's just walled garden + security theatre that'll keep them from allowing outside apps.
With a canonical source of truth, and set input/output expectations, the potential blast radius is quite small.
I don't think that's necessarily true. For instance, LinkedIn uses more memory than Gemma E2B inference does.
See Anywhere and Replit. Anywhere was the #1 or #2 app and was taken off the app store entirely before being put on and then taken off again.
Last I checked, Replit hasn't received an update on the iOS app store in over two months due to reviews denying them.
A kid playing Roblox can spend more than that in a good weekend.
I’m sure there are things on my phone it could replace (though I struggle to think of them) but there are plenty it can’t. My black magic camera app, web browsers, local send, libby/hoopla…
I can’t really think of any apps I use every day - or every week - that an LLM would replace. I’m not coding on my smartphone and aside from that an LLM is basically a more complex, somewhat inconsistent search engine experience right now for most people. Siri didn’t replace any of my apps, for instance. Why would chatGPT?
TL;DR: what apps would an LLM replace on my iPhone?
We should also remember that the effort of building and maintaining apps is dropping precipitously as LLMs get smarter, faster, and cheaper. OpenClaw signalled the direction in which we're heading, and within a year, Anthropic will no doubt have cheap and competent agents which can handle the maintenance autonomously in the background.
This is why SaaS valuations are getting hammered.
With people slowly abandoning dedicated computers they fully control (if we can even call windows/macOS that) and going towards mobile/tablet computing more every year, which is a far more locked down device run by companies that are becoming increasingly hostile to “side-loading,” I just don’t see how this can become reality.
Where can I get this amazing technology?
Apple doesn’t care about revenue from a random TODO app.
https://9to5mac.com/2026/03/01/apple-replacing-core-ml-with-...
Basically, a "toy" app to showcase where we are with coding agents on-device.
Come on folks, their IT hardware may be nice but supporting them is not worth it.
> 2.5.2 Apps should be self-contained in their bundles, and may not read or write data outside the designated container area, nor may they download, install, or execute code which introduces or changes features or functionality of the app, including other apps. Educational apps designed to teach, develop, or allow students to test executable code may, in limited circumstances, download code provided that such code is not used for other purposes. Such apps must make the source code provided by the app completely viewable and editable by the user.
Why is this related to local LLMs in app?
> execute code which introduces or changes features or functionality of the app,
https://www.macrumors.com/2026/03/30/apple-pulls-vibe-coding...
Seems pretty good to me!
This is not meant as a criticism, but people should be aware of their limitations.
The funny thing is that a lot of Google's internal training content uses an imaginary product "gShoe", and discusses the privacy implications of data that such a shoe might collect :D
At a glance, I see they do gather analytics about how much the app is used (model downloads, model invocations etc) without message content, pretty much just the model used.
I remember being excited when Apple got widgets because then I could add my 'Next Alarm time' to my home screen. Made my company work phone usable on trips.
I wonder when they are going to get NVIDIA cards or CUDA? Then they can actually run LLMs and not just trick people into buying it under the 30 year old idea of 'Unified Memory'.
They've had to be dragged kicking and screaming away from the NPU model only to admit that GPGPU tech was the right choice.
'Cool demo' -> Doesnt convert to tangible things.
Wont attempt to compete with companies better than them, but go their own route. "oh look it consumes low power!" (Things no one cared about).
They are the Nintendo of tech.
Threat found This web page may contain dangerous content that can provide remote access to an infected device, leak sensitive data from the device or harm the targeted device. Threat: JS/Agent.RDW trojan
What are the possibilities of an Android or iOS device where the OS is centered around a locally running LLM with an API for accessing it from apps, along with tools the LLM can call to access data from locally running apps? What’s the equivalent of the original Mac OS?
Do apps disappear and there’s just a running dialog with the LLM generating graphical displays as needed on demand?
Isn't the "edge" meant to be computing near the user, but not on their devices?
Can't wait until AI companies go from mimicking human thought to figuring how to licensing those thoughts. ;)
In a general sense, edge just means moving the computation to the user, rather than in a central cloud (although the two aren’t mutually exclusive, eg Cloudflare Workers)
For those that have lost their marbles: sure, people use words incorrectly, but that does mean we all have to use those words incorrectly.
In compute vernacular, "edge" means it's distributed in a way that the compute is close to the user (the "user" here is the device, not a person); "on device" means the compute is on the device. They do not mean the same thing.
qwen3-coder-next uses a lot less since it seems to only activate ~3B parameters at a time.
My guess is that this is still close to tech demo, and a lot of performance is left on the table.
The models are quite good. They aren't just a tech demo.
Apple is getting a base Gemini model (not a Gemma), and it will run on Apple private compute. Apple foundational models will remain the on device model
I want to test a hypothesis for "uploading" neural network knowledge to a user's brain, by a reaction-speed game.
You don't need a neural network. Traditional NLP is far better at this task. The keyword you're looking for is "phoenemizer"
I'm surprised traditional NLP being better than ML models for this task, can you point me to a benchmark analysis pointing out that non-neural Espeak-ng is better than ML models?
Also, I asked for a neural model for another reason as well, I still want semantic knowledge present, I want more than pronunciation, but before I use myself as a test subject, I want to make sure I get the proper pronunciation in case the highly speculative "uploading game" works... I don't want to early systematically mis-train myself on pronunciation...
Never paid an LLM provider and I have no reason to ever start.
[0] https://frame.work/products/desktop-diy-amd-aimax300/configu...
The only downside is that I suspect the Framework would be a decent bit quieter under load (not that this thing is abnormally loud). As well as you're limited to a single M.2 2230 internal SSD slot in this (I believe Micron recently launched a 4 TB model, but generally you'll max out at 2 TB without using an external enclosure).
I don't have anything against the Framework, I'm sure it's a great machine, but the Z13 is an incredible portable all-in-one device that can handle everything from general PC use to gaming to tablet/entertainment to LLMs & high perf.
I put my boards in mini itx rack mounts personally so framework is the only option.
Disappointing if you compare it to anything else from 2026, but fairly impressive for something that can run locally at an OK speed.
You need a relatively beefy phone to run this stuff on large amounts of text, though, and you can't have every app run it because your battery wouldn't last more than an hour.
I think the real use case for apps is more like going to be something like tiny, purpose-trained models, like the 270M models Google wants people to train and use: https://developers.googleblog.com/on-device-function-calling... With these things, you can set up somewhat intelligent situational automation without having to work out logic trees and edge cases beforehand.
It's a 100% replacement for free ChatGPT/Gemini.
Compared to the paid pro/thinking models... Gemma does have reasoning, and I have used the reasoning mode for some tax & legal/accounting advice recently as well as other misc problems. It's worked well for that, but I haven't tried any real difficult tasks. From what I've heard re. agentic coding, the open weight models are ~18-24 months behind Anthropic & Google's SOTA.
Qwen 3.5 122B-A10B should just fit into 128 GB with a Q4/5 and may be a bit smarter. There's apparently also a similar sized Gemma 4 model but they haven't released it yet, the 26B was the largest released.
I think this should be flagged.
The model itself works absolutely fine, though the iPhone thermal throttles at some point which really reduces the token generation speed. When I asked it to write me a business plan for a fish farm in the Nevada desert, it slowed down after a couple thousand tokens, whereas the Pixel seems to just keep going.