ARM's "Blackhawk" CPU Is a Plan to Have the Best Smartphone CPU Core
moorinsightsstrategy.com
moorinsightsstrategy.com
I feel like they are conflating 2 unrelated things?
There is "CPU inference" which is indeed pervasive and huge, and there is "generative ai llm inference" which is absolutely anemic on CPU, especially outside of tiny-context toy model demos, in spite of the popularity of llama.cpp and other cpu platforms .
I would hope smartphone users aren't running anything but extremely tiny llms on CPU. The IGP is not theoretically, but currently massively more efficient than any 2024 CPU could be.
Just to clarify: llama.cpp has also added some GPU acceleration over time.
Personally I would say "borderline-acceptable performance" with a 7B parameter model at 4-bit quantisation is nowhere near "great" or "supremacy". You can fit that on ~7GB of RAM, using your grandpa's 2016-era nvidia 1080.
To take the ARM-CPU-LLM-performance crown from Apple you have to beat a fully equipped M2 Max, which provides 96GB of memory with 400GB/s of bandwidth. That's enough to run 70B parameter models in 4-bit quantisation, and it can go toe-to-toe with a dual-nvidia-4090 system that draws a kilowatt of power.
Once the context grows, even M2 Max CPU alone really starts to chug.
However, minor tweaks to the design of GPU's could probably make them 5x quicker at directly processing quantized data (a 4 bit multiplier is so small you can afford to do thousands of them per clock cycle in a tiny silicon area), and I'm sure the next generation of GPU's will have that.
I don't really care about my phone being faster, but I would always take better battery life.
I’ve also noticed that there are emerging AI applications that cannot be run on-device by today’s smartphones. I would very much like an LLM Siri that doesn’t need a network connection.
Even the best 10nm processor can’t physically compete with a 3nm one on efficiency.
It will be interesting to see what happens with the Intel 18A process and who builds chips with them. Intel could also rapidly regain the performance crown if these are good processes too.
It never happens on same web app with 5 year old Samsung...
A web app though should be a situation where the memory usage will be roughly equivalent between the two platforms, so that seems reasonable.
If you want something more similar I would suggest the standard "Samsung galaxy S23" which has 8Gb compared to the IPhones 6Gb
2017 phone does NOT run out of memory, but 2023 phone model RUNS out of memory?
Newer phone is worse.
Their point wasn’t “iPhones have more RAM than comparable Android devices” (as you say, that comparison is entirely unreasonable based on those devices), it was “this is the RAM on the devices where I observed this behaviour”. If the iPhone in question has more RAM than the Android, yet still OOMs where the Android doesn’t, that’s a pretty strong argument that iOS isn’t magically thriftier than Android at managing memory.
If 6GB RAM iPhone runs out of memory, then how can it be better at memory management than Samsung phone with 3GB RAM, which does not run out of memory.
My problem is 16MB (8000*5000px) images in web browser which iPhone cannot handle. Most powerful phone every year, but 16MB is too much.
> Unfortunately the memory limit on iOS Safari is rather low. It’s limited to 384MB on version 15, it’s lower on earlier versions, and it’s probably device specific as well.
> As I said it depends on the device as well but it's always between 200-400MB
Its a decision, not a hardware limitation. Helps keep web bloat in check, but it’s a worse experience for the user. I’m sure Apple’s happy to suggest a native app without such limits tho
But no problem for very old low end android phone or IE11 in 10 year old laptop...
Even in terms of OS features, hardly anything that makes me feel like updating to the latest Android version, other than Google finally acknowledging Java isn't going anywhere and started following up on language updates (Android 12 onwards).
Naturally something that OEMs aren't that happy about.
I also bought a 120€ tablet with a Helio G99 CPU and it's so much faster than my old Galaxy Tab S3 and Lenovo Tab4 8 Plus.
To me it feels like we have reached the point where these devices are fast enough like it was when laptops matured, where performance was no longer an issue and upgrades were made when the device no longer worked or Microsoft made it obsolete due to Windows 11 not being allowed to run on it.
Oryon is based on the IP Qualcomm gained from the Acquisition of Nuvia, who developed this IP based on a very narrow license they got from ARM to develop for the server market.
Qualcomm is trying to transfer everything Nuvia has developed under that specific narrow ARM-license to the broad license of Qualcomm and use it for "powering flagship smartphones, next-generation laptops, and digital cockpits, as well as Advanced Driver Assistance Systems, extended reality and infrastructure networking solutions"
ARM has filed a lawsuit that this was never in scope of the license of Nuvia, the court-ruling is still pending on that one...
The related court-filing is worth a read: https://s3.documentcloud.org/documents/22273195/arm-v-qualco...
Which would probably lead to it being integrated into a nextgen Blackhawk architecture, merging both paths and making Qualcomm a licensee of Blackhawk.
But hey, I'm no lawyer ;)
They are staying even / besting apple in multicore through adding more cores than apple does.
The Cortex X4 Running at 3.3.Ghz on a N4 gets GB6 ~2250. #681.8/Ghz
The Apple A17 Running at 3.8 Ghz on a N3 gets GB6 ~2950. #776.3/Ghz
On a clock per clock basis, A17 is only about 14% faster. Consider X4 had 15% IPC uplift and resulted in real world 11% performance improvement on GB6. And they are claiming X5 would have the largest YoY IPC uplift, which I think we could consider to be 15%+ or 20%, a Cortex X5 would have similar if not slightly better than A17 performance on a Clock to Clock basis.
And it would be good enough for Microsoft / ARM PC.
Another thing I'd like to see is perf/watt.
You may want to have a word with AMD and Intel about that.
>many CPUs in the past could match Apple in clock-to-clock performance.
Are we talking about synthetic benchmarks or real world work loads.
Synthetic. Something like Geekbench.
It doesn't scale well beyond certain clock speed, which is why extreme overclocking dont get you a linear performance boost. But at the ranges mobile SoC it is about as linear as it gets. i.e Clocking your SoC from 3Ghz to 3.3Ghz will get somewhere around ~10% improvement in most workload.
>and target clock speed is a huge part of the CPU
Target clock speed is a huge part of considering for power usage.
Yes, target clock speed is a big part in considering power usage, but also fab process, memory architecture, etc. Adjusting for clock speed makes no sense, though adjusting for power consumption might.
And higher clock speed doesn't proportionally improve either real world metrics or benchmark results, so "benchmark score divided by clock speed" is a useless metric.
https://www.youtube.com/watch?v=iSCTlB1dhO0
The CPU peaked out at 14 Watts in multicore Geekbench. That's close to the peak CPU power consumption of the entire M1 chip in devices many times larger than an iPhone.
GeekerWan had it throttling 200-300MHz when simply running specInt/specFP. It essentially throttles down to the same speed of the iPhone 14 at slightly higher wattage.
For mobile devices, real-world peak CPU performance hasn't gotten much better than my aging iPhone 12 because most of the extra performance has come at the expense of heat/power.
But I am glad on HN people will fill the gap.