Microsoft Phi-3 Cookbook
github.com
github.com
https://github.com/microsoft/Phi-3CookBook/blob/main/md/04.F...
Scroll down to the end and the removed text is totally suspect. I wouldn't be to surprised if all of this was generated by an LLM then anything strange was edited by a human. Another reason not to leave everything to the LLM.
2024: the year of personal computers with neural processing units running small language models
How do NPU's work? Who builds them and how are they built? Are they capable of running a variety of SLM-like firmware?
And that's if the Snapdragon X Elite is actually on par with the Apple M3, like Microsoft claims.
Earlier X Elite benchmarks[1][2] showed that it was behind the M3, but hopefully Qualcomm have made some solid performance changes since then. Competition is good.
1. https://www.xda-developers.com/snapdragon-x-elite-benchmarks... Note: the link is comparing vs the M2
2. https://www.tomshardware.com/pc-components/cpus/early-snapdr...
MacBook Air 15,3" Apple M3 8-Core CPU & 10-Core GPU 24GB/1TB 2649 €
2. Where are you seeing the pricing for those Vivobook configs? I'm not able to pull up equivalent pricing in the US, so far.
That might be a pretty good deal, depending on the real world performance of the X Elite and the build quality of the ASUS laptop (I've had quite a few and it's been some great hits and some major misses.)
But the us price for 15" vivobook 16gb/1tb with elite (from announcement video) is 1299.
15" m3 with 16gb/1tb from apple.com is 1899
Either Surface line or Lenovo x line can be compared with Apple devices, anything else falls short.
- 13" or 15" screen
- Snapdragon X Elite (which still doesn't match M3 performance)
- 16 GB memory
- 512 GB SSD
Surface laptop price is $1399 (13") or $1499 (15"), which is $100 cheaper for the 13" vs the equivalent MBA, and $200 cheaper for the equivalent MBA 15". The Snapdragon X Plus, which is what it looks like you priced out, is quite inferior to the Elite and even the older Apple M1/M2 series chips.
This is all to say that neither is a bad buy, depending on your needs.
In multi-core from what I've seen X Plus is faster than base M3.
Single: 3088 Multi: 11595
Non-production Elite X From Anandtech[1]:
Single: 2800 Multi: 14400
Non-production Plus X From Anandtech[1]:
Single: 2425 Multi: 13100
Slower single perf for both Snapdragons. Decent 10% jump in multi over the M3. I am eager to see how the production units will pan out on benchmarks and sustained performance.
Qualcomm's been shady in the past with their so-called "Apple killer" chip benchmarks. I don't think these are "Apple killers", but I hope it pans out, for competition's sake.
1. https://www.anandtech.com/show/21364/qualcomm-intros-snapdra...
M4 iPad Pro: https://browser.geekbench.com/v6/cpu/6036233 S: ~3747 M: ~14740
Snapdragon X Elite: https://browser.geekbench.com/search?utf8=%E2%9C%93&q=snapdr... S: ~2400 M: ~14000
Note this is an M4 iPad Pro. I would imagine an M4 Mac would spec a little better, due to less heat/power constraints.
tl;dr: M4 stomps on the Elite in single-core with about 36% more performance. The baseline M4 and X Elite are about even on multi-core. X Elite has marginally better NPU TOPS performance if that's something that matters to you.
The Surface Pro has an OLED config now, which costs more, but the laptop and standard pro are still LCD.
X Plus and X Elite NPU hit 45 TOPS. M4 is 38 TOPS...
Also, as far I know, Elite does match M3. Actually, Qualcomm promise that even X Plus is able to match it, as they say Elite is better than M3.
We'll have to wait to see more benchmarks.
PS: There's 3 Elite series.
> Are we talking about AI/ML?
No, I was speaking to the usual single-core and multi-core benchmarks. Though you should buy based on what performance metrics you value most.While I have a basic understanding of TOPS metrics I don’t have a good enough understanding to speak much about it — especially when I’m not sure of what exact AI/ML workloads will be used on such platforms. I mean, how many tokens/sec and what wattage does that equate to?
> Also, as far I know, Elite does match M3.
I would disagree from what I’ve seen. > Actually, Qualcomm promise that even X Plus is able to match it, as they say Elite is better than M3.
The real world benchmarks will tell. I am super curious about the performance-per-watt which is something that really matters to me (heat, battery life, etc).Can the Elite outperform Apple’s chips? Can it do it at comparable wattage, or is it going to burn your lap doing so? Can’t wait to see.
The ones from today still have this issue.[2]
Beyond that, they've been pushing new ONNX features enabling LLMs via Phi for about a month now. The ONNX runtime that supports it still isn't out, much less the downstream integration of it into the iOS/Android runtimes. Heck, the Python package for it isn't supported anywhere but Windows.
It's absolutely wild to me that MS is pulling this stuff with ~0 discussion or reputation repercussions.
I'm a huge ONNX fan and bet a lot on it, it works great. It was clear to me about 4 months ago that Wintel's "AI PC" buildup meant "ONNX x newer Phi"
It is very frustrating to see an extremely late rush, propped up by potemkin blog posts that I have to waste time to find out are just straight up lying. Burnt a lot of goodwill that they worked hard to earn.
I am virtually certain that the new Windows AI features previewed about yesterday are going to land horribly if they actually try to land them this year.
[1] https://huggingface.co/microsoft/Phi-3-mini-4k-instruct-gguf... [2] https://x.com/jpohhhh/status/1793003272187351195
You have to build two in-development libraries, one from ToT, one of which is a dev branch to make it compile for iOS on a Mac temporarily.
The dev branch doesn't actually exist.
If you use the only branch by the author on the repo, it doesn't work.
The dev branch that doesn't work is a few commits on top of ToT from 2 months ago.
At the end of that non-existent road is a model that can't end messages properly, in MyThing.app that uses llama.cpp, or LM Studio, or Ollama, or MS cloud API.
I can't ship on that, and neither can anyone else.
I largely ignore benchmarks now, but on the other hand, while trying many models myself is easy for simple tests, really using a LLM for an application is a lot of work.
My use case is very simple: take 1000 word documents filled with two to three pages of information and pictures. And then output a set of requested information via prompting. Is there something off the shelf? Or do I have to make this?
Look at H2O.ai: https://github.com/h2oai/h2ogpt
The Phi-3 models are great though, especially the vision model has great potential for low latency applications (like robotics?)...