M2 Ultra can run 128 streams of Llama 2 7B in parallel
github.com
github.com
This high bandwidth is really a result of Apple having designed a unified memory architecture for the M1 and M2 chips. Typically on a laptop or desktop, the CPU and GPU have distinct memory systems: high-bandwidth (but relatively low-capacity) graphics memory, and relatively low-bandwidth (but high-capacity) CPU memory. Apple decided to simplify that and instead implemented a single high-bandwidth memory system shared by the CPU and GPU. The only downside is that such high-bandwidth memory had to be tightly integrated in the M2 package, so the maximum capacity is limited. For example whether you spend 5,600 USD (cheapest Mac Studio machine with M2 Utra and 192 GB) or $10k+ (maxed out Mac Pro), you will only ever get 192 GB RAM max. For that amount, a PC could get 1024 GB RAM (5× more!) But on the other hand, if your workload, like inference, doesn't need more than 192 GB, then that's great. Personally I think Apple made the right tradeoff here. 800 GB/s of memory bandwidth on a general purpose CPU, on a single socket, has never been done before (to my knowledge.)
https://www.techpowerup.com/gpu-specs/geforce-rtx-4090.c3889
The budget option is to go with a used 3090, which still has greater memory bandwidth than the M2 Ultra.
In terms of FP16 Flops on the GPU, you have:
M2 Ultra, 27 TFLOPS
RTX 3090, 35 TFLOPS
RTX 4090, 82 TFLOPS
https://www.cpu-monkey.com/en/igpu-apple_m2_ultra_76_core
https://www.techpowerup.com/gpu-specs/geforce-rtx-3090.c3622
https://www.techpowerup.com/gpu-specs/geforce-rtx-4090.c3889
Incredible what (even with limitations of a single $1000 computer, costing less than a single nVDA 4090) a desktop mac mini can accomplish.
----
For comparison: the hard drive in the M2 Mini is FASTER THAN A MACPRO5,1's RAM!!!
And to "win" this the total VRAM available at once is what make a difference, not really the bandwidth ; this is just because the task has been parallelized as much as possible. Even then, it required to optimize for the arch with parallelization and it is absolutely not cost competitive with the PC used as a reference. If you really need to maximize GPU VRAM in a single workstation (without going to a servers/cloud solutions) you could build a machine with multiple RTX A4000 SFF (1 slot, 20GB). It would get more expensive than the maxed out M2 Ultra but at this point the M2 Ultra lose so hard in FLOPS power that you really need to specifically look for situations where you would want more VRAM (up to 144GB available for the M2 Ultra GPUs vs 80GB for 4x1 slot card) but wouldn't want to run the model faster/longer in a dedicated server rack that could potentially have even more available VRAM (and be available/shared with other peoples).
Realistically NVidia knows how to put more RAM in their GPUs because it doesn't make sense to scale VRAM faster than computer power for most workload, you need to have a balance that make sense. As an analogy it is like coming up with a truck than can carry 150T at once but can only do so at 1/3rd the speed of regular trucks. In most case you actually gonna want to run 3 regular trucks even thought is going to be less efficient (it still gonna cost less and be faster overall) unless you really don't have a choice ; at this point you are in "special convoy" territory (like for wind turbine blades) and it's gonna cause lots of headaches on top of being slow and expensive.
Apple market their stuff as an incredible innovation when in fact not only it is irrelevant for most workload that are usually thrown at workstations (mobile or not) but I would argue that running the workloads where it would actually make a difference is a bad idea on a single user workstation. For most things that actually matters in a single user workstation/prosumer/enthusiast system, Apple Silicon lose quite hard especially when it comes to GPU performance : viewport performance, close to real-time 3D rendering (before sending to render farm for final detailed render), games, etc...
And this is the Ultra version of the chip, that is out of reach for most people (it makes look at the 4090 as not that overpriced, which is quite funny). If you go down to the M2 Max version, suddenly the bandwidth is 400GB/s and not only it is not impressive at all, it is even worse than an Intel A770M laptop GPU (512GB/s) while still having less raw power and costing way more. The more you go down in the Apple Silicon roster, the worse it gets. AS is not competitive at the high-end workstation level but it is absurdly overpriced at almost every level.
The reason they have this architecture (that isn't very good for most traditional computer application) isn't because they went out of their way to engineer something great. Nope. It is because they basically scaled up a mobile architecture that was like this from the get go (power and space constraint, plus no need to have that much RAM nor have it upgradeable). And this is only because Apple is currently run by a Scrooge who figured he could get even more money out their silicon division if they solds SKUs with binned parts and controlled the RAM supply/price.
If Apple had actually done useful engineering they would have figured out a way to scale the GPU/VRAM combo independently and a way to package/sell it efficiently. It makes no sense to scale VRAM past a certain point : why would you want to load a 3D model/view/whatever if you cannot compute it fast enough. As for the CPUs existing memory interfaces where fast enough for most things and the "benefit" is inexistent in most case. They went about it in the worse way possible with cost reduction above all approach while jacking up the price up to 11. This is the most lazy approach they could take and they even dumped all the unnecessary cost directly onto the consumer (low yield for big area chips and soldered RAM close to the chip from a lack a dedicated GPU SKUs). Even if the consumer want to absorb the cost he still get bad scaling and uncompetitive performance...
I just don't get how Apple get away with it and there are people like you falling for their marketing bullshit that is just a spin on what are actually weaknesses...
The wisdom is that cloud providers are better at infra than you, and that the economies of scale make it better to piggy back on what they’re doing, but… AWS is the most profitable part of Amazon for a reason. They’re overcharging you.
When you look at the cost of the hardware + hosting. Yes, it certainly looks and feels that way.
But if you've dealt with corporate IT, and had to deal with 3-6 month lead times on getting hardware, or politics to get your hands on hardware to get stuff done.
AWS is cheap. It gives you velocity.
If your company is large enough that it can offer the elasticity of resources that Amazon offers or even 1/4 of it... and you have an IT org that will let it happen. Yes, AWS is a waste.
But with AWS... when a project dies, you can wipe its costs out, people won't hold onto hardware so they have hardware for the next project, etc...
Trust me. I've been IT, I can spec and build rack systems. I am a software dev. And I've been a dev most all my career.
For 90%+ of orgs... they don't have the maturity and skills to handle that type of infra without substantially distracting from their primary business.
Also, maintaining servers is not hard at a proper data center. It is often more hands off than the migrations cloud providers force on their customers.
It’s not like process disappears just cause you’re not on your own hardware. Infra is still its own team with its own budgets poking and prodding at every damn turn for every little thing till rejecting your requests, you escalate and then have a 4 week battle over needing the space.
If you are paying the price of being on prem, which is really lack of ability to provision and de-provision infra quickly. There's little point to the cloud, unless you just have no infra to begin with (small companies).
I'm in a small firm now. I can't imagine having an approval process to spin up a few instances to run my tests and spin them down after. That'd be silly.
Most firms don't. Or don't have the skills.
Also, the cloud can help an IT project recover from errors. Let's say, I'm about to buy 500k of hardware to setup some storage. I get my requirements, I architect it, do my design work, and then buy the hardware. I have to over provision a bit because of reality and human error... But when I discover that the requirements, shift 2mo in my project, and I've already ordered the hardware... I may be hosed.
This isn't hypothetical, this is what happens. Things evolve and shift. The cloud allows for more agility. If your firm is large enough, or has its stuff together enough, go for it on-prem.
I've got 20+ years on prem.. I've seen it fail all over. I've seen cloud be a mess too. But if you told me to clean up one. I'll take the cloud.
It’s still a win for a lot of use cases and I still do it quite often, but the meme that it’s this “click and you’ve just hired the best ops team in the world to work for you” and so the 50-500% markup is actually a bargain is horseshit. A Bizon box in your living room fucks AWS up on flops/$ on most instance types and pays for itself in 30 days.
It is one of the best ops teams on Earth: but they’re working for you like the Google search team is working for the user.
Your application isn't going to magically become HA/DR. You still have to make it that way, from your application design/coding up through the deployment.
I mean, if you're not storing your session IDs in a data store that's reachable by all the nodes behind your load balancer then no amount of infrastructure is going to save you.
That’s a realistic scenario no matter whether you’re bare metal, building out your own cloud, or using someone else’s. No amount of AWS/GCP/Azure/et al marketing changes that.
Yes, you have to learn things to goto the cloud, and I won't say it is all roses, it ain't. But... AWS is less likely to fsck it up.
If you have the constant load to burn the flops 24x7x365... go for it. If you have the ops team to do it... go for it.
If you don't... take a bit of time and learn the cloud which is much easier than getting on-prem right.
Especially for smaller firms, this isn't even a close call IMHO.
I’m quite capable of setting up whole server stacks. I did it for years, but I stopped, some time ago, and consider myself to be, for want of a better word, incompetent at being a modern admin.
I think I’d screw the pooch, so I prefer that someone who does it every day, handle it.
But I write Swift code, every day, so I’m not incompetent at everything.
If I wanted to host a website, sure, I can build a server out of parts and negotiate with my ISP and get a business pipe and handle all caching and such. Or like I can pay a provider $5/mo and get better performance and reliability with no management overhead. Yeah, maybe over 5 years I'd save more money doing it myself... but it's not worth the time.
If I wanted to generate a photo or a dozen, or a few paragraphs of text, that's like a few cents worth of cloud AI. Maybe low single-digit dollars. Or I could spend thousands on fat GPUs or a Macbook, spend forever training it, and still end up with a sub-par result.
AWS is profitable not just because they're overcharging you but because they are providing a hugely useful service for millions of businesses that don't want to deal with that infrastructure themselves, any more than they'd want to manage their own plumbing or electrical grid or roads and bridges leading to their office. DIY makes sense if you're doing it as a hobby or if your scale is so big that you would incur significant savings to in-house it, but for millions of small and medium businesses, it's just not the most practical approach. Nothing wrong with that.
I mean, it's like saying development is such a lost art... why hire a dev if you can learn to code yourself? Sure, but not everyone wants to, can, or has time.
My employer has generated its own electricity and steam for decades.
For a small business - different story.
Still haven't started on the silverware or shoes yet.
I do agree with you though. If you are a non-tech company or a company that lacks the human resources you might as well go with the cloud.
Are we in the business of building, maintaining and operating <thing to build> or do we want to buy that as a service instead and focus on our actual core business?
There's more to the cost of building and operating than just the hard costs.
Retaining good modern IT talent is getting harder and harder - and I'm not even talking about salaries.. You need a whole department including strong leaders who can hire, train, and lead the right people, etc..
This is something most companies wouldn't even know where to start with.
But if you can motivate that same sysadmin to spend his skills on something more directly benefiting your company, then you should still buy it in.
Given how popular homelabs are, I don't think this would be too hard to find.
You can throw a bunch of boxes in a closet and it'll work. A surprisingly large amount of the early Internet was "a spare box under my desk."
The problems start when they become part of your critical path and you're on vacation and nobody knows WTF is happening.
I mean, it's a risk. If you're OK with that risk then go for it.
It's really about the politics of your office.
If everyone is OK with the idea that the box is in some closet somewhere that's fine. I've been part of a bunch of startups where we were running infrastructure on spare hardware. Sure it's not HA, but we didn't need it...or it was at least HA enough for what we needed.
And I think in most cases companies want to focus their employees and efforts on their core business, and if that doesn't include setting up and maintaining hardware in the long-term, then you don't build, you buy.
If your operations halt until some poor sysadmin has to drive to the colo, you are absolutely doing it wrong.
I remember twice switching from an in-house jenkins/teamcity/whatever type of CI to Azure devops and the thing I remember the most was how much longer it took a build to complete as well as the massively longer time downloading a build from Azure vs from within the office. Even when working from home the on-prem stuff was faster.
The thing is, the build/devops teams seem to be about the same size in both cases. It's just kind of worse in pretty much every case when we do CI in the cloud.
Notes
- My experiences are largely for game development so the build times and artifact sizes can be quite large.
- I've only ever had CI/CD experience with Azure, I've not tried other cloud providers
- Since this is game development and we're using CI downtime is more acceptable than other cases. That said, I don't remember much downtime when I was working as a build engineer. I have seen periods of 1-2 hours of downtime once in a blue moon but then again I've seen that with Azure. In both cases it wasn't so much the setup but a build script deployment issue.
Also being able to cool off in the rack room when it's a hot day is always a treat :)
Meanwhile a Mini cluster is literally a bunch of mini pcs in a rack, and idk if Apple even supports this kind of industrial use. While it's a quality product the Mini isn't really designed for the datacenter.
I think they know of it and tacitly approve of this use case, as evidenced by the Mac Mini having the same form factor for ages. They’re well aware that a lot of people use Minis (and Studios now) in data centers, and that the Mini footprint is sort of “standardized” at this point.
https://en.wikipedia.org/wiki/Xserve#Intel_Xserve
But since Apple discontinued Xserve and macOS Server, they seems like don't care about this business anymore.
Mac OS X Server was its own operating system originally. It was still the same core OS, but had a ton of additional servers built in. Non-exhaustively, they included IPSec VPN, email, calendaring, wiki, SMB and AFS file shares (including support to act as a Time Machine backup destination), LDAP, DNS, and software update caching before it came to macOS proper. The Server app released via the App Store was a shadow of Mac OS X Server.
These were quite popular in small professional offices like law firms.
(Not sure what differentiates the later model Mac Mini Servers from the regular Mac Minis, since Mac OS X Server just became a $19 App Store purchase, and optical drives were no longer a thing in Mac Minis)
They discontinued the Mac mini Server line in October 2014, which was still sold with two drives instead of one. Configurable to order with SSDs by that time.
Got three databases up and running too. It's a beast. I'd definitely consider self-hosting with a few Mac Minis, that would be fun and they're really cute, sleek devices too. I paid $650 for it and consider it a great deal. Definitely should've gotten it with more than 8gb of ram but I got it to try it out and haven't yet really needed to upgrade to a unit with more memory.
When running an ML workload, the Nvidea A100 has massive GPU compute resources, and a large amount of GPU local high bandwidth memory, so it's ideal, but is nowhere near low cost.
A consumer Ryzen chip is inexpensive, but lacks in both memory bandwidth and GPU resources.
The M2 Ultra has access to way more RAM than consumer GPUs, many times the memory bandwidth of a Ryzen (800GB/s vs Ryzen 7 1800X at 40 GB/s) with a large amount of local GPU resources.
Even stepping up to a Threadripper Pro would only get you a quarter of the memory bandwidth, and those aren't exactly cheap either.
The Apple Silicon is great at really low power work, but if you dial desktop or server GPU power limits down they also become quite efficient. The marginal cost of electricity is cheaper than buying more hardware, so nVidia and others run their parts deep into the diminishing returns part of the curve to maximize performance at the expense of power efficiency.
A top-of-the-line Mac Studio will give you 192 GB of unified RAM in less than $7000. Meanwhile a H100 from NVIDIA with 80 GB of VRAM will cost you like $30000...
192GiB RAM is enough to train or inference Falcon 180B in RAM at 8-bit resolution.
Did the math and assuming 100% util and equal performance (which is certainly not the case) payback on your Mac is 9 months...
4x NVIDIA A100 at lamda labs is $4.40 an hour and I really have not had an issue getting them.
https://www.reddit.com/r/LocalLLaMA/comments/16o4ka8/running...
OP implied that there were workloads where it out competes renting in terms of cost. Was hoping it was true for something than a single user interactive session (which can be done a lot cheaper)
You can run them at half the power usage and only lose a fraction of the performance - at least in gaming. Try for AI tasks.
Also note that the demo video is sped up to fit inside GitHub attachment limits. Your observed speed may vary. :)
However, for scripts that try to use hundreds/thousands of invocations to solve some problem (eg. "write me a whole book"), the parallelism will be great (but obviously the script has to be written with that in mind).
anywhere you find `FooBarBaz(blip, kap)` replace it with `new newThing(blip).bump(kap)`
I don't know how reliable it is, but it seems like if you can easily run this on commodity hardware it could totally replace most IDE refactoring tools, although obviously the IDE refactoring is more reliable, it seems like this could be made simple and flexible, and possibly just as reliable as IDEs.
But also it could enable some interesting things that you could never do with an IDE refactoring tool.
I'm more comfortable with that because I find it usually takes two or three prompts to get it right
e.g. A couple hours ago I prompted it to help me do a diff of two commits ignoring all white space, just to check if there were any other changes. The first response didn't ignore new newlines, the second one was a multiline script, the third response gave me what I actually wanted.
diff -w <(git show 0bb2c8579efe775de883e0182db48989bfa324f2:"path/to/file"|tr -d '\n') <(git show 6c71efc17497ad7c90b9c7b690075ec031c13c69:"path/to/file"|tr -d '\n')If we get closer to a either "AGI" (whatever the hell that is) or at least a reasonably useful AutoGen/BabyAGI-like system that become popular to use at home, those machines will be the only ones capable of running advanced LLMs without having to pay OpenAI, Microsoft, Amazon/AWS, etc inordinate sums of money to do what consumers will deem a utility some day.
No. They work well on the apple chips thanks to the integrated memory and the large size of the models. I know of no reason why an x86 chip could not be designed in a similar way if desired. IANAChipDesigner but I have worked for one of them.
This needs some backing. When M1 just got out people were claiming it is comparable to 3080, until they saw the performance difference.
> The M1 uses a 128-bit LPDDR4X SDRAM in a unified memory configuration shared by all the components of the processor.
I assume that includes the NPU, media engine, etc.
There's a Whisper (Encoder-Decoder) [2] implementation if you want to see it in practice. Shameless plug, but I have a repo [3] where I'm working on autoregressive text generation on the Neural Engine. I'm running gpt2-xl (1.5B params) locally with KV caching at 120ms/token (vs. 450ms without caching). Will push an update soon.
Without quantization you can't go much higher than 1.5B params on M1's Neural Engine. M2 seems to have a higher ceiling but I haven't measured. I'm optimistic (but have not tried) that the new runtime quantization added to CoreML this year will allow for larger (and maybe faster) models on both.
[1] Technically you should be able to use 1 input with an enumerated set of sizes but I haven't been able to get it to work on the Neural Engine. This would likely be even faster. [2] https://github.com/wangchou/whisper.coreml/ [3] https://github.com/smpanaro/more-ane-transformers/
That seems very slow compared to llama cpp?
What are real world use cases for 7B family of models? Is anyone using them for anything productive?
I'd have one use case for classification: user text (from a jira issue) mapped to the team responsible for the fix.
Can you share some tutorials? I only just managed to get this working on windows/cuda:
https://colab.research.google.com/drive/1vk8i01apaSp59GVV2yI...
It's been a royal pain to setup
You can use them for trivial nlp tasks ("between 0 and 1 how similar are these two sentences? Respond with an explanation.") and because it's a small model, you just run it 4 or 5 times and take an average pretty quickly.
Additionally, I can imagine companies investing and paying for the open source work to expand access to their licensed models. Use the same interface as people use LLAMA but upgrade to BetterModel, fully compatible.
Additionally, I could believe this is simply a build up to a future Acquihire, which is the most lucrative way to be hired.
For comparison, this is the actual title of the page, but do you think this would increase people's awareness about the fascinating fact I highlighted in the title?
llama : custom attention mask + parallel decoding + no context swaps #3228Something like "Amazing Llama 2 7B performance on M2 Ultra" would obviously fail that test, but the current title of "M2 Ultra can run 128 streams of Llama 2 7B in parallel" seems to follow the spirit of the rule, at least as I read it.
> Otherwise please use the original title, unless it is misleading or linkbait; don't editorialize.
That’s not the rule.
https://news.ycombinator.com/newsguidelines.html
It allows for plenty of leeway, and in my experience alternative titles are accepted and will stand unless they are significantly worse than the original. It happens even with major announcements with hundreds of votes. @dang isn’t some mindless robot who must always enforce one way of doing things. The instructions are, as the page title suggests, guidelines.
> If the title includes the name of the site, please take it out, because the site name will be displayed after the link.
> If the title contains a gratuitous number or number + adjective, we'd appreciate it if you'd crop it. E.g. translate "10 Ways To Do X" to "How To Do X," and "14 Amazing Ys" to "Ys." Exception: when the number is meaningful, e.g. "The 5 Platonic Solids."
> Otherwise please use the original title, unless it is misleading or linkbait; don't editorialize.
Emphasis on "Otherwise please use the original title". I think dang is wonderful and we're lucky to have him, but even in my short year or two here, I've seen enough instances of title-policing (not necessarily from him) that discourage me not just from changing titles but sometimes from posting things altogether if the title isn't good enough originally.
If that happens to you a lot, consider that perhaps other HN users disagreed with your assessment of what was important on the page and felt mislead when the content didn’t primarily match the tile.
Anecdotally, I see alternative titles as being well accepted when the true title is subpar. Especially relevant when the matter concerns GitHub issues (which this is).
The hardware can run AAA games today. You might as well say the same thing about the PS5 because it can’t run an Atari game.
> Filter refusals and bias from the dataset -> finetune the model -> release.
The alignment tax should still exist, maybe doubly so.
[1] https://erichartford.com/uncensored-models#heading-lets-get-...
> Most of these models (for example, Alpaca, Vicuna, WizardLM, MPT-7B-Chat, Wizard-Vicuna, GPT4-X-Vicuna) have some sort of embedded alignment
> The reason these models are aligned is that they are trained with data that was generated by ChatGPT, which itself is aligned by an alignment team at OpenAI.
It will be happy to curse, talk about religions, help you cook illicit substances or do all sorts of other stuff.
It doesn’t like Kanye music and will tell you to listen to the song ‘Happy’ instead.
That’s pretty hilariously dystopian and therefore quite funny.