2,279 karma · joined April 8, 2019
If you want to build agentic frameworks, use llama.cpp with its built in http server, and build the framework with python
There are 2 main reasons for running local LLMS.
1. Process private data/work with uncensored models.
2. Use a large amount of inference that would quickly blow through rate limits and/or run up API costs.
The thing that is critical for 2 is that a) you have to have a sweetspot between a pretty good model, which means largest parameter counts, and fast enough token generation where you can run agentic loops. The latter is needed because you aren't going go get the "intelligence" of larger models to form shell commands and run tools to figure stuff out, so the only way around that is to have custom agentic loops to force the model into doing what you want, which results in more text processing.
From my testing, Gemma4:31b is basically the only local model that can be relied upon to produce accurate results. Qwen models chase benchmarks, which results in MoE models (thus the A3B in the model, i.e 3 billion parameters are only active during inference). In general, these are good for very specific tasks, but fail to be accurate in considering cross task data, whereas Gemma, being fully active does a much better job. If you only need to do a very specific deterministic task, those models are pretty good.
As an aside though, if your task involves pure text processing (for example take html data, make it into a markdown document), you can also additive train Gemma270M quite easily all on CPU, and on a decent CPU it gets like 50-100 tok/sec, no need for any extra hardware.
The thing with Macs is that while they can run those models and larger models no problem, the tok/sec is very slow. This limits effectively what you can do with the models. On the M4 that the poster mentioned, Gemma:31b will run about 20 tok/sec. That means that when you wants to write a whole code file or process large context, you have to wait for it to do things. Compared to workflow with larger models, where file generation often takes like <10 seconds, it takes a while to adapt.
The only benefit of using Macs is the price for Mini and cheaper studios. However, once you reach the total cost of about 2.5k (note that the M4 statedin the article us about 2k), building a gfx card rig is the way to go. You can get 100 tok/sec on a 3090, and it will feel a lot like the cloud models.
Being able to run frontier models at <10 tok/sec is just an excersize in showing that you can afford a Mac. Its useless for any real tasks. Gemma4 is on par with a lot of the older frontier models like Gemini Flash, and you can run that at 100+ tok/sec on a 2x3090 rig which costs way less than even top of the line Mac Studio.
Basically an AGI should be able to exist as an algorithm initialized with random weights, and then learn everything from scratch. Training models on large data sets is not even remotely close to AGI.
I wouldn't have a problem with Apple fanboys if they just stated that they like Macs, but for some reason you have no problem with just straight up LYING. Or at least being so deluded that you state things that are so easy to prove false as fact.
I.e, 600 gb/sec is dogshit slow compared to vram speeds.
Everyone is living in this world where we still depend on each other to survive, but connected through technology that makes it seem like we don't. The problems that people face are magically solved behind the scenes.
As such, nobody really has a sense of community or belonging, so nobody really thinks past the virtue signaling aspects of how to preserve, protect, and grow that community.
And the people that do are forced to inconvenience themselves while they see others that don't care.
>~14 tokens/s
For anyone reading that has never ran local llms, please understand that anything under 100 tok/sec is worthless. You are faster typing stuff into Gemini free version that you get with a google account and copy/pasting it in (and you can easily build browser automation with playwright or any other js runtime to have this available in a chat window)
The critisism isnt about the policy on immigration. The critisim is about analogously giving a convicted felon a flamethrower to use on your house to get rid of unwanted bugs and thinking that its the right solution. All while also crying about small government all the time. That kind of thinking is pretty much on par for modern conservatives though.
The truth is llegal immigration in the sense of criminals coming here was never a problem. Most of the "illegal" immigrants were coming here by abusing the generous asylum system. And furthermore, they were coming here for work opportunity, not to leech from the system.
The border bill that Biden had in plan would have fixed that. Trump told Republicans to veto it because it would hurt his chances.
It be valid point if you cound point to a set of traditions and beliefs that defines what American culture is, but modern American culture is literally an average over contributions from different cultures due to all the historic immigration.
The largest reason US got to where it was at least in the world economy stage was because of the visa program. It was the country everyone wanted to move to, because you can apply your skills, make money, start a business, with very little restriction comparatively, buy a house, and live your dream life.
The more people wanna work, the more opportunities there are for work. Any time money changes hands, like in salary for work, this grows the economy, which further creates new jobs.
If someone is for this type of restricted immigration, then you should have a clear cut position on why you think things being more expensive, less available, and overall lower economic activity (which affects things like one being able to get a mortgate or other loan, and so on) is generally better for the country.
Because if they don't have those reasons, or have ones that are not rooted in reality than the position comes from xenophobia and racism in 99% of the cases.
And no, being nice to people that have no problem just straight up lying about shit is not the solution.
And 3090 is $1500 used, and 24gb gb of ram. 3x is $4500 for 72gb of ram.
You also don't need to run the models 24/7, nor is your computer gonna be doing inference all the time when you are coding.
To be specific, 5 models at 57 gb means you are using crap quantized models, which suck for any real agentic work. I mean, sure they give you some inference, but compared to the full parameter models like Qwen3.8 and Gemma4 that can run full agentic loops, you may as well just use cloud inference for the price.
You of course could "run" those larger models, but we both know that the tok/sec is dogshit on Macs for those.
And 57 gb is split across 3 cards quite easily, which will all be cheaper than your comparable Mac and way faster.
You really need to educated yourself on how running local models works and what the models like Gemma 4 are capable of, so you don't continue to waste money on Macs.
I mean, local LLMs already are really good if you own Nvidia Cards, not sure what you are waiting for. 2x3090 will blow any current Mac out of the water in tok/sec
A 4 GPU linux box with 3090s, which are $1500 a piece right now, will blow this thing out of the water. Even 2x3090 rig will run most of the good local models like Gemma4:31b at 100+ tok/sec
The VRAM of the GPUs are MUCH faster than the unified ram within Apple Silicon. The only difference is the initial model load, which takes longer from disk to VRAM due to PCIE limitations, but once the model is loaded, GPUs can prefill and and generate tokens way faster than any Apple Silicon.
So is your desire to upgrade to mac because you just aren't aware of how to set up a GPU rig, or is it something else?