HNHacker News
TopNewBestAskShowJobs

ActorNightly

2,279 karma · joined April 8, 2019

submissionscomments
ActorNightly··on Texas Data Center Map: See where data centers are operating or planned
My takeaway from this is apparently you can build a datacenter in a place where the weather reaches 110 degrees in the summer. AC has come a long way. I wonder what cost feasability of building essentially a biodome that is actively climate controlled that spans an entire town or city.
ActorNightly··on OpenAI's new reasoning technique alarms AI safety experts
I mean, its literally just more efficient loop of model generating reasoning text, for it to be fed back as context. There is nothing groundbreaking here. At the end of the day, its all just search.
ActorNightly··on My local model setup on an M4 Pro Mac Mini
Just ollama if you are feeling lazy. With ollama, pull the model (start with https://ollama.com/library/gemma3:27b), and `ollama run gemma3:27b`

If you want to build agentic frameworks, use llama.cpp with its built in http server, and build the framework with python

ActorNightly··on My local model setup on an M4 Pro Mac Mini
You are correct. In large part, the cost of something like Gemini on a very basic Google AI plan provides far more utility than local LLMs for coding assistance.

There are 2 main reasons for running local LLMS.

1. Process private data/work with uncensored models.

2. Use a large amount of inference that would quickly blow through rate limits and/or run up API costs.

The thing that is critical for 2 is that a) you have to have a sweetspot between a pretty good model, which means largest parameter counts, and fast enough token generation where you can run agentic loops. The latter is needed because you aren't going go get the "intelligence" of larger models to form shell commands and run tools to figure stuff out, so the only way around that is to have custom agentic loops to force the model into doing what you want, which results in more text processing.

From my testing, Gemma4:31b is basically the only local model that can be relied upon to produce accurate results. Qwen models chase benchmarks, which results in MoE models (thus the A3B in the model, i.e 3 billion parameters are only active during inference). In general, these are good for very specific tasks, but fail to be accurate in considering cross task data, whereas Gemma, being fully active does a much better job. If you only need to do a very specific deterministic task, those models are pretty good.

As an aside though, if your task involves pure text processing (for example take html data, make it into a markdown document), you can also additive train Gemma270M quite easily all on CPU, and on a decent CPU it gets like 50-100 tok/sec, no need for any extra hardware.

The thing with Macs is that while they can run those models and larger models no problem, the tok/sec is very slow. This limits effectively what you can do with the models. On the M4 that the poster mentioned, Gemma:31b will run about 20 tok/sec. That means that when you wants to write a whole code file or process large context, you have to wait for it to do things. Compared to workflow with larger models, where file generation often takes like <10 seconds, it takes a while to adapt.

The only benefit of using Macs is the price for Mini and cheaper studios. However, once you reach the total cost of about 2.5k (note that the M4 statedin the article us about 2k), building a gfx card rig is the way to go. You can get 100 tok/sec on a 3090, and it will feel a lot like the cloud models.

ActorNightly··on Apple caught off guard by AI demand for Mac Mini and Mac Studio
The reason Im so hostile is because it personally irks me when a technical forum like HN is filled with Apple-slop being passed off as technical discussion. Its fine to like Apple for what it is, the ecosystem, the industrial design, the advantages of dedicated hardware with specialized software to extend battery life. But instead, there is Apple-slop, which is basically starting with the assumption that whatever you can do on Mac is the standard, and that Macs are great because they fit that standard.

Being able to run frontier models at <10 tok/sec is just an excersize in showing that you can afford a Mac. Its useless for any real tasks. Gemma4 is on par with a lot of the older frontier models like Gemini Flash, and you can run that at 100+ tok/sec on a 2x3090 rig which costs way less than even top of the line Mac Studio.

ActorNightly··on 'Idiocracy' Predicted All of This
They are. For example, literally nobody should be having kids rn in US, because the next decade is gonna be pretty horrible. But people still are, because they are none the wiser.
ActorNightly··on FTC alleges Amazon illegally made $20B by rigging billions of ad auctions
Ah yes, we blame the most of the problems in society on advertisers, but when those advertisers get screwed over, its... bad?
ActorNightly··on Claude Fable 5.1 results on ARC-AGI
The entire premise of what that prize is is meaningless.

Basically an AGI should be able to exist as an algorithm initialized with random weights, and then learn everything from scratch. Training models on large data sets is not even remotely close to AGI.

ActorNightly··on Apple caught off guard by AI demand for Mac Mini and Mac Studio
>It was simply the best bang for the buck at the time,

I wouldn't have a problem with Apple fanboys if they just stated that they like Macs, but for some reason you have no problem with just straight up LYING. Or at least being so deluded that you state things that are so easy to prove false as fact.

I.e, 600 gb/sec is dogshit slow compared to vram speeds.

ActorNightly··on Europe's summer drought is so extreme that desertification is a growing threat
Its a case of technology outpacing our evolution.

Everyone is living in this world where we still depend on each other to survive, but connected through technology that makes it seem like we don't. The problems that people face are magically solved behind the scenes.

As such, nobody really has a sense of community or belonging, so nobody really thinks past the virtue signaling aspects of how to preserve, protect, and grow that community.

And the people that do are forced to inconvenience themselves while they see others that don't care.

ActorNightly··on The world may have less time than it thinks on climate change
[flagged]
ActorNightly··on The world may have less time than it thinks on climate change
The average view is that "climate change" == "higher/lower temperatures". Nobody really thinks of all the secondary effects.
ActorNightly··on Samsung's Processing-in-Memory (PIM)
Not if you have duplicates of rows on the first matrix, which can be done very efficiently if you build specialized hardware. Then its all just forward in parallel.
ActorNightly··on U.S. State Department pauses immigrant visa applications
Im sorry, but if you think that someone being effectively racist in a nice way is somehow better than being up front and calling them out for their belief, you are beyond lost. People like you deserve Trump. Youll ride your high horse of morality of being nice to everyone straight to your grave.
ActorNightly··on Run Qwen3.8 27B locally: real numbers from my Mac Studio
The standard is not what you can do with it, the standard is what is the free alternative. Local LLMs need to be able to beat the rate limits on all the free models like Gemini to be useful. The only way to do that is to have high enough tok/sec, especially for agentic loops.
ActorNightly··on Run Qwen3.8 27B locally: real numbers from my Mac Studio
> Qwen3.8 27B (Q4_K_M, 17GB) generates at ~14 tokens/s on my Mac Studio M3 Ultra

>~14 tokens/s

For anyone reading that has never ran local llms, please understand that anything under 100 tok/sec is worthless. You are faster typing stuff into Gemini free version that you get with a google account and copy/pasting it in (and you can easily build browser automation with playwright or any other js runtime to have this available in a chat window)

ActorNightly··on U.S. State Department pauses immigrant visa applications
Trump won because of apathy. https://www.reddit.com/r/Infographics/comments/1gms8gb/the_2...

The critisism isnt about the policy on immigration. The critisim is about analogously giving a convicted felon a flamethrower to use on your house to get rid of unwanted bugs and thinking that its the right solution. All while also crying about small government all the time. That kind of thinking is pretty much on par for modern conservatives though.

The truth is llegal immigration in the sense of criminals coming here was never a problem. Most of the "illegal" immigrants were coming here by abusing the generous asylum system. And furthermore, they were coming here for work opportunity, not to leech from the system.

The border bill that Biden had in plan would have fixed that. Trump told Republicans to veto it because it would hurt his chances.

ActorNightly··on U.S. State Department pauses immigrant visa applications
Yeah thats called xenophobia.

It be valid point if you cound point to a set of traditions and beliefs that defines what American culture is, but modern American culture is literally an average over contributions from different cultures due to all the historic immigration.

ActorNightly··on U.S. State Department pauses immigrant visa applications
Its also helpful to remind folk about why someone would hold this position.

The largest reason US got to where it was at least in the world economy stage was because of the visa program. It was the country everyone wanted to move to, because you can apply your skills, make money, start a business, with very little restriction comparatively, buy a house, and live your dream life.

The more people wanna work, the more opportunities there are for work. Any time money changes hands, like in salary for work, this grows the economy, which further creates new jobs.

If someone is for this type of restricted immigration, then you should have a clear cut position on why you think things being more expensive, less available, and overall lower economic activity (which affects things like one being able to get a mortgate or other loan, and so on) is generally better for the country.

Because if they don't have those reasons, or have ones that are not rooted in reality than the position comes from xenophobia and racism in 99% of the cases.

ActorNightly··on Nvidia agrees to acquire Hugging Face for $13B
hugging face is basically a real easy way to run llms. They also provide a bunch of libraries to do things like split compute across all your resources.
ActorNightly··on New Mac mini, featuring M6 and M5 Pro
Name one feature that Mac has that you think is so great.
ActorNightly··on New Mac mini, featuring M6 and M5 Pro
Except there isn't. Or they never make it to top.

And no, being nice to people that have no problem just straight up lying about shit is not the solution.

ActorNightly··on New Mac mini, featuring M6 and M5 Pro
Running agentic loops doesnt mean just being able to execute them. When your llm is so slow that you are faster writing the code yourself with free gemini that comes with google account, local llm is no longer worth it. Tok/sec is everything.

And 3090 is $1500 used, and 24gb gb of ram. 3x is $4500 for 72gb of ram.

ActorNightly··on Musk's SpaceX to build $100B launch facility in Louisiana
... untill Musk no longer matters. After that, that project is going in the trash faster than you can blink. Its a collosal waste of time, money, and resources, and its never going to work.
ActorNightly··on Qwen 3.8-Flash-Next releasing tomorrow (125B a6B)
GFX vram is still faster.
ActorNightly··on Apple introduces M6 and M5 Ultra
Running local models isn't about the cost, its about owning your own data and being able to run uncensored models. If you just want to code with an assistant, local inference is by far not worth it.

You also don't need to run the models 24/7, nor is your computer gonna be doing inference all the time when you are coding.

ActorNightly··on New Mac mini, featuring M6 and M5 Pro
I dunno why Apple fans do this thing where they pretend that their specific workflow is the standard for how things work, and because they can do it so well on their Macs, that means Macs are the best.

To be specific, 5 models at 57 gb means you are using crap quantized models, which suck for any real agentic work. I mean, sure they give you some inference, but compared to the full parameter models like Qwen3.8 and Gemma4 that can run full agentic loops, you may as well just use cloud inference for the price.

You of course could "run" those larger models, but we both know that the tok/sec is dogshit on Macs for those.

And 57 gb is split across 3 cards quite easily, which will all be cheaper than your comparable Mac and way faster.

You really need to educated yourself on how running local models works and what the models like Gemma 4 are capable of, so you don't continue to waste money on Macs.

ActorNightly··on New Mac mini, featuring M6 and M5 Pro
> Perhaps in a few years when both local-running LLMs are good enough on regular consumer hardware there'll be a good reason to upgrade.

I mean, local LLMs already are really good if you own Nvidia Cards, not sure what you are waiting for. 2x3090 will blow any current Mac out of the water in tok/sec

ActorNightly··on Nitter and XCancel receive cease and desist notices
Its bad, but in the end nobody really cares that much about what is good or bad.
ActorNightly··on Apple introduces M6 and M5 Ultra
Honest question - why are you so stuck on Macs for local inference?

A 4 GPU linux box with 3090s, which are $1500 a piece right now, will blow this thing out of the water. Even 2x3090 rig will run most of the good local models like Gemma4:31b at 100+ tok/sec

The VRAM of the GPUs are MUCH faster than the unified ram within Apple Silicon. The only difference is the initial model load, which takes longer from disk to VRAM due to PCIE limitations, but once the model is loaded, GPUs can prefill and and generate tokens way faster than any Apple Silicon.

So is your desire to upgrade to mac because you just aren't aware of how to set up a GPU rig, or is it something else?

← PreviousPage 4 of 34Next →