HNHacker News
TopNewBestAskShowJobs

c7b

1,732 karma · joined September 3, 2022

submissionscomments
c7b··on Hands-On with the AMD Ryzen AI Halo
Agreed, the Strix Halo doesn't feel good enough for $4k. It was supposed to be $2k (and was so until 6-12 months ago), which felt like a great deal. Not the best AI chip but you get what you pay for. A tinkerer's dream that could maybe even fit into a birthday gift budget for a lucky teenager. I hate to say it, but I hope they fail with their $4k box.
c7b··on The revenge of the philosophy majors
Bit of a tangent, but it's fun to think about how much it takes to become a -er, -ian or -ist in a given field. Philosophy is probably one of the hardest, you need to be seen as up there with the all-time greats. In history or physics you probably need to be faculty, in economics you need to have a PhD, in engineering you don't even need a degree but you need to be practicing,...
c7b··on YC CEO says he ships 37K LoC AI code per day. A developer looked under the hood
I think we're on the same team, I find the general attitude towards security in the current AI scene scary. I was just hoping for a bit more ammunition than what that article gave us.
c7b··on YC CEO says he ships 37K LoC AI code per day. A developer looked under the hood
Hmm the list is a bit underwhelming. Basically, it's unnecessary requests, bloated JS, unoptimized images and generally poorly structured code. I would hate if that was where the average website is headed, but realistically, we were already headed there before LLMs. From the headline I was expecting CVEs, broken UX flows / business logic, leaked secrets.
c7b··on Price per 1M tokens is meaningless
Even more important in a local context is the difference between token generation and prompt processing speed. We tend to focus on the former, but for multi-turn/agentic workflows the latter can dominate.
c7b··on AMD Ryzen AI Halo – $4k AI Dev Kit
You're the one making things up. An M3 Ultra with 128GB RAM doesn't exist, the M3 Max has 410GB/s bandwidth [0]. I was of course talking about the M4 Max with 546GB/s, which was closer to twice the price of a Strix Halo mini PC in a typical configuration when it was still available. And memory bandwidth isn't everything, NVidia's lead in software is substantial, look up any tests comparing them side-by-side.

[0] https://en.wikipedia.org/wiki/Apple_M3 [1] https://en.wikipedia.org/wiki/Apple_M4

c7b··on AMD Ryzen AI Halo – $4k AI Dev Kit
The M4 Max with 128GB RAM has 546GB/s memory bandwidth [0], compared to Strix Halo's 250 (on the label, I've yet to see a benchmark that tops 220). It's not available at 128GB RAM anymore, at least in my shop, but when it was not so long ago it was about 4,7k, or a little over twice the price of a cheaper Strix Halo PC (around 2,2k a few months ago).

[0] https://en.wikipedia.org/wiki/Apple_M4

c7b··on AMD Ryzen AI Halo – $4k AI Dev Kit
But ideally they would be competitive, right? If your goal is LLM or Diffusion inference or - god forbid - training, you're going to get way better performance on DGX Spark. The difference is more stark than 250 vs 273 GB/s bandwidth delta would suggest.

Now I think it's totally fine to have a less capable offering, and the Strix Halo is still a mighty capable machine for inference on mid-size MoEs. At 2k it was a tinkerer's dream. But the performance difference should be reflected in the price. This is roughly a doubling of the price compared to less than a year ago without adding any notable features, it's appalling.

c7b··on AMD Ryzen AI Halo – $4k AI Dev Kit
This is just a little under the price of NVidia's DGX Spark with CUDA or a Mac with 128GB and twice the memory bandwidth. The point of Strix Halo used to be that it was half the price of those way more capable machines. You'd be crazy to buy the AMD chip at this price. But the hardware market is generally crazy right now, so I'm sure this will sell as well, unfortunately.
c7b··on Kimi K2.7 Code is generally available in GitHub Copilot
Sounds like you got it sorted, but more generally this may be interesting: https://github.com/kyuz0/amd-strix-halo-toolboxes
c7b··on Kimi K2.7 Code is generally available in GitHub Copilot
I'm sure there will be a fix for it, but it illustrates an important broader point I should probably have made above: if you opt for local AI today, expect to run into some issues. Expect to learn a bit about the tools you're using, the not-so-fun way. I'm not recommending it to non-technical friends (yet).
c7b··on Kimi K2.7 Code is generally available in GitHub Copilot
I'm saying, either you have a problem with the copyright issues related to AI training or you don't. If you do, neither Qwen nor Claude are acceptable, if not then both are. They have similar moral standing to me.

Btw, ethically sourced, open source LLMs exist! Check out eg Olmo by Allen AI: https://allenai.org/olmo

c7b··on Kimi K2.7 Code is generally available in GitHub Copilot
I'm not sure what you're trying to say. Is that a good or a bad thing? Model distillation is presumably part of the reason why Qwen is so good, yes. As a consumer, that's a good thing I would say. It's a natural counterbalance to the monopolistic tendencies of other tech segments.

If you have ethical concerns, model distillation feels like an arbitrary line to draw. Why is the first type of piracy ok, the second not? You should restrict yourself to ethical open source models. Which is btw where I genuinely hope the future of local models is going to lie. Open weights is not enough, we need fully open source models to be sustainable. Even for simple things like updating the knowledge cutoff. How we are going to distribute the training effort will be an interesting problem where I don't see an obvious solution yet. Maybe the blockchain/federated learning people can suggest something. Or university consortia, or some public sector solutions. Or something really boring - I for one would absolutely be willing to pay for DRM-free weights of an open source model (even if I could pirate them for free).

c7b··on Kimi K2.7 Code is generally available in GitHub Copilot
Unified memory feels like the future of consumer hardware, agreed! Do check out r/StrixHalo
c7b··on Kimi K2.7 Code is generally available in GitHub Copilot
Cool! Anything you want to share? I haven't looked much into my system prompt yet, do you have any tips?
c7b··on Kimi K2.7 Code is generally available in GitHub Copilot
Web search, MTP (speeds up generation), uncensored models. Lots more things on my bucket list (eg various things related to image generation).

Not gonna lie, if you're coming from ChatGPT/Claude Code, you'll mostly be adding back features you've taken for granted, or solving problems you wouldn't have had. But sometimes you do get some extra utility, like uncensored models, which have become my go-to. Not because I'm doing anything saucy, but I hated how I'd become trained to pre-emptivly self-censor my prompts. The guardrails in open weights models are no less strong than in proprietary ones, subjectively even a bit stronger in Qwen. But luckily there's an entire sub-discipline of model ablation. Another advantage would be better control over image generation (although I can't attest to that, yet).

c7b··on Kimi K2.7 Code is generally available in GitHub Copilot
Personally, I would always max out the RAM you can fit into your budget. You might get lower bandwidth (= slower generation) than you do on a Mac if you choose a Strix Halo or DGX Spark, but there are always new tweaks being discovered to speed things up. That being said, with 32GB you should be able to fit an ok quant of 35B-A3B or 27B with some context, with 64GB you should be golden.
c7b··on Kimi K2.7 Code is generally available in GitHub Copilot
Gotta say, I've lost all interest in cloud-based AI products. Too many cool features and workflows that I was once excited about that I can't or don't use anymore for a variety of reasons (price hikes, subjectively nerfed, disappeared altogether, replaced,...) for me to even remember. It's tiring.

I've set up a small rig, mostly settled on Qwen3.6 and I'm slowly adding features myself. It probably can't compete with Claude. I don't even know, I've stopped checking. It's providing a ton of value to me as is, and it only keeps getting better. All it takes is to realize that it doesn't actually matter if the grass is (maybe even objectively) greener somewhere else. Feels so good to know that it won't change under my feet. I've got this amazing, highly extensible tool, and it's mine.

c7b··on Leanstral 1.5
I'm not sure I understand the Weights policy. This site says the weights are Apache-licensed, suggesting it's open weights. But I can't find a download link. Their Huggingface profile seems to only provide an earlier snapshot [0]. Any pointers on whether/where we can or will be able to download the weights?

[0] https://huggingface.co/mistralai/Leanstral-2603

c7b··on We Are the Last People Who Know How It Works
Desire to learn is deeply ingrained in humans and hard to root out. The education system does a pretty good job at it, unfortunately, but nothing ever is 100% effective.
c7b··on We Are the Last People Who Know How It Works
Moral panic. We all have private tutors on every subject at our disposal now. It takes some special kind of mental gymnastics to conclude from this that no one will learn anything anymore henceforth. The article is just a long-winded way of saying "Kids today have it too easy", or, equivalently, "I'm getting old".
c7b··on Qwen 3.6 27B is the sweet spot for local development
This. Do consider local LLMs, but set aside a dedicated machine for it. Connect via VPN or reverse proxy. If it's not a Mac them I'd also put a server distro on it. No need for a desktop environment, save your RAM.
c7b··on Qwen 3.6 27B is the sweet spot for local development
You could fit a Q4 GLM5.2 in 512GB and still have some space for context (372-475GB for the model): https://unsloth.ai/docs/models/glm-5.2

But yeah, there's a bit of a dearth of models that could fully utilize memory in the 128-256GB bracket at the moment. But things move so fast in this space, I wouldn't base my decision on a generation of models that's just a few months old.

c7b··on Qwen 3.6 27B is the sweet spot for local development
My 2c: you don't need the Strix Halo desktop, the chip comes in many rigs, most of them cheaper, the performance difference isn't worth it. It used to be half the price of a DGX Spark or a Mac with 128GB RAM. If you can still find it at that price I'd say it's the best bang for your buck. Otherwise, Macs have 2-3x the memory bandwidth of the DGX Spark, depending on the chip, so I'd prefer them. Unless you're planning on building a cluster. The DGX Spark has two 100GB/s connectors, ideal for clustering. But I haven't checked what else you could get for the price of two DGX Sparks.
c7b··on DSpark: Speculative decoding accelerates LLM inference [pdf]
The question is also what game they're playing. Deepseek came out of a hedge fund. I think it's no coincidence that their publications tend to have a large impact on AI stock prices.

Destroying the growth story of overvalued stocks is an interesting investment strategy. It's not even new. Shortsellers understandably get terrible rep from execs, but their actions are more often in the public interest than you'd think. Normally it's exposing fraud, but here we get the really fortunate side benefit of what could eventually amount to the most significant contribution to the general software community since Linux.

c7b··on Early adversity leaves lasting molecular imprint across the body: primate study
Almost anything can be made sense of from an evolutionary perspective. Often even the opposite of what's being observed. Can be a fun game to play. The corollary is it's not useful for vetting theories for plausibility.
c7b··on GLM-5.2 – How to Run Locally
Can someone explain the math to me? Why is 1-bit only ten percent less memory than 2-bit?
c7b··on GLM-5.2 – How to Run Locally
You can get a 128GB Strix Halo for under $3k. Used to be under $2k. Even if you believe it'll be completely obsolete for AI in two years, it'll still be good for many other things. Games for at least several more years, a great home server and/or desktop almost indefinitely. Plus, we might actually reach good enough levels for some AI use cases, if we're not already there.

And never underestimate the potential for enshittification. Your local rig will only deliver better performance over time as more and more tweaks come out. With cloud services expect the opposite to happen as subsidies run out. It's entirely possible that they will intersect on a bang per buck basis within two years.

c7b··on The frontier is open-source today
What's your hardware stack?
c7b··on Vacation With An Artist – Mini-Apprenticeships with Artists in Their Studios
Why? Afaik, apprentices are employees, with all the benefits that that entails.
← PreviousPage 4 of 20Next →