Any artifacts or blogs I can check out? I'm curious how you manage to make them all useful in parallel. I have a hard enough time getting one instance of Qwen3.6-27B being useful full time haha.
I think this is overselling their capabilities. I've used Gemma 4 and Qwen 3.6 quite a bit on my strix halo home server. They're great models and the dense variants are significantly better, but they're still very far behind the frontier. If you boot up Gemma 4 MoE and OpenCode/Pi and expect to perform anything like Claude Code or Codex you're going to be very disappointed.
My strix halo board is feeling more useful and less toylike with the recent performance gains combined from MTP, better quantization, and generalized performance improvements across the stack. For example, I can run Unsloth's Gemma4-31B 4-bit QAT model with around 30tg and 200pp. I don't find that to be too slow at all. Particularly because it's nearly full accuracy and good enough for a lot of different stuff I throw at it.
I think it also helps that I'm using my machine to do home server stuff. It excels at all of the traditional workloads. Then I can lean on the AI to help with automation here and there. I find it deeply satisfying.
The problem with this question is that it encompasses a huge spectrum of capabilities and expectations. If you can only run an 8B model and expect it to be good at vibe coding / one shotting things you're going to have a bad time.
If you're able to run a model on the scale of ~30B, you can find that with a reasonably scoped and well defined task they do very well. I've found both Gemma4-31B and Qwen3.6-27B to be the best in this range at the moment. You can swap in the MoE models for faster inference, but they are noticeably worse at most tasks. They can one-shot / vibe code tasks with small scope, but still do much better with guidance.
If you really want frontier-like capabilities, you'll probably need at least 128GB of memory and either huge compute or a lot of patience. Most people just don't have either the money or the patience to make these local models work.
The patience required for local model usage goes far beyond just waiting for tokens though. It takes a lot of effort to get things configured and working properly for your workflow and hardware.
I think it heavily depends on what you're asking the model to do. Qwen3.6, both 27B and 35B-A3B, do agentic tool use very well. Their decision making is sus, but the dense model is decent in that way. A 4-bit quant for either of those can run on many home systems with a bit of configuration.
The biggest issue I've noticed is that the chat templates for open models are really hit or miss. The default Qwen3.6 chat template mostly works these days, but depending on your workload it may cause major issues. There are plenty of "fixed" chat templates on hugging face, but people report mixed success. It really seems to depend a lot on what the tool you're using expects.
I don’t know why you’re getting downvoted. It’s true. Averaged across a wide variety of benchmarks Fable is the only Anthropic model that performs better than GPT 5.5 xhigh.
I wonder if model distillation will continue to work as well as it has. Given hidden reasoning, the ever expanding number of expected capabilities, a serious compute shortage, the looming possibility of model collapse, and dramatically higher API costs I would guess that it's getting much harder to do.
Support requests have always been the weakest link in the security chain for big corps. I've had accounts of mine turned over with 2FA disabled by humans before. I guess we shouldn't be surprised that the LLMs are doing the same thing.
The simple fact that 2FA can be removed by low level support staff drives me mad. It defeats the whole purpose of the process.
Most of the examples they've chosen seem.. not good? What an odd mix of bad game engine and AI slop. I can't imagine that this stuff makes good training data for real-world applications.
I’m not sure I understand your question. Every interaction you have with a model in a web page does the same thing in the backend. It feeds the whole conversation history, perhaps with a bit of processing, into the model so it can process the next generation. Filling the context window is how these models retain coherence.
I also cut off JetBrains recently after a long relationship with their tools. I agree with the points made by the author. The tools are clunky resource hogs for seemingly no reason. I was really excited when JetBrains announced Fleet and promised a lightweight UI with the old analysis engines as lighter background processes. It seemed like it would solve a lot of the problems I had with their IDEs. That never materialized though. They say that Fleet integrated into Air, but Air is not an IDE. So now we're just left with the diminishing value of their traditional IDE offering and some floundering attempts to get into the AI market. What a shame.
Why is nobody on HN talking about this? Unless this is a very advanced fake, it seems like the first proof that humanoid robots are actually capable of real labor.
If Wall Street was so wise they would only reward meaningful layoffs. Laying off 10% of a company by stack ranking every team accomplishes nothing. Particularly if the company just hires the same number of cut people next quarter.
If a tree has a dead branch, you cut it off. Cutting off 10% of the leaves evenly distributed among branches will remove some dead leaves, but it leaves the source of the problems unaddressed.
Well put. I too am optimistic that, in the long term, good will prevail and we'll be stronger because of the suffering. I also agree that there's happiness and meaning to be found in presence and local life. However, it feels quite hard to let it wash over me when I spend so much time at work. The hippie lifestyle is very tempting, but I want stability and a family.
Is the US one of the best places for career growth and income? I'm 30. I've been in the tech industry for several years. During the COVID tech boom I would have agreed. I made insane amounts of money for a new grad. Then I was laid off, cut, or just downright fired for unethical reasons a few times. The combination of reduced demand for software folk, the further loss of autonomy and meaning thanks to LLMs, and that blight on my resume has made it very difficult to believe this is a great place for a career. I know I'm not alone in thinking this. Many of my techy friends, and strangers that I've met, share a bleak sentiment about the future of our careers. It seems that negativity stretches far beyond tech lately. White collar work seems to be more hopeless than ever before.
I've retreated to a public servant tech role. I was drawn to the theoretical stability of this position and the idea that might effort might do some genuine good for my local community. After several months of being here I'm skeptical that I'll be able to do any good because I'm no good at the internal politics. The stability is somewhat comforting, but only in the sense that I will not starve to death. Inflation seems like it will continue to outpace my potential earnings.
Does anybody have more insight into the demand for electricians during data center construction? This article is really light on the details. I was researching it recently and got the impression that the majority of electricians hired during DC construction are much more specialized than the average residential electrician. It also seems like large data center construction typically demands a magnitude of several hundred electricians during the peak of construction. Which to me sounds like a lot less demand for the average electrician than some of the news outlets have been claiming.
I sense that the frustration you feel is that professors are able to make choices based on their values, but the average person is not. That is broadly speaking, of course.
I think it is a great shame that we live in a modern world where we do we must to survive regardless of how it makes us feel. I suspect it is the root of much suffering.
I hope the industry starts competing more on highest scores with lowest tokens like this. It's a win for everybody. It means the model is more intelligent, is more efficient to inference, and costs less for the end user.
So much bench-maxxing is just giving the model a ton of tokens so it can inefficiently explore the solution space.
I've seen plenty of people look at those metrics and they certainly do tell a story of growing inequality and instability. To me, it seems more obvious that those issues are largely unaddressed by the people in power because they're more concerned with growing their wealth than taking care of their people. I suspect that's obvious to Americans given their overwhelming distrust for institutions, politicians, etc. Unfortunately Americans seem to lack the ability to discern who actually cares about them. By seeking change we've ironically bolstered the opposition to our basic human needs.
I was surprised by the bit about Costco selling the outlet-tier trash. I don't currently have a membership, but I've generally understood their position to be quality at cost.
I don't know if Next.js, TanStack, etc are more abstract than Rails, Django, etc. They're undoubtedly more complex though. I also find it hard to believe that it's some sort of conspiracy by management to make developers more fungible. I've seen plenty of developers choose complexity with no outside pressure.
I think the unfortunate truth is the simplest. Web development has long been detached from rationality. People are drawn to complexity like moths to a flame.
Qwen3-coder-next is way worse than Sonnet 4.5. Also, despite he lack of "coder" in the name Qwen3.5 is much better at coding than Qwen3-coder-next so you might want to check that out.
I don't know how well it performs, but you can extend Qwen3.5 to 1 million token context using YaRN. Also, Nemotron 3 Super was recently released and scales up to 1 million token context natively.