HNHacker News
TopNewBestAskShowJobs

sosodev

2,960 karma · joined January 28, 2019

I like software.
submissionscomments
sosodev··on IronGlass Brings Legendary Soviet Cinema Lenses to Mirrorless Cameras
These prices are insane. You can buy all (most?) of the lenses they’re recreating for a fraction of the price and adapt them to a mirrorless camera no problem. I bought a Helios 44-2 recently for $100 and adapted it to my camera for like $15.
sosodev··on MSA: Memory Sparse Attention
I spent some time trying to understand this paper and I think calling this a new attention mechanism is a bit misleading. As a dead comment pointed out this is much closer to RAG. It's not exposing all 100M tokens directly to the model while doing each prediction. However, the RAG mechanisms have been integrated directly into the model architecture and that means it can have higher accuracy and lower latency. The higher accuracy is because it isn't storing text, but rather the actual in-memory representations (K/V, compressed tensor representations, routing keys, etc) of each document so it can search and utilize them more effectively. Given that it's computing up to 100x the context space it, like RAG, cannot process that volume in realtime. They explicitly state the the model needs to do offline encoding before handling inference. So you shouldn't expect to just send 100M tokens over an API and start getting a response.

I also think some of the benchmarks are misleading. Getting a RAG system to do an attention benchmark and then comparing it against a model without RAG just isn't fair. It is obviously better but it's not apples to apples. Some of the benchmarks compare against model+RAG and there the delta in performance is much smaller.

sosodev··on Apple Business
I think it's hard to know where to draw the line between derivative product and something unique. If we follow your logic that TSMC hasn't done anything new, then aren't all computer manufacturers just rehashing the ENIAC or whatever? Is a Tesla just a better model T? No, arguably we would say that these products are new to market because they've integrated new technologies in unique ways and often expended massive capital on R&D to do so. TSMC is no different.
sosodev··on Apple Business
TSMC. They dominate the semiconductor market because they're consistently first to market with the world's most advanced chip fabrication.
sosodev··on “Collaboration” is bullshit
Can we actually align incentives at scale? It seems to me that if it were possible we would live in a utopia.
sosodev··on Tinybox – A powerful computer for deep learning
Most people are using something in the llama family for inference. Llama server is my go to. Unsloth guides describe how to configure inference for your model of choice.
sosodev··on Tinybox – A powerful computer for deep learning
What models are you testing? A 120b model with hybrid attention should fit within 80gb of VRAM fine at a 4-bit quant. Also, 4-bit quants that are done well are generally fine. They certainly don’t make the model unusable.
sosodev··on New farm bill would condemn pigs to a lifetime in gestation crates
I didn't intend to. I think that domesticated animals have long had a harmonious relationship with humans so I find it a bit difficult to believe that it's always an ethical dilemma. Pets are just the most obvious lens to identify that.

I also think we need to be careful with the idea that we should entirely avoid suffering because it's impossible to do.

sosodev··on New farm bill would condemn pigs to a lifetime in gestation crates
I'm skeptical of this claim because there's clearly a growing population that hates the idea of putting anything they don't understand in their bodies. Genetically modified vegetables, food dyes, vaccines, etc.

I find it hard to believe you could convince a large portion of Americans to eat lab grown meat just to save a buck.

sosodev··on New farm bill would condemn pigs to a lifetime in gestation crates
Are all pets suffering?
sosodev··on New farm bill would condemn pigs to a lifetime in gestation crates
This is a false dichotomy. The choice is not lab grown or suffering. Farmed animals could live happy, healthy lives and then be culled in a humane way.

The problem is that it costs slightly more and our society is more concerned with cost than animal suffering.

sosodev··on How to run Qwen 3.5 locally
I’ve tried it via openrouter. It’s very good, but for some tasks frontier models are still significantly better.

For me, the 122b model is good enough on my own hardware that the downsides can be worked around for the sake of privacy and cost savings.

sosodev··on Something is afoot in the land of Qwen
I’ve been running it via llama-server with no issues. Running the latest Bartowski 6-bit quant
sosodev··on Something is afoot in the land of Qwen
Around 20ish tokens a second with 6-bit quant at very long context lengths on my AMD AI Max 395+

I’m trying to use local models whenever possible. Still need to lean on the frontier models sometimes.

sosodev··on Something is afoot in the land of Qwen
Some of the early quants had issues with tool calling and looping. So you might want to check that you're running the latest version / recommended settings.
sosodev··on Something is afoot in the land of Qwen
In my experience Qwen3.5 is better even at smaller distillations. From what I understand the Qwen3-next series of models was just a test/preview of the architectural changes underpinning Qwen3.5. So Qwen3.5 is a more complete and well trained version of those models.
sosodev··on Something is afoot in the land of Qwen
I've noticed that open weight models tend to hesitate to use tools or commands unless they appeared often in the training or you tell them very explicitly to do so in your AGENTS.md or prompt.

They also struggle at translating very broad requirements to a set of steps that I find acceptable. Planning helps a lot.

Regarding the harness, I have no idea how much they differ but I seem to have more luck with https://pi.dev than OpenCode. I think the minimalism of Pi meshes better with the limited capabilities of open models.

sosodev··on Something is afoot in the land of Qwen
I really hope this doesn't hinder development too much. As Simon says, Qwen3.5 is very impressive.

I've been testing Qwen3.5-35B-A3B over the past couple of days and it's a very impressive model. It's the most capable agentic coding model I've tested at that size by far. I've had it writing Rust and Elixir via the Pi harness and found that it's very capable of handling well defined tasks with minimal steering from me. I tell it to write tests and it writes sane ones ensuring they pass without cheating. It handles the loop of responding to test and compiler errors while pushing towards its goal very well.

sosodev··on Claude's Cycles [pdf]
The compute speed is definitely correlated with the memory consumption in LLM land. More efficient attention means both less memory and faster inference. Which makes sense to me because my understanding is that memory bandwidth is so often the primary bottleneck.

We're also seeing a recent rise in architectures boosting compute speed via multi-token prediction (MTP). That way a single inference batch can produce multiple tokens and multiply the token generation speed. Combine that with more lean ratios of active to inactive params in MOE and things end up being quite fast.

The rapid pace of architectural improvements in recent months seems to imply that there are lots of ways LLMs will continue to scale beyond just collecting and training on new data.

sosodev··on Claude's Cycles [pdf]
The model that processes search results is tiny and dumb. You shouldn't compare it to the frontier models that are solving complex math problems.
sosodev··on Claude's Cycles [pdf]
My understanding, from listening/reading what top researchers are saying, is that model architectures in the near future are going to attempt to scale the context window dramatically. There's a generalized belief that in-context learning is quite powerful and that scaling the window might yield massive benefits for continual learning.

It doesn't seem that hard because recent open weight models have shown that the memory cost of the context window can be dramatically reduced via hybrid attention architectures. Qwen3-next, Qwen3.5, and Nemotron 3 Nano are all great examples. Nemotron 3 Nano can be run with a million token context window on consumer hardware.

sosodev··on I baked a pie every day for a year
I challenge each and every one of you to make a pie by the end of the month.

I made one, for the first time in my life, last week. It brought me tremendous joy not only to make it, but to have something nice to share with friends.

sosodev··on Nano Banana 2: Google's latest AI image generation model
I think you're under-estimating how much personal taste applies in that industry. Yes, there's a lot of free content but it's often low quality and/or difficult to find for a particular niche. The OF pages, and other paid sites, are curated collections of high quality stuff that can satisfy particular cravings repeatedly with minimal effort.

A big part of it also the feeling of "connection" with the creator via messages and what not, but that too can be replicated (arguably better) by AI. In fact, a lot of those messages are already being generated haha.

sosodev··on Nano Banana 2: Google's latest AI image generation model
Doesn't Grok allow users to create lewd content or did they roll that back?

Also, I suspect that we'll soon see the same pattern of open weights models following several months behind frontier in every modality not just text.

It's just too easy for other labs to produce synthetic training data from the frontier models and then mimic their behavior. They'll never be as good, but they will certainly be good enough.

sosodev··on Step 3.5 Flash – Open-source foundation model, supports deep reasoning at speed
I see. It seems the looping is a bug in the model weights but there are bugs in detecting various outputs as identified in the PR I linked.
sosodev··on The Future of AI Software Development
Thanks for the additional info. I suspected that MiniMax M2.5 might be a bit too much for this board. 230B-A10B is just a lot to ask of the 395+ even with aggressive quantization. Particularly when you consider that the model is going to spend a lot of tokens thinking and that will eat into the comparatively smaller context window.

I switched from the Unsloth 4-bit quant of Qwen3 Coder Next to the official 4-bit quant from Qwen. Using their recommended settings I had it running with OpenCode last night and it seemed to be doing quite well. No infinite loops. Given its speed, large context window, and willingness to experiment like you mentioned I think it might actually be the best option for agentic coding on the 395+ for now.

I am curious about https://huggingface.co/stepfun-ai/Step-3.5-Flash given that it does parallel token generation. It might be fast enough despite being similar in size to M2.5. However, it seems there are still some issues that llama.cpp and stepfun need to work out before it's ready for everyday use.

sosodev··on Step 3.5 Flash – Open-source foundation model, supports deep reasoning at speed
Have you tried Qwen3 Coder Next? I've been testing it with OpenCode and it seems to work fairly well with the harness. It occasionally calls tools improperly but with Qwen's suggested temperature=1 it doesn't seem to get stuck. It also spends a reasonable amount of time trying to do work.

I had tried Nemotron 3 Nano with OpenCode and while it kinda worked its tool use was seriously lacking because it just leans on the shell tool for most things. For example, instead of using a tool to edit a file it would just use the shell tool and run sed on it.

That's the primary issue I've noticed with the agentic open weight models in my limited testing. They just seem hesitant to call tools that they don't recognize unless explicitly instructed to do so.

sosodev··on Step 3.5 Flash – Open-source foundation model, supports deep reasoning at speed
I think there are multiple ways these infinite loops can occur. It can be an inference engine bug because the engine doesn't recognize the specific format of tags/tokens the model generates to delineate the different types of tokens (thinking, tool calling, regular text). So the model might generate a "I'm done thinking" indicator but the engine ignores it and just keeps generating more "thinking" tokens.

It can also be a bug in the model weights because the model is just failing to generate the appropriate "I'm done thinking" indicator.

You can see this described in this PR https://github.com/ggml-org/llama.cpp/pull/19635

Apparently Step 3.5 Flash uses an odd format for its tags so llama.cpp just doesn't handle it correctly.

sosodev··on The Future of AI Software Development
I was testing the 4-bit Qwen3 Coder Next on my 395+ board last night. IIRC it was maintaining around 30 tokens a second even with a large context window.

I haven't tried Minimax M2.5 yet. How do its capabilities compare to Qwen3 Coder Next in your testing?

I'm working on getting a good agentic coding workflow going with OpenCode and I had some issues with the Qwen model getting stuck in a tool calling loop.

sosodev··on Dario Amodei – "We are near the end of the exponential" [video]
Isn't it just the usual feedback loop that happens with popular podcasters? They have connections and get a few highly popular guests on. As long as their demeanor is agreeable and they keep the conversation interesting other high profile guests will agree to be on and thus they've created a successful show.
← PreviousPage 4 of 25Next →