HNHacker News
TopNewBestAskShowJobs

anana_

232 karma · joined December 5, 2022

submissionscomments
anana_··on Bonsai 2 27B: Near-Lossless Compression in a 9x Smaller Footprint
The benchmarks they chose are rather cherry picked to not include long context or difficult ones that involve long horizon work or many agent turns, as I suspect this is where the model shows more differences compared to the full fat one
anana_··on Nvidia is the central bank of AI
I'm curious what the actual rate of replacement/useful life of these GPUs running AI inference 24/7 is.

If these cards burn out in less than the ~5 years of depreciation that accounting puts them at, well then there will be problems.

anana_··on Qwen3.8 27B scores 52 on Artificial Analysis
Agreed. Luckily, this model also scores high in AA non-hallucination, so it knows what it doesn't know -- perfect for situations where it can just tool call a web search.
anana_··on Qwen3.8 27B scores 52 on Artificial Analysis
And to read the tea leaves a little:

3.8 actually performs slightly worse than 3.6 on AA-Omniscience Accuracy, which could imply that they traded out world knowledge for capability in other areas.

It also produces nearly twice as many tokens per task as 3.6 (and by extension, time), which may be a tradeoff required to achieve correctness at this parameter size.

anana_··on Qwen3.8 27B scores 52 on Artificial Analysis
For more context, this puts it on par with models like GLM 5.2 and GPT 5.6 Luna, which are far larger
anana_··on Qwen 3.8 27B
When stuff like this: https://doublespeed.ai/ exists I don't find that hard to believe at all, although it cuts both ways
anana_··on Qwen 3.8 27B
Seems like MTP is available immediately!
anana_··on Qwen3.8-27B
As was the case with GLM 5.3, it seems that there is still much juice to be squeezed from post-training
anana_··on Qwen 3.8 27B
Monstrous benchmarks! Hoping it is not benchmaxxed.
anana_··on GLM-5.3: Frontier coding with emergent cyber capabilities
What a week for AI model releases
anana_··on Qwen 3.8
Apparently agentic performance in Gemma was improved recently: https://x.com/googlegemma/status/2077449152062247219

Too little too late imo

anana_··on Ornith-1.0: self-improving open-source models for agentic coding
They keep mentioning a 31B dense model, but there are no benchmarks or weights for it anywhere?
anana_··on Qwen-AgentWorld: Language World Models for General Agents
It looks like the purpose of this model is to i. generate environmental sim data for doing RL on other models or ii. act as a foundation model (they trained it to select actions as well as predicting the next state in the same loop?)

Either way, neither are intended for end consumers.

anana_··on Qwen-AgentWorld: Language World Models for General Agents
I believe the benchmark listed is about simulating the environment for the various tasks, rather than doing them. It seems that the point of this model is to generate sim data to improve other models with
anana_··on Ask HN: Has anyone replaced Claude/GPT with a local model for daily coding?
Unfortunately on Strix Halo or any similar unified memory set up, dense models are gonna be dirt slow due to the tiny memory bandwidth... But I agree, 27B is superior.
anana_··on Ask HN: Has anyone replaced Claude/GPT with a local model for daily coding?
Perhaps try a different model? Just from anecdotal experience, I find that the Gemma models smaller than 31B do not tool call as often as they should.

Some of the benchmarks appear to back this up [0]

Of course, a lot depends how you are using it (inference parameters, harness, prompting, etc.), but the model is quite important too.

[0]: https://artificialanalysis.ai/models/open-source/small?model...

anana_··on CrankGPT
I have one too and it never occurred to me to use it for anything other than games. Would be interested in seeing how you did it!
anana_··on LinkedIn is scanning browser extensions
They do now - https://support.mozilla.org/en-US/kb/use-sidebar-access-tool...
anana_··on Kimi K2.6: Advancing open-source coding
Hypothesizing here, but maybe the idea is sort of a form of technological/economic warfare? Releasing performance equivalent yet more cost efficient open weight models should in theory drive the cost of inference down everywhere.

This I assume will make it more difficult for US AI labs to turn a profit, which might make investors question their sky high valuations.

Any sort of melt down in the AI sector would almost certainly spread to the wider US market.

In contrast, in China, most of the funding for AI is coming directly from the government, so it's unlikely the same capital flight scenario would happen.

anana_··on I ran Gemma 4 as a local model in Codex CLI
I'm not saying it's the latest Qwen iteration - that would be Qwen3.6.

I'm saying it's the latest iteration of the finetuned model mentioned in the parent comment.

I'm also not suggesting that it's "the latest and greatest" anything. In fact, I think it's rather clear that I'm suggesting the opposite? As in - how can a small fine tune produce better results than a frontier lab's work?

anana_··on I ran Gemma 4 as a local model in Codex CLI
It's rather surprising that a solo dev can squeeze more performance out of a model with rather humble resources vs a frontier lab. I'm skeptical of claims that such a fine-tuned model is "better" -- maybe on certain benchmarks, but overall?

FYI the latest iteration of that finetune is here: https://huggingface.co/Jackrong/Qwopus3.5-27B-v3

anana_··on I run multiple $10K MRR companies on a $20/month tech stack
Upon rereading, I'd agree. Fits with the tone of the rest of the write up.
anana_··on I run multiple $10K MRR companies on a $20/month tech stack
> Sometimes you need the absolute cutting-edge reasoning of Claude 3.5 Sonnet or GPT-4o

Dead giveaway

anana_··on Something is afoot in the land of Qwen
https://huggingface.co/Qwen/Qwen3.5-27B

I wasn't aware of that, which page mentions that?

anana_··on Something is afoot in the land of Qwen
I've had even better results using the dense 27B model -- less looping and churning on problems
anana_··on Warren Buffett dumps $1.7B of Amazon stock
They own GEICO...