HNHacker News
TopNewBestAskShowJobs

skohan

9,265 karma · joined January 26, 2017

spencerkohan (at) gmail
submissionscomments
skohan··on GLM-5.3-Flash
I'm using the unsloth dynamic Q4 and getting good results. I was running Q5, but Q4 gives more context headroom so I can run two agents in parallel with ~100k context each with 32GB vram.
skohan··on GLM-5.3-Flash
3.8 doesn't have a minimal thinking mode, only low, medium and xhigh.
skohan··on GLM-5.3-Flash
That sounds like something is off - I'm using UD-Q4_K_XL on pi with xhigh thinking, and unless I'm vastly underestimating the complexity of the script that's the kind of task I would expect to take a couple of minutes (getting ~30t/s decode). What server are you running, and are you using the recommended parameters from qwen/unsloth?
skohan··on GLM-5.3-Flash
Your assumption is incorrect.
skohan··on GLM-5.3-Flash
I'm using pi inside a self-made harness. I've found going super lightweight with context (AGENTS.md is maybe 20 lines) and letting the model discover what it needs to gives the best results.
skohan··on GLM-5.3-Flash
I've had the exact opposite experience. I've been using 3.8 for my daily driver since last week, and I've gradually been giving it more and more complex tasks as it continues to deliver high quality results. Now I am basically handing off large complex features, and 3.8 is doing the planning, task breakdown, implementation and review with just a few notes from my side.

The tradeoff is time (especially on RDMA4 hardware) - it does take a long time and spend a lot of tokens to get to the result, but I've found I can trust the results enough that I can queue a lot of work, essentially have it running all the time and achieve a decent velocity.

It's the first small local model I've felt like I can do real work with.

skohan··on Apple introduces M6 and M5 Ultra
> interesting features

Can you elaborate at all?

skohan··on GPT-5.6 Sol Pricing Cut by 50%
Well I wouldn't say a billion times better. I've actually been having a surprising amount of success working with local models. And my investment has only been the equivalent of 4 months of a Max x20 subscription.

Experimentation and privacy are definitely advantages, but it's also quite a lot of fun.

skohan··on Qwen3.8 27B scores 52 on Artificial Analysis
This also seems like quite an esoteric use-case to me, but I guess some people might need to know when text follows a fibonacci sequence in terms of word count.
skohan··on GPT-5.6 Sol Pricing Cut by 50%
This is a big part of the reason I went local-only. Subscription limits are horrible for having a decent workflow.
skohan··on GPT-5.6 Sol Pricing Cut by 50% on OpenRouter
I would happily buy up a load of datacenter GPU's at deep discount
skohan··on Qwen3.8 27B scores 52 on Artificial Analysis
Glimmer benchmarks around Qwen 3.6 27B levels no?
skohan··on Qwen3.8 27B scores 52 on Artificial Analysis
I found Glimmer underwhelming in terms of coding - I tried it as a drop-in replacement for 3.6, and the output was noticeably worse. 3.8 has been a significant step up so far from early testing.
skohan··on Qwen3.8 27B scores 52 on Artificial Analysis
It's only been a couple days, but I haven't seen looping issues with 3.8 so far, compared to 3.6 which did occasionally have this problem.
skohan··on Qwen3.8 27B scores 52 on Artificial Analysis
At that point aren't you just edge-case testing?

Surely most of your use-cases are not novel tasks that combine obscure domains.

It seems to me the real way to evaluate the value of a model is how it performs in your real-life workflows.

skohan··on Qwen3.8 27B scores 52 on Artificial Analysis
Depends on your use-case. Over the past couple days, I've found 192k context more than enough for coding. There's more thinking for sure compared to comparably sized models (running on xhigh), but I've found the results are so much better that the entire session consumes less tokens on average since weaker models need more review passes.
skohan··on Qwen3.8 27B scores 52 on Artificial Analysis
Imo it makes sense for things to move in the direction of small, focused models that excel in one area. I use LLMs for technical work 99% of the time, I could care less about general world knowledge, or if the model is good at creative writing.

With good orchestration and delegation you can get surprisingly far with small models running on consumer hardware.

skohan··on Qwen3.8 27B scores 52 on Artificial Analysis
I'm running 3.8 27B locally, and the results from the past few days have been excellent. I find raw speed is less of an issue when you can trust the model more to reach the right result.
skohan··on Mark Zuckerberg attacks 'closed' AI rivals as Meta returns to open models
That could be a move towards regulatory capture. Enact standards that they (and OpenAI) largely control, block access to Chinese models in the US, and effectively prevent challengers from catching up.

At the same time they slow down the arms race, so they can back off on training Capex without losing their lead.

skohan··on Mark Zuckerberg attacks 'closed' AI rivals as Meta returns to open models
Isn't this meta release kind of a counterexample to that?
skohan··on Muse Glimmer: 30B-parameter model optimized for always-on local agent workflows
There are exciting developments underway in analog computing.
skohan··on Muse Glimmer: 30B-parameter model optimized for always-on local agent workflows
The power of small models isn't only that you can run them on local hardware. You can also fully own your data and workflow, and choose/fine-tune models for your specific use-case.

for clarity, I'm not agreeing with GP that small models will mean doom for data center projects

skohan··on Muse Glimmer: 30B-parameter model optimized for always-on local agent workflows
I'm using Qwen3.6 27B Q4, max context with pi on 32GB VRAM (although I'm testing out Glimmer on a feature implementation literally right now). Pi is great because it has minimal context added by the agent.

Looking forward to the 3.8 27B release to compare.

skohan··on Muse Glimmer: 30B-parameter model optimized for always-on local agent workflows
> write code to do this in this file

I haven't had to micromanage to this level. I usually start with a spec for a feature, which will be as detailed as I am opinionated about the feature. But it's usually on the level of a high-level context, plus some key implementation details (technology choices, key requirements, maybe an interface/API specification to 80% detail), and then the project already has high-level policies documented about e.g. how to structure files within the project.

Then I do a planning phase, task breakdown, and implementation of subtasks all within the model. I do read through it, but mostly the quality is good and I might make a couple notes. Then I do a review phase, which usually picks up a couple things. I'm moving towards less manual review of results and more automation as I learn what I can and can't trust the model with.

There's definitely a capability gap vs. larger models, but honestly I kind of prefer this workflow, as I stay more in touch with how the codebase is structured.

And it's great to be able to experiment as much as I want without worrying about how many tokens I'm burning or how close I am to a usage limit.

skohan··on Meta Muse Glimmer – open weights 30B local coding model
Yes they have quants for 32GB and 20GB use-cases (including mmproj and kv cache + context)
skohan··on Meta Muse Glimmer – open weights 30B local coding model
It should - the kquant-dynamic variant is targeted towards 32GB. Downloading it now to give it a try.
skohan··on Muse Glimmer: 30B-parameter model optimized for always-on local agent workflows
If you don't mind exfiltrating all your IP to the API provider
skohan··on Muse Glimmer: 30B-parameter model optimized for always-on local agent workflows
Still great if they want to play in this space. Having competition for the 24-32GB VRAM target is only good for the end user.
skohan··on Muse Glimmer: 30B-parameter model optimized for always-on local agent workflows
I've been coding using the LLM server in my living room for the past few weeks, and I haven't had this much fun with tech for ages
skohan··on Muse Glimmer: 30B-parameter model optimized for always-on local agent workflows
I could imagine pulling out all the stops to get a release over the finish line a week early if you're worried about being surpassed by another release
← PreviousPage 2 of 34Next →