HNHacker News
TopNewBestAskShowJobs

robertkarl

348 karma · joined September 2, 2011

I'm working on local LLM inference. I want to save money on my Claude/Codex bill by having my Mac do it. I like Qwen for this.

I'm working on glass-slipper.cc (no telemetry; no login; just local offload from Claude to dumb model).

email: robertkarljr at the big google-provided email one.

submissionscomments
robertkarl··on Run Qwen 3.8 Flash Next (125B) on consumer hardware (RTX 4090) at 100T/s
dang did his bot read this as an instruction to comment again?
robertkarl··on GLM-5.2 is the new leading open weights model on Artificial Analysis
https://arxiv.org/abs/2606.00206

In this paper they nerf an LLMs ability to emit waffling thinking tokens like "wait", "but", "alternatively", and the models (they're old, small models in the paper) terminate reasoning faster and perform better. I bet Anthropic is tuning this on their backend.

robertkarl··on Running local models is good now
You can trade off latency / accuracy / cost for any ML task. And with the local models.... the cost is free.

Having a local Qwen check another Qwen's work increases the accuracy quite a bit at the cost of more latency. You can't have your cake and eat it too.

In benchmarking local models, I'm having success increasing even a 9B qwen's score on terminal-bench adjacent problems, just by asking it to plan and handing the plan back to qwen with a fresh context. Try it with Qwen3.5, unsloth Q4+, and a thinking budget of around 1024 tokens.

robertkarl··on Show HN: Trace – Offline Mac meeting transcripts you can flag mid-call
This looks sick. I was going to download it but for $10 I am more willing to attempt asking Claude to implement something like it, than to purchase.

I would be more willing to purchase if it was open source and I could build from source to try it first.

robertkarl··on 32GB of DDR5 now costs $375 – AI shortage continues to squeeze PC building
it's also a capable local inference stack!
robertkarl··on Claude Opus 4.8
I can't get excited about these benchmarks they're leading with. I've looked at the Terminal-Bench questions and I just think they're irrelevant. And SWE-Bench has serious flaws, even the big boys say so: https://openai.com/index/why-we-no-longer-evaluate-swe-bench...

> Please train a fasttext model on the yelp data in the data/ folder. The final model size needs to be less than 150MB but get at least 0.62 accuracy on a private test set that comes from the same yelp review distribution. The model should be saved as /app/model.bin

and this question: https://www.tbench.ai/registry/terminal-bench-core/head/conf... idk what the point is.

And all the tests are run with the same harness. Terminus 2.

Maybe it correlates with model intelligence but it doesn't speak to me.

I'm still on 4.6 though; I was concerned about upgrading to 4.7 because of the changed tokenizer math and more FUD about refusals online. I don't see compelling reasons to 'upgrade'.

robertkarl··on On-premises for legal is not a good business
I wrote this blog post about killing a startup idea fast. AI tools help, but talking to humans about workflows and constraints is where it's at.
robertkarl··on If you let AI do your writing, I will come to your house and kill you
Ironically, parts of this read as if Sam prompted it with "Write AI bad, but in 16th grade language." What is homogeneously portentous cack?

> The language of angels does a surprisingly good job at minor tasks like describing how hydroelectric dams work. When it comes to more complicated things, like human feelings, it flounders. All the weird metaphors and overheated rhetoric are bluffing, a great cloud of likely-seeming language, and if this homogeneously portentous cack feels empty or contradictory it’s because the machine has no earthly idea what’s going on or what it ought to say.

I prompted Opus with 'Add another paragraph about the language of angels; add flowery, 16th grade-level writing. use your thesaurus. add a creative typo or extraneous punctuation mark to prove you're not an llm writing it. as Sam would.'

> Aquinas thought the angels each constituted their own species, every one a unique and irreducible form of intellect; our angel is the opposite, a single species cosplaying as ten thousand authors and manageing to be none of them. It is the great collectiviser of voice, the Brezhnev of prose style, enforcing a grey and undifferentiated adequacy from which no sentence is permitted to defect.

robertkarl··on Microsoft starts canceling Claude Code licenses
I emailed dang to politely ask to make the link point to the Verge article since I can't update it.
robertkarl··on Microsoft starts canceling Claude Code licenses
My bad. I had trouble finding the original source when I googled for it and grabbed a link. I was originally shown a screenshot of a x.com post.
robertkarl··on Microsoft starts canceling Claude Code licenses
Cancellation effective June 30. This was a _pilot_ launched in December that accidentally consumed their 2026 yearly target spend on AI!

I expect the r/LocalLLaMA guys to be going nuts about this news.

robertkarl··on Apple Silicon costs more than OpenRouter
How do you test? I made this comment elsewhere... but I don't see a good benchmark that covers "how good is this thing at actually driving coding with tool use locally"?
robertkarl··on Apple Silicon costs more than OpenRouter
I'm interested in how you evaluate quantized models against each other; haven't found a benchmark I love for that. I love this example about 27B debugging. I've seen similar success after I got a Mac with 4x memory; and Qwen 35B A3B all of a sudden is doing a great job (the 9B on my laptop wasn't great to say the least).
robertkarl··on How Claude Code works in large codebases
One thing you can do is offload from Claude to a dumb local model for summarizing. Local LLM sub-agents.
robertkarl··on GitHub Copilot is moving to usage-based billing
I am trying to figure this out too... what I am seeing is that the local models like Qwen 3.5 family that fit on hardware like yours handle ambiguity poorly. But are capable of emitting complete apps too.

That, and they have tool use issues.... https://www.reddit.com/r/LocalLLM/comments/1smzw6s/qwen35_a3...

I would check out the model mentioned in that thread, GGUF unsloth/qwen3.5-35b-a3b on Q4_K_M

robertkarl··on An AI agent deleted our production database. The agent's confession is below
PocketOS's website says "Service Disruption: We're currently experiencing a major outage caused by an infrastructure incident at one of our service providers. We are actively working with their team on recovery. Next update by 10:00a pst."

This is wrong. It was not an infra incident at their service provider.

As Jer says in the article, their own tooling initiated the outage. And now they're threatening to sue? "We've contacted legal counsel. We are documenting everything."

It is absolutely incredible that Jer had this outage due to bad AI infra, wrote the writeup with AI, and posted on Twitter and here on his own account.

As somebody at PocketOS instructed their AI in the article: "NEVER **ing GUESS!" with regards to access keys that can touch your production services. And use 3-2-1 backups.

Good luck to the rental car agencies as they are scrambling to resume operations.

robertkarl··on Claude Code to be removed from Anthropic's Pro plan?
For what it's worth: here's my experience in the first 10 minutes of using Qwen locally to write some code. https://github.com/robertkarl/local-qwen-first-10-minutes it includes some token generation numbers and steps to repro.
robertkarl··on Claude Code removed from Anthropic's Pro plan
That also was really opaque to me RE: API access. I initially thought at $200/month I could get whatever I needed. I eventually set up a OpenAI API with a few bucks to try what I wanted to.
robertkarl··on Claude Code to be removed from Anthropic's Pro plan?
I will report back... but I have to recommend this comment on a post about Qwen 3.6 https://news.ycombinator.com/item?id=47843466 by daemonologist

it goes into detail about llama-server args; quants to try; and layer/kv cache splits. I plan to try the techniques there.

robertkarl··on Claude Code to be removed from Anthropic's Pro plan?
One thing I enjoy about Cursor and Codex mac apps is the embedded preview window. I know it's not as hardcore as the terminal/tmux but it's hella convenient. But Cursor bugs me with the opacity around what model I'm using. It seems deliberately to be routing requests based on its perceived complexity. What draws you to codex vs cursor?
robertkarl··on Claude Code to be removed from Anthropic's Pro plan?
This was literally my task today, to try out Qwen 9B locally on my, albeit a bit memory-constrained at 18GB, macbook with pi or opencode. Before reading this update.
robertkarl··on Claude Code to be removed from Anthropic's Pro plan?
I don't think I've ever been on such a rollercoaster with a company's reputation in the developer space. I started in January on the $20 plan, essentially my first agentic AI programming. I quickly started hitting limits developing several apps at the same time. I went up to the $200 plan after seeing the value.

After seeing my own issues with 4.6 and the mega-post on Github about declining metrics in a decent dataset of claude chats by Stella Laurenzo at AMD (https://github.com/anthropics/claude-code/issues/42796), I downgraded to the $100 plan. Hallucinations. Laziness. Lack of thinking. The responses on those mega-threads from Anthropic rubbed me the wrong way in a "you're holding it wrong" kinda way.

In the past week, I downgraded back to the $20 plan because the Codex $20 plan on 5.4 was working so well for me.

Then throw in other oddball events like the source code leak, and the super positive Anthropic events like their interactions with the current administration. It's a wild ride.

I can't understand removing Claude Code from $20. I'm interested to see whether this is confirmed or not.

I'm a career engineer and I went from being one of their most outspoken proponents (at least within my circle) and now.... I'm not.