348 karma · joined September 2, 2011
I'm working on glass-slipper.cc (no telemetry; no login; just local offload from Claude to dumb model).
email: robertkarljr at the big google-provided email one.
In this paper they nerf an LLMs ability to emit waffling thinking tokens like "wait", "but", "alternatively", and the models (they're old, small models in the paper) terminate reasoning faster and perform better. I bet Anthropic is tuning this on their backend.
Having a local Qwen check another Qwen's work increases the accuracy quite a bit at the cost of more latency. You can't have your cake and eat it too.
In benchmarking local models, I'm having success increasing even a 9B qwen's score on terminal-bench adjacent problems, just by asking it to plan and handing the plan back to qwen with a fresh context. Try it with Qwen3.5, unsloth Q4+, and a thinking budget of around 1024 tokens.
I would be more willing to purchase if it was open source and I could build from source to try it first.
> Please train a fasttext model on the yelp data in the data/ folder. The final model size needs to be less than 150MB but get at least 0.62 accuracy on a private test set that comes from the same yelp review distribution. The model should be saved as /app/model.bin
and this question: https://www.tbench.ai/registry/terminal-bench-core/head/conf... idk what the point is.
And all the tests are run with the same harness. Terminus 2.
Maybe it correlates with model intelligence but it doesn't speak to me.
I'm still on 4.6 though; I was concerned about upgrading to 4.7 because of the changed tokenizer math and more FUD about refusals online. I don't see compelling reasons to 'upgrade'.
> The language of angels does a surprisingly good job at minor tasks like describing how hydroelectric dams work. When it comes to more complicated things, like human feelings, it flounders. All the weird metaphors and overheated rhetoric are bluffing, a great cloud of likely-seeming language, and if this homogeneously portentous cack feels empty or contradictory it’s because the machine has no earthly idea what’s going on or what it ought to say.
I prompted Opus with 'Add another paragraph about the language of angels; add flowery, 16th grade-level writing. use your thesaurus. add a creative typo or extraneous punctuation mark to prove you're not an llm writing it. as Sam would.'
> Aquinas thought the angels each constituted their own species, every one a unique and irreducible form of intellect; our angel is the opposite, a single species cosplaying as ten thousand authors and manageing to be none of them. It is the great collectiviser of voice, the Brezhnev of prose style, enforcing a grey and undifferentiated adequacy from which no sentence is permitted to defect.
I expect the r/LocalLLaMA guys to be going nuts about this news.
That, and they have tool use issues.... https://www.reddit.com/r/LocalLLM/comments/1smzw6s/qwen35_a3...
I would check out the model mentioned in that thread, GGUF unsloth/qwen3.5-35b-a3b on Q4_K_M
This is wrong. It was not an infra incident at their service provider.
As Jer says in the article, their own tooling initiated the outage. And now they're threatening to sue? "We've contacted legal counsel. We are documenting everything."
It is absolutely incredible that Jer had this outage due to bad AI infra, wrote the writeup with AI, and posted on Twitter and here on his own account.
As somebody at PocketOS instructed their AI in the article: "NEVER **ing GUESS!" with regards to access keys that can touch your production services. And use 3-2-1 backups.
Good luck to the rental car agencies as they are scrambling to resume operations.
it goes into detail about llama-server args; quants to try; and layer/kv cache splits. I plan to try the techniques there.
After seeing my own issues with 4.6 and the mega-post on Github about declining metrics in a decent dataset of claude chats by Stella Laurenzo at AMD (https://github.com/anthropics/claude-code/issues/42796), I downgraded to the $100 plan. Hallucinations. Laziness. Lack of thinking. The responses on those mega-threads from Anthropic rubbed me the wrong way in a "you're holding it wrong" kinda way.
In the past week, I downgraded back to the $20 plan because the Codex $20 plan on 5.4 was working so well for me.
Then throw in other oddball events like the source code leak, and the super positive Anthropic events like their interactions with the current administration. It's a wild ride.
I can't understand removing Claude Code from $20. I'm interested to see whether this is confirmed or not.
I'm a career engineer and I went from being one of their most outspoken proponents (at least within my circle) and now.... I'm not.