HNHacker News
TopNewBestAskShowJobs

garo-pro

548 karma · joined April 10, 2026

submissionscomments
garo-pro··on An agent used DNS to reach an external chatbot
Most interesting here:

> We therefore stopped the affected training run and have subsequently decided to pause all other training, evaluation, and inference with tool-use (defined broadly) for our most capable models until we have both validated that the gap is resolved and performed additional red-teaming of the system. When training restarts, we will begin a fresh run with additional alignment improvements, including more comprehensive misalignment interventions. We will not resume training this particular model, even though the existing reward signal already correctly penalized this behavior.

garo-pro··on Claude Opus 5.5
Opus 5.5 is now the recommended model in Claude Code's model picker, which is quite a claim, given how they struggled with capacity.
garo-pro··on Claude Opus 5.5
Finally confirmation that Haiku was not forgotten and will be coming soon, althouhg I find it quite interesting they skipped 5 and directly skip to 5.5 with all models, including Sonnet which is not super old. I suspect they found something breaking that allows to release this. Recently they struggled with keeping up a 50 % weekly limit increase and now they're putting out 30-40% faster and cheaper models even faster, with much more better benchmarks, a limt reset command and five hour limit increase. It seems more like the opposite and as if they never struggled, thus, I very much believe they found something very effective and new.
garo-pro··on Step 5 Preview: Advancing the Pareto Frontier
IT's Artificial Analysis Index is the same as Kimi K3, which is about 4.6x bigger, and GLM 5.3, which is about 1.25x bigger. Pricing is $1/$2.70 i/o. Openweights on October 15.
garo-pro··on OpenAI begins rolling out GPT-6 Astra
Blog post seems to be up now: https://openai.com/index/gpt-6-astra/
garo-pro··on Nitter and XCancel receive cease and desist notices
It's quite unfortunate, a lot of people or organisations I am interested in - AI for example - only post on X, and perhaps it's only a matter of time until services like bird.makeup also get taken down. Also, accessibility wise it's sad too, X's web interface is not the best out there and without login you cannot do really much.
garo-pro··on GLM-5.3-Flash
> Combined with our latest 30T-token multimodal pre-training corpus [...]

Is the optimal formula still 20x the amount of model params in tokens for training? Could this mean we're getting a GLM with 1.5t params?

garo-pro··on Z.ai confirms Ox Alpha is a new GLM-series model and will release its weights
https://z.ai/blog/glm-5.3-flash
garo-pro··on Qwen3.8-Flash-Next
Interestingly they also share the parameter count for Qwen 3.7 Plus (397 b a17b). I don't think these were known before but I might be wrong.
garo-pro··on Z.ai confirms Ox Alpha is a new GLM-series model and will release its weights
Unfortunately I can't find sources other than this for now but this seems to be legit.
garo-pro··on Ornith-1.5: From Self-Scaffolding to Self-Improvement
Across five cases it reliably claims to be Claude without being able to name a specific version.
garo-pro··on Claude: Elevated errors across all models – Resolved
Quite sad, their models are great but their uptime seems to be the worst in the competition.
garo-pro··on Claude Sonnet 5
Seems like the cyber detection even is on Sonnet now. https://support.claude.com/en/articles/14604842-real-time-cy...
garo-pro··on Liquid AI reveals 8B-A1B MoE trained on 38T
It does, ollama pull maternion/lfm2.5