HNHacker News
TopNewBestAskShowJobs

hagen8

31 karma · joined February 25, 2026

submissionscomments
hagen8··on Sharing AI progress in mathematics
This is what they are trying to do with Lean
hagen8··on 2026 Eclipse Webcams
Mountains close to leon?
hagen8··on DeepSeek V4 Flash 0731
Cached input tokens are what drives most costs.
hagen8··on Building an Advanced Agentic Harness
Wrong. They are commonly used by millions.
hagen8··on Building an Advanced Agentic Harness
Check out academic papers about:

1. Hierarchical skills, workflow, skill learning 2. Meta Harness, self-learning harnesses 3. Trace/trajectory representation 4. Common agentic benchmarks

But first more basic things like 5. Blog posts form anthropic 6. How Claude Code/PI/ Hermes!! agent works 7. Agent sessions/ Forking/ Hooks

hagen8··on When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation
Check out https://agents-last-exam.org/ there is still room for improvements!
hagen8··on Show HN: Run an 80B Qwen in 4.3 GB of RAM on a Mac, and a 35B on an iPhone
There are certain physical limits. Calculations need to be done. Either less calculations are necessary for the intelligence, or u accept less intelligence. But there is a limit in what u can do with specific hardware.
hagen8··on Elevators
Most importantly, after entering the elevator. First press the close button and then the floor. That way u, safe the time of pressing a button as the door is already closing.
hagen8··on AI companies are shredding rare books
Where are the sources for that?
hagen8··on IRGC claims it destroyed Amazon's Bahrain data center
This is the claude code frontend-skill.
hagen8··on Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber
Just switch the model, its not that much effort tbh. And u can also get a cheaper model than 2.5 lite for the same intelligence
hagen8··on Human mathematicians are being outcounterexampled
This will soon happen with theoretical physics, computer science, and everything which can be verified cheaply. Then, we will have long running projects augmented by agents for 2-4 years while AI companies are collecting data of human workflows. After that we will see AI being able to do those projects by themselves. This will lead to super fast human progress and cheap products. The price of things will be bound by energy and natural resources. Interesting times are ahead of us
hagen8··on Running Gemma 4 26B at 5 tokens/sec on a 13-year-old Xeon with no GPU
Inference costs will go down massively once they use the upcoming GPUs. I estimated that a model like GLM5.2 will be around 0.03USD/M output tokens in 2 years when the Feynman GPUs will be available in 2028. And this did not even consider architectural efficiency improvements. In mid 2027 we will already see a 10x reduction once everyone has switched to the Ruby architecture.

It will be feasible for everyone to have 20 different agents running at all times. A new world is coming

hagen8··on Running Gemma 4 26B at 5 tokens/sec on a 13-year-old Xeon with no GPU
Some ppl don't like to hear it. But I would assume that token costs when using an inference provider are cheaper than electricity of using locally.

If we just take into account output token generation for simplicity. With 5tps u get 18k tokens an hour. That would costs around 0.005USD from an inference provider.

I estimate that the server consumes probably around 500W during inference.

In Germany where 1kwh cost around 0.3USD, 18k tokens inferred locally would therefore cost 0.15USD which is 30x the costs of using an inference provider.

But for ppl who worry about their data, running locally might still be good. However, they should be aware, that it is much less efficient than using an inference provider.

The efficiency gap will also significantly increase as new GPUs will make inference much more efficient.

EDIT: I first thought it'd be 180k token, but thanks to someone mentioning in the comments, it is 18k. I guess with that, it will be tough unless u got electricity almost for free. Also, the inference providers are probably still using H200/H100 for those small models. Once they use GB300 or next year the new Ruby GPUs, inference will be cheaper by a factor of 30. By then, running local models will mostly be about privacy.

hagen8··on GPT-5.6
In my opinion Opus is waaayy better in agentic orchestration. It feels like it can natively deal with multiple subagents whereas gpt needs to be taught extensively.
hagen8··on Herdr: Agent multiplexer that lives in your terminal
This is way to complex... Why don't just use some harness which manages all that and give u a good UI?
hagen8··on 1M context is now generally available for Opus 4.6 and Sonnet 4.6
Well, the question is what is contributing to the usage. Because as the context grows, the amount of input tokens are increasing. A model call with 800K token as input is 8 times more expensive than a model call with 100K tokens as input. Especially if we resume a conversation and caching does not hit, it would be very expensive with API pricing.
hagen8··on 1M context is now generally available for Opus 4.6 and Sonnet 4.6
Did u use the API or subscription?
hagen8··on Redox OS has adopted a Certificate of Origin policy and a strict no-LLM policy
They will sooner or later change that policy or get very slow in keeping up.
hagen8··on GPT-5.4
But does it use the same agent harness? Because the harness determines the behavior a lot.