HNHacker News
TopNewBestAskShowJobs

martianvoid

73 karma · joined May 7, 2026

submissionscomments
martianvoid··on Fixing GRPO's credit assignment problem without evaluating every step
I think maybe they have enough compute to try out all the promising directions at a smaller scale first and then port them to larger run

They also have incredibly smart people that can understand which improvements and directions are even worth pursuing for, I think these people are being paid in millions

martianvoid··on Solving Factorio Quality
I am nerd snipped
martianvoid··on LLM Agents Can Easily Tamper with Their Own Traces
Isn't this trivial? Any coding agent can and will perturb/modify it's own trace/session files if provided the right incentive structures, sometimes I just ask claude code to remove sections from it's own session that didn't contributed anything and were unnecessarily bloating the context
martianvoid··on Llama.cpp banned him for accidentally tagging in a fork PR
related: https://news.ycombinator.com/item?id=49859982
martianvoid··on Anthropic: The Situation Report
I don't care if this is a PR stunt or a highly exxaggerated contribution of claude in the entire ebola tracking scenarios, I feel very optimistic about the future around how these models are just becoming a daily life tool for everyone and not just programmers or tech bros
martianvoid··on Prism Inference
> Input Output

> Kimi K3 $3.00 $15.00

> Qwen3.8 2.4T-A95B $2.00 $6.00

Why is kimi priced at twice that of qwen 3.8? If I remember correctly both of them are of the same 2.5T param family. I wonder kind of optimizations or hardware they have that make such pricing differences for the same sizes of models

martianvoid··on OpenAI expects to burn $280B by 2030
Where is all this money even going??
martianvoid··on Show HN: Claude Style Patch – A dropin Claude.MD section for improving Claudish
Would be great if there were before and after comparisons of some outputs in the README, makes it more easier to understand how much value this will be adding If I want to use it
martianvoid··on A Story About Joe
Now imagine there are 2 futures from here on, either joe becomes so much mature and internalizes so much context around the product and your team that your boss realizes it's faster and much cheaper to remove you as the middle man and promote joe to your level and your boss will directly work with joe or either joe starts asking for more money per PR because joe realizes how valuable he is to you and your company
martianvoid··on Microduck
This actually seems awesome, just going through the video on their site made me burst out with 100 different ways in which I would tinker with it, spend time in learning different gaits, some RL or some tricks to have it jump or do a back flip

I can see some parallel between how arduino gave hobbyist a mini computer they can use in any way imaginable, I can see this kind of thing being used to solve unimaginable problems by hobbyist or RL computer scientist or heck any UG student with claude subscription

martianvoid··on Ask HN: What's your biggest regret in life?
Investing in people early
martianvoid··on Reduce Claudish: Paul-Graham-writing-style-in-Claude.md
> Thinking, tool calls, code, etc. remain completely unconstrained.

I wonder how much of these tricks actually work? As a side effect does it make the model produce thinking tokens to remember not to do that, “ohhh wait the user instructions says I should not use PG style in my thinking switching back to my …”

martianvoid··on A joke domain purchase turned in geopolitical warfare
Slightly unrelated but I have no idea how wind speed predictions are made up high in the altitudes, is it just interpolation of data that are get from various radiosonoids across various timestamps and finding a regular or seasonal pattern?
martianvoid··on ARC-AGI Leaderboard
It's actually crazy to see the difference between opus 5 and the next best model on ARC AGI 3 when you actually look at the ARC AGI problems
martianvoid··on Harness Engineering for Self-Improvement
I am very much skeptical on the joint optimization of the harness and the model weights, I mean we can argue that claude code itself was a joint optimization of the claude, they released the first claude model and then the claude code harness then improved claudes performance on the harness then improved the harness after the model improved, but this was a process managed by humans all the time, it would be interesting to see claude improving it's own weights or harness in the future (or it might already be happening at anthropic right now)

I just didn't find the SIA paper to be testing this out rigorously or either provide sufficient evidence that it works

martianvoid··on Show HN: Stage CLI – An easier way of reading your AI generated changes locally
This is just git with extra steps...
martianvoid··on LAWS: A new transform operation turning LLM inference into cheap cache lookups
Is there any codebase associated with this that we can check and tinker with? or is this just a promo thing?