HNHacker News
TopNewBestAskShowJobs

CodeReclaimers

1 karma · joined February 7, 2026

submissionscomments
CodeReclaimers··on My LLM optimization loop reward-hacked its own benchmark (and other lessons) [pdf]
Agreed. The wrinkle I thought was worth writing up is: there's no learned reward model here and no training at all. The "reward" is wall-clock executiion time and the model is frozen; the search is happening at inference time, not in an RL loop. So the usual "the proxy is a fuzzy approximation that degrades under optimization pressure" story doesn't apply.

This was on a ~200-line surface I thought I'd locked down, and it still got gamed in a way I might not have caught right away if it wasn't a nearly impossible run time (~45usec). So anyways...you apparently don't need a soft proxy or a lot of steps for this kind of thing to show up.

CodeReclaimers··on Show HN: Yesify – We raised $40M to return the word "yes" from an API endpoint
Looking forward to the day when "Yesify is down, resulting in half the internet not working" is a real headline. :)