HNHacker News
TopNewBestAskShowJobs

bisonbear

62 karma · joined September 17, 2025

Building evals for AI coding agents, on your repo. Tests pass. Nobody's measuring the rest. http://stet.sh email ben@stet.sh
submissionscomments
bisonbear··on Claude Opus 4.7
coming more in line with codex - claude previously would often ignore explicit instructions that codex would follow. interested to see how this feels in practice

I think this line around "context tuning" is super interesting - I see a future where, for every model release, devs go and update their CLAUDE.md / skills to adapt to new model behavior.

bisonbear··on How We Broke Top AI Agent Benchmarks: And What Comes Next
working on something similar to evaluate model performance over time using tasks based on your own code. obviously this is still susceptible to the same hacking mechanics documented here, but at a local level, it's easier to detect/fix, and should give a stronger signal of subjective harness/agent/context performance than these large generic benchmarks

also I keep hearing complaints that opus is nerfed, but IMO it's nice to have objective data to back that. I feel like half of the nerfing complaints are people getting past honeymoon phase...

bisonbear··on Ask HN: How do you know if a tweak to your AI skill made it better?
a bit heavier weight, but seems worthwhile if working in an org where many people consume the skill:

- find N tasks from your repo that serve as good representation of what you want the agent to do with the task - run agent with old skill/new skill against those tasks - measure test pass rate / other quality metrics that you care about with skill - token usage, speed, alignment, ... - tests aren't a great measure alone - I've found them to be almost bimodal (most models either pass/fail) and not a good differentiator - use this to make decisions about what to do with the skill - keep skill A, promote skill B, or keep tweaking

I've also had success with an "autoresearch" variant of this, where I have my agent run these tests in a loop and optimize for the scores I'm grading o

bisonbear··on Anatomy of the .claude/ folder
PRs for AGENTS.md are necessary, but not sufficient, exactly because of non-determinism. You can LGTM the AGENTS.md change, but it's so hard to know what downstream behavioral effects it has. I feel like the only way to really know is by building a benchmark on your repo, and actually A/B testing the AGENTS.md change. I'm building something in the space - happy to share if it's something that sounds interesting to you
bisonbear··on Lat.md: Agent Lattice: a knowledge graph for your codebase, written in Markdown
Very cool, interested to read more once you post! FWIW I've been building eval infras that does something adjacent/related — replaying real repo work against different agent configs, and measuring the agent's quality dimensions (pass/fail, but also human intent alignment, code review, etc.). If you want to compare notes on the harness design, or if having an independent eval of lat vs. no-lat on quickjs would be useful, happy to chat :)
bisonbear··on Ask HN: How are you keeping AI coding agents from burning money?
cost control is a policy problem - we certainly don't need to use opus 4.6 for a simple test refactor, but many people (including myself) default to it anyways. we need a way to measure cost / performance for agents on individual repos, with individual types of tasks, to get a better sense of what tasks can be trusted to cheaper agents, and what tasks must be routed to the SOTA
bisonbear··on Lat.md: Agent Lattice: a knowledge graph for your codebase, written in Markdown
managing agents.md is important, especially at scale. however I wonder how much of a measurable difference something like this makes? in theory, it's cool, but can you show me that it's actually performing better as compared to a large agents.md, nested agents.md, skills?

more general point being that we need to be methodical about the way we manage agent context. if lat.md shows a 10% broad improvement in agent perf in my repo, then I would certainly push for adoption. until then, vibes aren't enough

bisonbear··on Anatomy of the .claude/ folder
I'm also thinking on how we can put guardrails on Claude - but more around context changes. For example, if you go and change AGENTS.md, that affects every dev in the repo. How do we make sure that the change they made is actually beneficial? and thinking further, how do we check that it works on every tool/model used by devs in the repo? does the change stay stable over time?
bisonbear··on Toward automated verification of unreviewed AI-generated code
I'm becoming convinced that test pass rate is not a great indicator of model quality - instead we have to look at agent behavior beyond the test gate, such as how aligned is it with human intent, and does it follow the repo's coding standards.

I wrote a short blog about this phenomenon here if you're interested https://www.stet.sh/blog/both-pass

also +1 on placing heavy emphasis on the plan. if you have a good plan, then the code becomes trivial. I have started doing a 70/30 or even 80/20 split of time spent on plan / time implementing & reviewing

bisonbear··on GPT‑5.4 Mini and Nano
I agree with your analysis but not the conclusion.

Evals are broken - OpenAI showed that SWE Bench Verified was in the training data - models were able to reconstruct the changes from memory (https://openai.com/index/why-we-no-longer-evaluate-swe-bench...)

However, this doesn't mean we should completely give up on benchmarking. In fact, as models get more intelligent, and we give them more autonomy, I believe that tracking agent alignment to your coding standards becomes even more important.

What I've been exploring is making a benchmark that is unique per-repo - answering the question of how does the coding agent perform in my repo doing my tasks with my context. No longer do we have to trust general benchmarks.

Of course there will still be difficulties and limitations, but it's a step towards giving devs more information about agent performance, and allowing them to use that information to tweak and optimize the agent further

bisonbear··on Speed at the cost of quality: Study of use of Cursor AI in open source projects (2025)
Really interesting study. One thing I keep coming back to is that tests have no way of catching this sort of tech debt. The agent can introduce something that will make you rip your hair out in 6 months, but tests are green...

My theory is that at least some of this is solvable with prompting / orchestration - the question is how to measure and improve that metric. i.e. how do we know which of Claude/Codex/Cursor/Whoever is going to produce the best, most maintainable code *in our codebase*? And how do we measure how that changes over time, with model/harness updates?

bisonbear··on Show HN: Agentic Docs Templates, keep AI coding agents disciplined
curious how you measure/track how this actually impacts the coding agent?
bisonbear··on Structure Dictates Behavior: golden signals for agentic development teams
For agentic development teams, I see there being two ways to measure performance:

How good is the human at using the agent, and how good is the agent itself?

I agree with the thesis here that the traditional DORA metrics don't have as much signal in an agentic world. I like the metrics mentioned in the article - another one I would propose is "number of turns" - the idea being that, if the agent goes off course, the human has to spend more turns course correcting the agent, whereas if the agent is aligned, there are just a few turns in the conversation.

For the "measuring the agent itself" part, I'm convinced that traditional benchmarks are broken, and that we need a way to measure our coding agents on our tasks, and anything else is irrelevant/noise.

bisonbear··on Many SWE-bench-Passing PRs would not be merged
yikes, using AI without tests is not fun. with testing at least you have some confidence that the AI isn't going completely off track, without them you're pretty much flying blind

having linters is super important IMO - I never try to make the AI do a linter's job. let the AI focus on the hard stuff - architecture, maintainability, cleanliness, and the linter can handle the boring pieces.

I also definitely see the AI making changes that are way larger than necessary. I try to capture that in the eval by comparing a "footprint risk" which is essentially how many unnecessary changes did the AI make vs the original PR.

I would certainly like to move beyond using PRs as a sole source of truth, since humans don't always write great code either. Maybe having LLM-as-a-judge looking for scope creep/bloat would be a decent band-aid?

bisonbear··on Many SWE-bench-Passing PRs would not be merged
yea I'm down - feel free to send me an email ben@benr.build
bisonbear··on Many SWE-bench-Passing PRs would not be merged
I've been working on building out "evals for your repo" based on the theory that commonly used benchmarks like SWE-bench are broken as they are not testing the right / valuable things, and are baked into the training data (see OpenAI's research on this here https://openai.com/index/why-we-no-longer-evaluate-swe-bench...)

Interestingly, I had a similar finding where, on the 3 open-source repos I ran evals on, the models (5.1-codex-mini, 5.3-codex, 5.4) all had relatively similar test scores, but when looking at other metrics, such as code quality, or equivalence to the original PR the task was based on, they had massive differences. posted results here if anyone is curious https://www.stet.sh/leaderboard

bisonbear··on Personal Computer by Perplexity
sounds like it's another openclaw-as-a-service provider?
bisonbear··on Ask HN: How are people doing AI evals these days?
assume you're referencing coding agents - I don't think people are. If they are, it's likely using

- AI to evaluate itself (eg ask claude to test out its own skill) - custom built platform (I see interest in this space)

I've actually been thinking about this problem a lot and am working on making a custom eval runner for your codebase. What would your usecase be for this?

bisonbear··on Selection rather than prediction
Intuitively makes sense, but in my experience, a more realistic workflow is using the main agent to sub-agent delegation pattern instead of straight 7x-ing token costs.

By delegating to sub agents (eg for brainstorming or review), you can break out of local maxima while not using quite as many more tokens.

Additionally, when doing any sort of complex task, I do research -> plan -> implement -> review, clearing context after each stage. In that case, would I want to make 7x research docs, 7x plans, etc.? probably not. Instead, a more prudent use of tokens might be to have Claude do research+planning, and have Codex do a review of that plan prior to implementation.

bisonbear··on Show HN: A-MEM – Memory for Claude Code that links and evolves on its own
curious how this is different from claude-mem?

https://github.com/thedotmack/claude-mem

bisonbear··on Show HN: Share Claude Code and Codex CLI Transcripts
pretty cool. I've been testing claude/codex head to head, looks like you pass their security audit

would be cool to see/extend the ttl on the transcripts

https://agentexports.com/v/g92c990f0cfb9962a#lClk4hHKdmv52Nx... https://agentexports.com/v/ga07365c8abedbd2a#5boQrM0ZUz78LIF...

bisonbear··on Build Software. Build Users
I've experimented with something similar - my flow is to have the subagents "initialize" a persona for the task at hand, and then have the main thread simulate a debate between the personas. Not sure if it's the best approach but it's helpful to get a diversity of perspectives on an issue
bisonbear··on Spacetime as a Neural Network
Yep! Honestly it was just a random rabbit hole - I recently read Lee Smolin's Time Reborn (highly recommend btw, super fascinating read) and was curious as to what his more recent work was about, which lead me to come across The Autodidactic Universe paper. With the AI hype train full steam ahead, the paper felt newly relevant, especially as it seems that we're starting to hit plateaus for model intelligence and looking to other areas (e.g. world models) for further advancement in the field
bisonbear··on Spacetime as a Neural Network
I posted this to reddit and got a bunch of great additional reading recs from the community there! https://www.reddit.com/r/ArtificialInteligence/comments/1pyr...
bisonbear··on Show HN: Learning a Language Using Only Words You Know
checked out the tool and think it's a cool idea! one piece of feedback though - I actually feel like the inverse product would be more helpful for me. What I mean is replacing ~95% of english text with words (Chinese in my case) that I can understand, and leaving the remaining ~5% (words I definitely don't know) in English.

At least for me, there's large value in consuming bigger volumes of Chinese to get me used to pattern-matching on the characters, as opposed to only reading a smaller amount of harder characters that I'm less likely to actually encounter

bisonbear··on Show HN: Learning a Language Using Only Words You Know
I have personally had success with using Kimi for Chinese creative writing making the same assumption that Moonshot, as a Chinese company, has more/better Mandarin language pretraining data
bisonbear··on Show HN: Learning a Language Using Only Words You Know
As a fellow Mandarin learner - this is super cool! Intuitively makes a lot of sense for the "full immersion" component of language. I love to see exciting uses of AI for language learning like this instead of just more slop generation :)

I haven't dug into the github repo but I'm curious if by "guided decoding" you're referring to logit bias (which I use), or actual token blocking? Interested to know how this works technically.

(shameless self plug) I've actually been solving a similar problem for Mandarin learning - but from the comprehensible input side rather than the dictionary side:

https://koucai.chat - basically AI Mandarin penpals that write at your level

My approach uses logit bias to generate n+1 comprehensible input (essentially artificially raising the probability of the tokens that correspond to the user's vocabulary). Notably I didn't add the concept of a "regeneration loop" (otherwise there would be no +1 in N+1) but think it's a good idea.

Really curious about the grammar issues you mentioned - I also experimented with the idea of an AI-enhanced dictionary (given that the free chinese-english dictionary I have is lacking good examples) but determined that the generated output didn't meet my quality standards. Have you found any models that handle measure words reliably?

bisonbear··on One agent isn't enough
good question - however I don't think these are necessarily mutually exclusive.

I have repeatable workflows that harness the benefits of multiple agents. Repeatable workflows drive consistent results for single agents. Using multiple agents allows you to fully explore the problem space.

An example of using these concepts harmoniously would be creating a custom slash command that spawns sub-agents that each have custom prompts, causing them to do more exploration. The commands + agent prompts make the flow repeatable + improvable

bisonbear··on Ask HN: Is building a calm, non-gamified learning app a mistake?
I've been exploring the "AI as conversation partner for immersion" use case for a project I'm building and find it pretty helpful for a few reasons

1. Effectively infinite engaging comprehensible input at your level 2. Fantastic way to practice new vocabulary and grammar patterns (AI can provide correction for mistakes) 3. Somewhat fun - if you view chat as a choose your own adventure, the experience becomes more interesting

bisonbear··on Ask HN: Is building a calm, non-gamified learning app a mistake?
I've been using these fundamentals (calm, non-gamified, emphasis on focus & flow) for building a Mandarin language learning via chat with AI. My goal was to give the user a focused tool (i.e. chat with an AI at your level) and let them experiment & play at their own pace.

However, due to the more user-driven approach to this learning method (output-focused, user has to put in effort to chat with the AI and get feedback), there is more friction with using the tool. This isn't necessarily a bad thing - in fact, more friction can lead to more meaningful experiences. That being said, I believe the market will push tools to be low friction and low effort (i.e. gamified apps) that are focused on consumption rather than tools that require more user effort.

just my 2c from a fellow builder. if curious, check it out here! would love any feedback

https://koucai.chat

← PreviousPage 2 of 3Next →