HNHacker News
TopNewBestAskShowJobs

bisonbear

63 karma · joined September 17, 2025

Building evals for AI coding agents, on your repo. Tests pass. Nobody's measuring the rest. http://stet.sh email ben@stet.sh
submissionscomments

I compared Opus 4.8 vs. Opus 5 on 25 of my tasks to see what the difference was

stet.sh·19 pts·bisonbear·
0

I compared 5 popular token saving methods in Codex and found that none delivered

stet.sh·2 pts·bisonbear·
0

I ran Sonnet 5 vs. Opus 4.8 head to head on 24 tasks to see what's different

stet.sh·1 pts·bisonbear·
0

I evaluated GLM 5.2 against the frontier on tasks from real repos

stet.sh·2 pts·bisonbear·
2

I benchmarked Opus 4.8 vs. GPT 5.5 on 2 open source repos

stet.sh·3 pts·bisonbear·
0

I used autoresearch to improve my AGENTS.md, measured against real tasks

stet.sh·8 pts·bisonbear·
7

A brief investigation into the GPT-5.5 regression claims

stet.sh·1 pts·bisonbear·
0

The Opus 4.7 reasoning curve - Medium is the best default?

stet.sh·1 pts·bisonbear·
0

GPT-5.5 low vs. medium vs. high vs. xhigh: the reasoning curve on 26 real tasks

stet.sh·2 pts·bisonbear·
0

GPT-5.5 vs. GPT-5.4 vs. Opus 4.7 on 56 real coding tasks from 2 open source repo

stet.sh·4 pts·bisonbear·
0

I ran Opus 4.7 vs. Old Opus 4.6 vs. New Opus 4.6 on 28 Zod tasks

stet.sh·2 pts·bisonbear·
0

Coding evals are broken. CI is green while AI code quality goes unmeasured

stet.sh·1 pts·bisonbear·
0

Agents.md is the highest-leverage code you're not testing

stet.sh·1 pts·bisonbear·
0

Your AI coding benchmark is hiding a 2x quality gap

stet.sh·3 pts·bisonbear·
0

Things I Learned at the Claude Code NYC Meetup

benr.build·2 pts·bisonbear·
0

Claude vs. Codex in the Messy Middle

benr.build·1 pts·bisonbear·
0

Spacetime as a Neural Network

benr.build·11 pts·bisonbear·
5

One agent isn't enough

benr.build·18 pts·bisonbear·
2

Context Engineering: The New Skill for Working with AI Agents

benr.build·1 pts·bisonbear·
0

The New Math of Building with AI

benr.build·2 pts·bisonbear·
0