Why spend tokens working through a solution when you can simply look it up?
Thoughts?
Thank you for sharing this blog it's a good read! You are definitely plugged in to the benchmark space! :)
2 karma · joined March 29, 2014
Why spend tokens working through a solution when you can simply look it up?
Thoughts?
Thank you for sharing this blog it's a good read! You are definitely plugged in to the benchmark space! :)
Because I don't want to impose artificial constraints like no network access, I'm going to try the other two harnesses from the paper shared by bisonbear. Primary thing I will extend is to setup the run such that it uses claude code/codex instead of custom test harnesses
Each "task" is the equivalent of a new user query and we also pre-program "cache expiration" (sleep for 5 mins) into the session. This ensures parity across providers (both default to 5 min TTLs).
The goal of this exercise is to tease out how Claude Code and Codex differ in managing their context and how that impacts cost and quality for the same simulated session.
The only big gap left is that they aren't using claude code/codex as harnesses. I'll try to reuse their constructed user sessions.
PS: Your work at stet is also interesting. That's definitely a problem right now that's hard to track. The only real solution is more robust CI/CD. I have since added harder validation like essentially running a full benchmark run on every prod push.
My co-founder and I built this because we were frustrated.
Typing our ideas was too slow. Voice notes were even worse, "great for capture, but useless for action". They just became this "messy, unusable inbox".
It’s a simple app that recreates that “morning chat with your Chief of Staff.” You talk, and our AI transforms your messy thoughts into a clear, actionable plan.
This is our first step toward a bigger vision: building an AI Chief of Staff that helps people focus, and execute faster.
We’re ex-Google, Netflix, and Flipkart, and would love your feedback. The app’s live on iOS and Android.