HNHacker News
TopNewBestAskShowJobs

achalpandey

2 karma · joined March 29, 2014

vachiai.com Reduce your claude code API bill.
submissionscomments
achalpandey··on How are you measuring Claude Code and Codex performance?
Hmm, I am trying to benchmark cost/quality for real world sessions. In that scenario "model resourcefulness" and efficiency is actually a good thing.

Why spend tokens working through a solution when you can simply look it up?

Thoughts?

Thank you for sharing this blog it's a good read! You are definitely plugged in to the benchmark space! :)

achalpandey··on How are you measuring Claude Code and Codex performance?
UPDATE: turns out "some" models know how to game the premise. They simply lookup the solution to the exact solved SWE bench problems! Haha

Because I don't want to impose artificial constraints like no network access, I'm going to try the other two harnesses from the paper shared by bisonbear. Primary thing I will extend is to setup the run such that it uses claude code/codex instead of custom test harnesses

achalpandey··on How are you measuring Claude Code and Codex performance?
Here's my current plan, the "session" will be made up of multiple SWE bench tasks stitched together.

Each "task" is the equivalent of a new user query and we also pre-program "cache expiration" (sleep for 5 mins) into the session. This ensures parity across providers (both default to 5 min TTLs).

The goal of this exercise is to tease out how Claude Code and Codex differ in managing their context and how that impacts cost and quality for the same simulated session.

achalpandey··on How are you measuring Claude Code and Codex performance?
Thank you! Both of those papers are super new and super relevant.

The only big gap left is that they aren't using claude code/codex as harnesses. I'll try to reuse their constructed user sessions.

PS: Your work at stet is also interesting. That's definitely a problem right now that's hard to track. The only real solution is more robust CI/CD. I have since added harder validation like essentially running a full benchmark run on every prod push.

achalpandey··on How are you measuring Claude Code and Codex performance?
Looking for feedback and thoughts. Here's a link to my one-page spec: https://docs.google.com/document/d/e/2PACX-1vRu5Fv5-KTJDnCEx...
achalpandey··on Show HN: Speak your mind and get a prioritized action plan instantly
Hey HN,

My co-founder and I built this because we were frustrated.

Typing our ideas was too slow. Voice notes were even worse, "great for capture, but useless for action". They just became this "messy, unusable inbox".

It’s a simple app that recreates that “morning chat with your Chief of Staff.” You talk, and our AI transforms your messy thoughts into a clear, actionable plan.

This is our first step toward a bigger vision: building an AI Chief of Staff that helps people focus, and execute faster.

We’re ex-Google, Netflix, and Flipkart, and would love your feedback. The app’s live on iOS and Android.

Link: https://link.vachiai.com/PaKL/iwh2dlhb