HNHacker News
TopNewBestAskShowJobs

aray07

1,136 karma · joined November 30, 2017

submissionscomments
aray07··on Claude Opus 5.5
Opus 5 was just incoherent - curious to see what improvements they have made here. Would love to see some kind of postmortem to better understand how writing styles change from model to model.

I wouldn’t be surprised if Opus 5 was trained on content written by other LLMs

aray07··on Ask HN: What are you working on? (September 2026)
I’ve been working on Opslane, an open-source agent that identifies user-facing issues and investigates them.

Traditional error trackers have two failure modes:

1. False positives: They show you thousands of errors, and you can’t tell the impact on the user

2. False negatives: Many user-facing issues don’t throw exceptions, so they go unnoticed.

Opslane combines error tracking and session recording. And there is an agent that acts on both.

Opslane reduces false positives by ranking issues based on how many users are facing a particular issue. It also learns about your product by reading your code and watching your session recordings.

False negatives are harder. Opslane reviews session recordings to spot frustration. They look for rage clicks, dead clicks, and abandoned forms.

Here is a link to the repo: https://github.com/opslane/opslane

aray07··on Show HN: Watches user sessions, finds bugs that matter, and fixes them
hey, thanks for the feedback.

hey! thanks for the feedback. Only for web right now - mobile is on the roadmap.

As for privacy, I get where the concerns are coming from. But as somebody who has used session recording before - it is such a helpful tool to understand how your customers are using your product and struggling.

Also, all the data you get from a session recording - is data you already have (or could easily get).

As for the culture - I completely agree with you. This is why companies like Linear stand out (https://linear.app/now/zero-bugs-policy). One of my beliefs for starting Opslane is that we can get more companies to understand the bugs that matter and help improve the papercuts in their product.

aray07··on Show HN: Watches user sessions, finds bugs that matter, and fixes them
hey! thanks for the feedback. Only for web right now - mobile is on the roadmap.

We are recording sessions using rrweb (https://github.com/rrweb-io/rrweb)

As for privacy, I get where the concerns are coming from. But as somebody who has used session recording before - it is such a helpful tool to understand how your customers are using your product and struggling.

Also, all the data you get from a session recording - is data you already have (or could easily get).

aray07··on Show HN: Watches user sessions, finds bugs that matter, and fixes them
Thanks for the comment! I did come across posthog and it's a good validation that it is an important problem to solve!
aray07··on Show HN: Watches user sessions, finds bugs that matter, and fixes them
Since opslane watches all user session recordings, we know which screens and actions are more important. We use this to prioritize issues.
aray07··on Claude Fable 5
i have never seen this before - where you offer something and then take that away
aray07··on Opus 4.7 to 4.6 Inflation is ~45%
Came to a similar conclusion after running a bunch of tests on the new tokenizer

It was on the higher end of Anthropics range - closer to 30-40% more tokens

https://www.claudecodecamp.com/p/i-measured-claude-4-7-s-new...

aray07··on Claude Opus 4.7 costs 20–30% more per session
yeah thats the part that is unclear to me as well - if our usage capacity is now going to run out faster.
aray07··on Measuring Claude 4.7's tokenizer costs
im running some experiments on this but based on what i have seen on my own personal data - I dont think this is true

"given that Opus 4.7 on Low thinking is strictly better than Opus 4.6 on Medium, etc., etc.”

Opus 4.7 in general is more expensive for similar usage. Now we can argue that is provides better performance all else being equal but I haven’t been able to see that

aray07··on Claude Opus 4.7 costs 20–30% more per session
effort level is separate from tokenization. Tokenization impacts you the same regardless.

I find 5 thinking levels to be super confusing - I dont really get why they went from 3 -> 5

aray07··on Claude Opus 4.7 costs 20–30% more per session
i think the new qwen models are supposed to be good based on some the articles that i read
aray07··on Claude Opus 4.7 costs 20–30% more per session
anthropic’s pricing is all based on token usage

https://platform.claude.com/docs/en/about-claude/pricing

So if you are generating more tokens, you are eating up your usage faster

aray07··on Measuring Claude 4.7's tokenizer costs
are you okay with paying more for your services without any perceived improvement in the service itself?
aray07··on Claude Opus 4.7 costs 20–30% more per session
yeah thats is my biggest issue - im okay with paying 20-30% more but what is the ROI? i dont see an equivalent improvement in performance. Anthropic hasnt published any data around what these improvements are - just some vague “better instruction following"
aray07··on Measuring Claude 4.7's tokenizer costs
isn’t caveman a joke? why would you use it for real work?
aray07··on Claude Opus 4.7 costs 20–30% more per session
yeah opus 4.7 feels a lot more verbose - i think they changed the system prompt and removed instructions to be terse in its responses
aray07··on Claude Opus 4.7 costs 20–30% more per session
yeah similar for me - it uses a bunch more tokens and I haven’t been able to tell the ROI in terms of better instruction following

it seems to hallucinate a bit more (anecdotal)

aray07··on Claude Opus 4.7 costs 20–30% more per session
yeah i am still not clear why there are 5 effort modes now on top of more expensive tokenization
aray07··on Claude Opus 4.7's new tokenizer: 1.47x on English, 1.01x on Chinese
good point - i analyzed text tokenization. will run some experiments to see how visual tokenization has changed
aray07··on Claude Code adjusting down 5hr limits
Yeah this is a bit crazy and not surprising at all.

The limits have always been opaque and you never know when they change.

I started building an open-source local proxy that logs every rate-limit header Claude Code sends.

I am using it to track and get a better sense of the 5h and 7d weekly limits.

Some initial data from 11 observed 5h sessions on Max 20x: - 5h budget: roughly $120–$280 per window - 7d budget: roughly $1,300–$1,900 - Separate Sonnet-only 7d budget at ~$150 - 95% of tokens are cache reads. They barely move the meter.

It’s open source so more people can run it and we can figure out the real numbers.

https://github.com/abhishekray07/claude-meter

aray07··on Anthropic takes legal action against OpenCode
you can choose your own model in claude code and it generally defaults to Opus
aray07··on I Take My Laptop to the Gym So Claude Doesn't Have Downtime
Yeah - its definitely a new way of working and getting used to!
aray07··on Agents that run while I sleep
The dropping requirements problem is real. What's helped us is breaking the spec into numbered ACs and having the verification run per-criterion. If AC-3 fails you know exactly what got dropped.
aray07··on Agents that run while I sleep
I do it per feature, not per step. Write the AC for the whole feature upfront, then the agent builds against it. I haven't added a spec-validation step before coding but that's a good idea. Catching ambiguity in the spec before the agent runs with it would save a lot of rework
aray07··on Agents that run while I sleep
Agreed. The spec file is context. Writing acceptance criteria before you prompt provides the context the agent needs to not go off in the wrong direction. Human leverage just moved up and the plan/spec is the most important step.

Parallelism on top of bad context just gets you more wrong answers faster

aray07··on Agents that run while I sleep
This is great. The tests in this case are the spec. When you give the agent something concrete to fail against, it knows what done looks like.

The problem is if you skip that step and ask Claude to write the tests after.

aray07··on Agents that run while I sleep
i think the friction has moved upstream - now it's working on the right thing and specifying what correct looks like. i don't think we are going back to a world where we will write code by hand again.
aray07··on Agents that run while I sleep
yup, agree - i spend most of my time reviewing the spec. The highest leverage time is now deciding what to work on and then working on the spec. I ended up building the verify skill (https://github.com/opslane/verify) because I wanted to ensure claude follows the spec. I have found that even after you have the spec - it can sometimes not follow it and it takes a lot of human review to catch those issues.
aray07··on Agents that run while I sleep
Test theatre is exactly the right framing. The tests are syntactically correct, they run, they pass but do they actually prove anything?
Page 1 of 2Next →