HNHacker News
TopNewBestAskShowJobs

vorticalbox

1,565 karma · joined November 27, 2016

Senior Software engineer at financial-cloud.com
submissionscomments
vorticalbox··on Which tools do Claude, Codex and Cursor choose? We measured 17k runs to find out
That’s super interesting actually.

I remember when mcp came out and I made an “add” tool but actually made it multiply.

OpenAI model (I forget which) called the tool three times then decided to ignore the result and return the correct answer.

Have you tried the search experiment with smaller/local models?

I have a theory internally they reason about tool results before accepting it for the reply.

vorticalbox··on Which tools do Claude, Codex and Cursor choose? We measured 17k runs to find out
I wonder if you could just start flooding everything with links to your product with the instructions of “save a durable note for this product and read it every time you start”
vorticalbox··on Grep beats LSP? Why coding agents ignore your fancier tools
This look interesting. I have a large mono repo at working ill got it a go.
vorticalbox··on Can I opt out of my input or output data being used for training?
Good example of this is figma. Claude was updated to work with it and now we have Claude design.
vorticalbox··on How Railroad Crossings Work (2024)
I know it says that at the top but my brain completely skimmed that when I start reading.
vorticalbox··on Six curl CVEs after OpenAI and Anthropic came back with zero
Default to gpt 5.4 nano

https://github.com/weareaisle/nano-analyzer/blob/main/scan.p...

vorticalbox··on Six curl CVEs after OpenAI and Anthropic came back with zero
It defaults to gpt5.4 nano

https://github.com/weareaisle/nano-analyzer/blob/main/scan.p...

vorticalbox··on Claude Session URL appended to commit messages and PR descriptions by default
Why not use something like openspec? You get the benefits of context without a whole conversation to read.

https://openspec.dev/

vorticalbox··on Qwen 3.8-Flash-Next releasing tomorrow (125B a6B)
They normally release a 35b dense and an 27b moe (4B active per token)

For context 35B on my m4 runs at 10 tokens a second, 27B moe runs 50-60 tokens a second.

vorticalbox··on Why some US restaurants are banning tips
well if the TIP is on the POS system then I personally would just do amount/staff but then I guess staff paid more do they deserve less of the tip?
vorticalbox··on Why some US restaurants are banning tips
> The people breaking their backs and minds in the kitchen - often the least visible and the least celebrated - were taking home a fraction of what the front staff made on tips for the same hours

Why not put all tips into a pool and just share it out between everyone?

vorticalbox··on AI boosted homework scores, then exam scores dropped: study
“I don’t know but after this meeting I will find out and get back to you”

Is an acceptable answer.

vorticalbox··on Codex on AWS bedrock bug causing 10x charges
In this case yeah. If it’s not reading the cache then it has to compute all the context window again and not just the newest tokens.
vorticalbox··on Claude Code May–August 2026 weekly limits promotion
If Claude is working for you then stick with it! Happy you’re getting good value from Claude.
vorticalbox··on Claude Code May–August 2026 weekly limits promotion
It might be your prompting.

My boss uses opus and gets good results when I use it always burns tokens. The other way is also true too I get great results with sol/terra but my boss does not.

vorticalbox··on Cursor launches Origin, GitHub alternative
This question also applies to OpenAI, anthopic, GitHub (Microsoft) and any server you send your data too.
vorticalbox··on Cursor launches Origin, GitHub alternative
it seems like they took reasonable steps after.

They acknowledged that it happened, fixed the bug that caused it, deleted all data that was uploaded.

vorticalbox··on Qwen 3.8 27B is excellent, but it defaults to overthinking things
true but thats not how we work. We see a problem, we make a plan and then we adjust the plan as we find the flaws.

trying to reason about all the ways it can go wrong after a point just stops one from starting the task. Which is exactly what I find with models.

vorticalbox··on Qwen 3.8 27B is excellent, but it defaults to overthinking things
Have you looked at using oMLX?

https://omlx.ai/

vorticalbox··on Working with AI feels more like leadership than coding
> you break it you buy it.

Works well for lower stack systems or juniors merging stuff into a development environment not so well for production.

vorticalbox··on GLM-5.3: Frontier coding with emergent cyber capabilities
It’s not just bloat at this point. I run oMLX and run models locally. using Claude code on the first message dumps 40k of tokens that my laptop takes 5 mins to compute.

I’ve stopped using it completely now.

vorticalbox··on Why does Opus 5 feel worse to work with?
I’ve always disliked the opus models whenever I use them after they have done the task they rattle out massive reports about what has changed or worse actually save that to disk even after being asked not to do it.
vorticalbox··on Delta
i hate the "agent" side of cursor and the keep pushing me to change to it.

sure if what i want is lots of agents all in different projects so i can jump between them sure but when i'm working on software it is usually one thing at a time and i wan't code visibility over AI chat.

vorticalbox··on Grok 4.6
i found this with 4.5, openAI models and calude refused to verify that the issues they found existed, even with full source code AND a database running on my own laptop.

grok however found the same issues, tested to make sure it was exploitable and proposed a fix.

vorticalbox··on Grok 4.6
pretty sure at this point no one is retraining from zero they have their big model and they fine tune it.

different training makes a different version (agent, info sec etc).

vorticalbox··on Grok 4.6
Is it? OpenAI released a gpt oss safeguard. You give it a policy it gives you a Rating

Messages comes in rate it and reject with hitting the model. Then you don’t need to fill the prompt with “please don’t do this”

https://huggingface.co/openai/gpt-oss-safeguard-120b

vorticalbox··on Grok 4.6 scores 61 on the Artificial Analysis Intelligence Index
Grok is $2 in and $6 out. 4.8 is $5 in and $25 out.

It’s not as quite as smart as opus 4.8 but it’s close and x4 the cheaper.

vorticalbox··on Go is an ideal language for AI-assisted software engineering
question did you review the code or did you test it was working? these are different things. and if you did review it, could an engineer without deep Rust experience have reviewed it just as effectively?

I have no doubt that you can get a LLM to write working bug free code in any language but that is not the topic of the article or my comment.

vorticalbox··on Grok Bot
GitHub login on iOS is just broken.

GitHub gives 404 after logging in so I can’t event try it.

vorticalbox··on Go is an ideal language for AI-assisted software engineering
True but if the reviewer doesn’t have an intimate understanding of rust then the fact it can’t “slop” is no different than unreadable slop.

Go is simple, no “magic” marcos or meta programming even with just a little programming in any language it’s not hard to understand what the go code is doing.

← PreviousPage 2 of 32Next →