HNHacker News
TopNewBestAskShowJobs

o10449366

2,510 karma · joined January 31, 2019

submissionscomments
o10449366··on 9 Ads per Minute: FIFA Cup 26 – "the price of the beautiful game"
What? Lakers are literally sponsored by Wish.
o10449366··on Claude Fable 5.1 and Claude Mythos 5.1
I experience this with many of the "advanced" models. I find they're actually the most efficient on their low/medium settings, occasionally high. Anything higher than that, and they start inventing more task list items than they check off. They seem to think that every personal project needs extensive adversarial analysis and guardrails and will invent non-issues without being asked.
o10449366··on Show HN: Bullshit Detector – agent skills that fact-check videos and articles
god i wish these low effort projects were banned from Show HN

the entire premise of a "bullshit detector" that's entirely vibe-coded is laughable

o10449366··on Professor's invisible prompt trap catches 32/35 students cheating with AI
Unfortunately, I think it's just a continuation of a trend that already started pre-covid. I was surprised that most juniors and interns I worked with seemed to lack the ability to "self-help" and Google basic questions they had. I think it reflects a schooling system in the US that doesn't teach any real critical thinking - just studying until you "know" the answer, without actually teaching students the intuition of how to get there. The result is that when they encounter something not even novel, but slightly unfamiliar they don't even make an attempt to start or immediately tunnel vision down the wrong path without taking a first step back. Funnily, the latter approach feels similar to how LLMs perform on truly novel problems, too.
o10449366··on Elevated errors on Claude Opus 5
Claude has always noticeably degraded under heavy load. Opus goes from "decent to work with" to "dumb intern" depending on whether you're working at 3 AM west coast or 10 AM - 5 PM. It's part of why I cancelled my subscription - "max" plans and "extra high" effort are meaningless when there's so much variability between model availability, model performance, and harness bugs every day and every week.
o10449366··on Fable 5 vs. GPT-5.6 Sol on an NP-Hard Problem: Does /goal help?
I think like the other commentator noted, I don't do this blindly, as I also wouldn't if I was instructing a junior engineer. I have a rough sense of how long it would take me to research all angles of the task. 5 minutes for a complicated feature is too short for a comprehensive plan. Letting it run for a few hours is also probably unproductive. But 10-60 minutes depending on complexity with followup review is a nice sweet spot. Just my $0.02.
o10449366··on Fable 5 vs. GPT-5.6 Sol on an NP-Hard Problem: Does /goal help?
/goal has replaced plan mode for me. This is the pattern I use for 95% of my AI work now:

1. Read X feature of Y and tell me when you fully understand it (if there's any detail missing in the summary, repeat until the context is primed)

2. What time is it?

3. /goal Spend X minutes from $time writing a technical design doc on $feature. There must not be any vague language or ambiguity in the document. Read carry_forward_requirements.md and testing_best_practices.md and explicitly incorporate them into the document you write. The document should be executable for a contextless implementer when done and include specific code and document references and changes needed. Spend the full X minutes working on and reviewing this document - do not quit early and wait

Even just spending 10 minutes forcing GPT to write a design doc results in much more robust plans than plan mode, in my experience, and saves time I would spend iterating on the initial plan mode draft anyway.

o10449366··on CursorBench 3.1
I feel like this benchmark reiterates my disbelief that anyone uses the latest Anthropic models for any productive work. They seem to be the best at burning tokens and spawning unnecessary subagents even for well-defined and tightly scoped tasks.

Can we get a count of people that have had Claude read irrelevant documents or perform unnecessary web searches even when told not to from the beginning?

I'm starting to wonder if this increased token usage is inadvertently bleeding into how Anthropic actually trains their model, especially leading up to IPO. As older models are deprecated and users are forced onto newer models, if the default is less efficient and more token expensive that directly results in higher "profit" for Anthropic in terms of the consumption their users have to tolerate - lest they jump to a competitor.

o10449366··on Microsoft starts canceling Claude Code licenses
I switched from Anthropic to OpenAI after spending ~$40K in equivalent token costs using Claude over 3 months.

I found Opus 4.7 to be slow and wasteful with token usage. It's shocking how inefficient it is with tasks like bash tool usage and web searching, delegating them to a dozen subagents only to get stuck and never return until you esc and intervene. That, in addition to all of the broken tooling Anthropic built in to limit token usage like the broken monitoring tool made managing Claude a chore. I was happy to pay $200/month for Opus 4.5 when they had more capacity, but 4.7 felt like a huge step back and no longer worth the price and inconvenience.

I remember an OpenAI employee comment on the GPT5.5 release post about how they specifically geared it towards long-horizon tasks and its been a breathe of fresh air in that regard. I have five two-week long sessions going right now and there's been no degradation in performance or efficiency. It's much better at carrying rules/learnings forward even in long-running sessions and grounding/refreshing itself in verified facts when it loses context.

Its funny because in two weeks I've gotten way more done with GPT5.5 with way fewer tokens and way less handholding. I think this goes to show how important tooling and the harness is and how a capable model like Opus 4.7 can be severely handicapped by bad product decisions.

o10449366··on SingleRide: Longest route on NYC Subway without visiting the same station twice
The longest route on the NYC subway is precisely when you have an appointment and the train decides to stop because there's electrical issues, someone jumped on the track, "there's a train stopped ahead of us", the express decides to go local instead, someone is holding the door, your route involves the F/G or any line that serves less affluent neighborhoods...
o10449366··on DeepClaude – Claude Code agent loop with DeepSeek V4 Pro
Idk, my recent experience with Claude is that 4.7 barely knows how to use basic bash tools - how to properly check when programs have finished running, even basic stuff like how to run pytest suites and read the failed tests from the output without re-running the suite to specifically look for them. It's shockingly dumb for all of the tooling they've built into Claude Code (the useless Monitoring tool that blocks bash polling/sleeping that actually works, etc.).

I finally get fed up and started using GPT 5.5 the past 4 days and its a breath a fresh air despite feeling much more minimal. With Claude I had to write so many hooks to enforce behaviors it wouldn't remember and it lacked common sense on. GPT 5.5 does a much better job with things like knowing the AWS CDK CLI can hang on long CloudFormation deployments and it should actively check the deployment status using CloudFormation API rather than hanging for 30+ minutes - and it does this all without asking.

Maybe there's better tooling built into Codex too, but at least on the surface level it seems like how smart the model is makes a significant difference because Claude has more tools than I can count and still struggles to use "grep".

Edit: Like just now - I can't tell you how many times I day I see this sequence:

"Sorry, I'll run in parallel"

"Error editing file"

"File must be read first"

Repeat 10x for the 10 subagents Claude spawned and then it gets stuck until you press escape and it says "You rejected the parallel agents. Running directly now"

o10449366··on Uber torches 2026 AI budget on Claude Code in four months
this is the most tired hn comment ever

"X is just Y - why is it so complicated?"

its lazy and boring to read these on every thread about a disliked big company

o10449366··on Uber Torches 2026 AI Budget on Claude Code in Four Months
I spent $24,096.47 in "API" costs with my $200 Claude Code Max subscription in April.

I'm building my own saas. I spent 6 months writing the code by hand before using Claude, and that was fine, but its much faster to give the exact specs to Claude and have 3-4 sessions working in parallel with me. When you validate changes with exact test specs there's much less correction you need to do. I always hit my weekly limit and it's far cheaper for me to use this than to hire someone and spend time onboarding them.

o10449366··on Show HN: GhostBox – Borrow a disposable little machine from the Global Free Tier
As unreliable as GitHub actions are, this is what ruins nice things (free for public repos) for the rest of us.
o10449366··on Claude.ai and API unavailable [fixed]
I've been on the $200 plan for 3 months, but this will be my last month. I got great use out of 4.5 for a while, but 4.6 felt like a half step back (conflated with all the random hidden config changes during it), and 4.7 is genuinely terrible.

It's impossible to tell these days whether 4.7 is stuck because it's thinking and Anthropic suppressed all output (seriously, 4.7 will just start making changes without explaining any reasoning - how is that an upgrade?) or because the underlying infrastructure is having issues.

4.5 -> 4.7 feels like going from working with a coach-able, junior engineer that does well with clear guidance to working with a cocky mid-level that will spend too long on pointless tangents and make confidently incorrect changes without any discussion.

o10449366··on GitHub Copilot code review will start consuming GitHub Actions minutes
Cache create is 202,746,985 and cache read is 11,998,411,722 from claude-code-monitor
o10449366··on OpenAI models coming to Amazon Bedrock: Interview with OpenAI and AWS CEOs
Lol, spinning up swat teams because someone high up decides "drop everything this is my pet priority now" is politicking. It looks good for the leaders, meanwhile its the engineers pulling the all nighters and dealing with having to maintain systems that are operationally compromised from day 0 because there's no proper planning/scoping involved other than "Big Man says this needs to be done in 2 weeks"
o10449366··on GitHub Copilot code review will start consuming GitHub Actions minutes
Yeah, I'm sure the numbers are a bit inflated compared to API, but with my Claude $200/month subscription I've supposedly consumed 12,160,410,828 tokens in April for a cost of $22,733.03.
o10449366··on Email could have been X.400 times better
Yes, I throw away maybe 98% of the physical mail I receive. But I have to sort through it for the 2% that's maybe important and for some reason MUST be received by physical mail.

There's obviously a lot of physical waste in this day and age and it's arguably small by comparison, but the paper waste from mail spam still disgusts me.

o10449366··on An update on recent Claude Code quality reports
Yeah and that statement also speaks to their test rigor if they make a change that big without thoroughly testing the edge case they're modifying.
o10449366··on An update on recent Claude Code quality reports
Resuming from sessions are still broken since Feb (I had to get claude to write a hook to fix that itself), the monitoring tool doesn't work and blocks usage of what does (simple sleep - except it doesn't even block correctly so you just sidestep in more ridiculous ways), and yet there seems to be more annoying activity proxies/spinner wheels (staring into middle distance)... Like I don't know how in a span of a few months you lose such focus on your product goals. Has Anthropic reached that point in their lifecycle already where their product team is no longer staffed by engineers and they have more and more non-technical MBAs joining trying to ride the hype train?
o10449366··on ChatGPT Images 2.0
When you program, do you consider using your prior knowledge of programming cheating?
o10449366··on Claude Token Counter, now with model comparisons
So it sounds like you were using it on auto mode then if it went ahead and fixed the vulnerabilities without additional turns? If so, that isn't really a single prompt.
o10449366··on Qwen3.6-Max-Preview: Smarter, Sharper, Still Evolving
I have the M3 Max MBP with 128 GB of memory and the 40 core GPU. What's the best local model I can run today for coding?
o10449366··on Claude Token Counter, now with model comparisons
What was your single prompt? That seems highly unlikely.
o10449366··on DaVinci Resolve – Photo
Softbox lighting and it looks off because obviously no one lights their work desk like they would for a professional photo shoot.
o10449366··on Backblaze has stopped backing up your data
I've recently been looking for online backup providers and Backblaze came highly recommended to me - but I think after reading this article I'll look elsewhere because this kind of behavior seems like the first step on the path of enshittification.
o10449366··on US and Iran agree to provisional ceasefire
Is it terrorism? radicalization seems like a pretty natural human response when your family/home/community gets indiscriminately obliterated by missiles from the sky.
o10449366··on US and Iran agree to provisional ceasefire
gee, i wonder who helped put those religious fanatics in power and have helped them retain power by raining down death from the sky
o10449366··on Amazon holds engineering meeting following AI-related outages
Paywalled
Page 1 of 11Next →