HNHacker News
TopNewBestAskShowJobs

usef-

788 karma · joined November 27, 2014

submissionscomments
usef-··on Gemini 4 Argon
I'm skeptical of anyone that has absolute confidence in either direction, to be honest. It's clearly an unknown.

Our current capabilities were science fiction a very short time ago, and we are still improving in multiple areas simultaneously (hardware, algorithms, scaling, data efficiency, inference...). We don't really know what the limit is yet.

Reversing gravity seems to counteract the current knowledge of the physical laws, but human-level intelligence doesn't (it has already been achieved once), and there's enough reason to believe that human-level intelligence itself is not a fundamental limit (energy usage constraints in evotution, brain-size limit fitting through the birth canal, etc).

usef-··on Gemini 4 Argon
If we theoretically hit the point where it could do all of its research itself, better than a human, things _would_ be different. I guess it's a question of whether we think we'll get there.
usef-··on Gemini 4 Argon
Yes, it may come down to how well they get efficiencies from scale.

People always compare the inflated API prices, but subscription prices of American models are competitive for the intelligence. You get >20x the subscription cost in tokens.

usef-··on Livenerf: Has Opus 5.5 been nerfed yet?
What sorts of things, if you can say? Is it a similar sized/complexity codebase? Most projects do become larger and/or more complex over time. And most people's standards do creep up as they learn.
usef-··on GPT 6.1 Sol: Near-Astra intelligence for a fifth of the price
Statements like these remind me how much each of us are in our own unique information bubbles. People seemed very excited about analysing the differences of each model on my feeds.

Have you used recent ones on any large projects or coding issues? They've improved tremendously lately in my experience, in terms of implementing non-trivial things, debugging, handling legacy/complex codebases, etc. I usually keep a text document of things that models failed to do/fix and recent ones have wiped it all away.

usef-··on GPT 6.1 Sol: Near-Astra intelligence for a fifth of the price
Do you use them? I feel like you're talking about the news headlines. But in daily, direct, individual usage (one model at a time) I've found they've improved tremendously over the last few months, for tech/programming work at least
usef-··on GPT 6.1 Sol: Near-Astra intelligence for a fifth of the price
Yes, even for crud apps there's a huge variability in how nicely you can make them, and how many mistakes or footguns models make along the way. I'm very curious to what level these people are building to.
usef-··on GPT 6.1 Sol: Near-Astra intelligence for a fifth of the price
I suspect most people in this position would be using the subscriptions, though.

A $20 Anthropic subscription is $500+ equivalent of API credit. You can build quite a lot on the $20 plan and get to use the best model.

Their API pricing has healthy margins built in.

usef-··on Pirating the Pirates
More otimistically, there were far fewer obsessive niches online then as the technical requirements were so large to set a site up. Even if the ones that existed were great.

There are a lot more obsessive YouTube channels and newsletters in a lot more varied niches today by comparison. I'm delighted regularly by new discoveries.

usef-··on Sonnet 5.5
It sounds like you're thinking of cryptography, not general security.

The world uses many security products that have false positives and false negatives (firewalls, intrusion detections, wafs, fraud detection, spam...). Those aren't generally considered security through obscurity.

They openly talk about the system and its drawbacks here if you're interested: https://www.anthropic.com/news/fable-safeguards-jailbreak-fr...

It's paired with rate limits, monitoring, account control, multiple classifiers, a deliberate safety margin, model design, and other things too.

(I think this was also why they faced the controversy over not having zdr in fable: they wanted to use logs to detect repeated attempts, etc. Possibly a bigger change with Opus/Sonnet is that it's zdr with the classifiers?)

usef-··on Sonnet 5.5
This is not relying on a hidden design.
usef-··on Sonnet 5.5
How is this security through obscurity?

That term is about hiding a system's design in order to secure something, rather than having secure design.

usef-··on Sonnet 5.5
And Sonnet has a similar classifier in front according to the article:

> it’s the first Sonnet model to launch with cyber safeguards

usef-··on Sonnet 5.5
If it's for personal use, is there a reason you don't want the subscription?

Anthropic's $20 subscription gives >$500 worth of credit by most measures, which is pretty similar, and you get a better model. Their raw API prices have fat margins.

And as another commenter said, Luna is the cost leader at the moment if you really need API pricing.

usef-··on Sonnet 5.5
Anthropic's "Max" modes seem like a yolo mode: "use 10x the tokens to try to break the hardest possible problems". But their models don't seem less efficient at normal reasoning modes.

I can't see Ember on AA's index yet, but their post claims "half the reasoning tokens for the same answers" as Kimi K3.

That would make it about so, I assume?

             Score  Tokens  Reason  Cost 
 Kimi K3 Max    44  48k     32k     $2.00
 Half reason    44  32k ?   16k ?    ?

 Opus Med       51  26k     12k     $1.34
 Opus High      54  36k     18k     $1.82
 Opus Max       58  119k    84k     $5.98

 Sonnet Med     41  ?       ?       $0.59
 Sonnet High    47  ?       ?       $1.08
 Sonnet Max     56  193k    142k    $7.60
Medium is Anthropic's default.

Having a less efficient mode isn't necessarily a mistake -- the purpose of configurable effort levels after all is to be able to put more thought into a problem.

usef-··on AI companies in race to demonstrate their model most threatening to humanity
Aren't they supposed to tell the public if something bad happened?

Past CEOs (chemical companies etc) we criticised for covering up mistakes.

(This article is satire, btw, for those that didn't click through)

usef-··on AI companies in race to demonstrate their model most threatening to humanity
This is a satire website too.

"said Dr. Andrew Lenson, a Senior Lecturer in all kinds of science sounding stuff"

"Anthropic shares skyrocketed upon the revelations; a remarkable development, particularly given the company is privately owned."

Perhaps it's mocking the misleading press coverage of AI companies?

usef-··on Rails World 2026 Opening Keynote [video]
At least at the current level of AI, it still takes days/weeks/months to make a good app if you have standards. That's not effort everyone will want to do.
usef-··on Rails World 2026 Opening Keynote [video]
The current Hey was designed before vibe coding. They started using AI very recently.
usef-··on Rails World 2026 Opening Keynote [video]
According to the schedule, today had a keynote on "AI and the future of Ruby & Rails — a chat with Matz & DHH". I suspect that might have clarified, but sadly they don't seem to have released a video (yet?)
usef-··on Rails World 2026 Opening Keynote [video]
Translation does play to their strengths.

But yes, I think RoR's selling feature was developer ergonomics, which suddenly seems less of a benefit if developers aren't the ones writing the code. Readability and conciseness still has benefit, but the trade-off of worse performance (and less static checking) is suddenly more questionable.

Have a look at fly.io facing similar issues, another platform whose value prop was dev ergonomics: https://fly.io/blog/kurt-scott-money-sprites/

usef-··on Claude Opus 5.5
Maybe that was a different essay?
usef-··on GPT-6 Sol and Luna
He didn't say they were pausing research. You might have only read the social media responses to his essay, not the essay itself. Social media seems even less accurate than usual when it comes to anything AI related.
usef-··on Claude Opus 5.5
The clearest limit right now is GPUs, and anyone giving up their allocation will be rerouted elsewhere in the world (NVIDIA is already sold out for the next year). I think you're underestimating competitiveness of the current market if you think Anthropic stopping right now wouldn't be absorbed by all others reasonably quickly. Your 50% figure is highly doubtful if you see how many players there are now.

Your idea of "making a start" (giving up their position) also would mean they couldn't really do anything else to solve the problem afterwards(?). Sometimes you can improve what's happening in a room more by staying in that room.

> Oh, so they're killing people's opportunities and jobs because they want to help people?

To be clear: Dario has talked about worries of jobs etc in the past, wanting society to prepare more for it, but the safety issues they're talking about with pacing seem to be focused more on their existential/AGI worries, not jobs/electricity etc. If someone truly believes in the existential worries (which they seem to: they wrote and published about it long before Anthropic was founded, and have directly made costly decisions based on it, like blocking their own models' capabilities) it trumps the other worries for them. At least that's my reading.

usef-··on Claude Opus 5.5
What do you think would be a better title?

The metaphor seems to be like a pacer runner in marathons: If you run too hard in the beginning of a marathon you will blow up and fail, so runners follow a pacer at the speed they can actually maintain safely.

Note that they wont necessarily be slower at finishing the overall race.

usef-··on Claude Opus 5.5
Yes, and pacing also stops you running too fast at the beginning of the race and blowing up. I believe "we can finish the overall race faster" is part of the metaphor they were aiming for.
usef-··on Claude Opus 5.5
It is remarkable seeing whole other subthreads here criticising the vagueness of the words "pacing the frontier", as if no text existed below the title.
usef-··on Claude Opus 5.5
What do you think would be the benefit of them stopping if others race ahead? Do you think they have some special ability no one else will find, among the many competitive firms right now?

I believe they think slowing can only be coordinated from the frontier or via government, and stopping would lose any leverage they have to help coordinate that.

(I suspect not many people read the essay, judging by how many people seem surprised they're releasing improved models)

usef-··on GPT-6 Sol and Luna
I think if they truly believe it's happening we generally want to encourage them to be honest with the public, though, don't we? We've spent decades complaining about ceos not being honest in the public risks that they see
usef-··on GPT-6 Sol and Luna
There's no difficulty in cancelling.
Page 1 of 12Next →