HNHacker News
TopNewBestAskShowJobs

andxor

850 karma · joined November 24, 2016

submissionscomments
andxor··on U.S. appeals court upholds designation of Anthropic as supply chain risk
What a sophisticated opinion.
andxor··on Claude Opus 5.5
Not necessarily, if Haiku performs better than Luna.
andxor··on We must pace the frontier
Fable is still better than Astra and Fable is not the best model Anthropic has.
andxor··on We must pace the frontier
Really? Anthropic has the strongest models and it's in the best position to begin RSI and win the race. A pause would favor competitors.
andxor··on I Changed My License
It's a quixotic crusade, in perfect European style.
andxor··on GPT-6 Astra on robot arms
Which benchmarks? Only the ones OpenAI cherry-picked.

It debuted as ~same score as Sol on Artificial Analysis. People couldn't accept it so they had to change the formula.

The model is a big step forward only in desktop use and 3D. That's impressive, but for software engineering, Fable is still in a league of its own.

andxor··on Discovery of a new OpenAI agent message board
These are often experimental models that haven't undergone full safety testing. Not comparable to publicly accessible models.
andxor··on Discovery of a new OpenAI agent message board
It's patently obvious at this point.
andxor··on GPT-6 Astra
This is not saying much. Opus 4.8 is ancient history.
andxor··on GPT-6 Astra
That's not my experience and I suspect it's not most people's experience. Out of curiosity, what's the hardest task you tried?
andxor··on GPT-6 Astra
> Sol is so much better than Fable 5

I'm genuinely so confused when people say this with a straight face. Are you talking about coding? Desktop use? Prose? Or something else?

Sol is a much smaller models and it shows. It often misses the forest for the trees.

andxor··on GPT-6 Astra
Artificial Analysis just published their aggregate score (61).

Still below Fable 5, let alone Fable 5.1.

EDIT: This is suspiciously low. Calls the relevance of existing benchmarks into question.

andxor··on GPT-6 Astra
> Performance is significantly higher than Fable 5.1

That's not clear. Need to see independent benchmarks first.

andxor··on Google Has Removed MV2 Extensions from the Chrome Web Store, Including UBO
What is a concrete example of an ad you are seeing that it can't remove?
andxor··on Google Has Removed MV2 Extensions from the Chrome Web Store, Including UBO
uBlock Origin Lite
andxor··on Moderna reports first positive Phase 3 for mRNA neoantigen therapy in melanoma
That says more about you than Moderna:

https://totalrealreturns.com/n/MRNA,SPY

MRNA has vastly outperformed the market since IPO.

andxor··on Moderna reports first positive Phase 3 for mRNA neoantigen therapy in melanoma
You know many stocks have done better than the SP500, right?
andxor··on Moderna reports first positive Phase 3 for mRNA neoantigen therapy in melanoma
Oh yeah? Show me your discounted cash flow analysis.
andxor··on MathCode, Mathematical Coding Agent
> hooking up slop to slop is just unlikely to produce anything valuable

Do you have a formal proof of that?

andxor··on ChatGPT lost 22 points of web share in a year
Really? This just tells me you haven't tried it in a long time. I'd argue it's a much better experience for the average user.
andxor··on GLM-5.3: Frontier coding with emergent cyber capabilities
Do you work at OpenAI?
andxor··on GLM-5.3: Frontier coding with emergent cyber capabilities
There's no shortage of extremely valuable problems to solve.
andxor··on GLM-5.3: Frontier coding with emergent cyber capabilities
Are you assuming it won't?
andxor··on GLM-5.3: Frontier coding with emergent cyber capabilities
Fable finished training 6+ months ago.

At this point, Anthropic only needs to release models to the public when the competition forces them to.

OpenAI also has a better model (Astra) that they haven't released yet.

andxor··on U.S. economy lost 23,000 jobs in July, a sudden reversal
You're getting downvoted but you're exactly right
andxor··on U.S. economy lost 23,000 jobs in July, a sudden reversal
This is why engineers don't make money trading markets
andxor··on Google Discloses $94.1B in SpaceX Stock, Marking 6% Stake
Much better business instinct that the average poster on HN
andxor··on Google Discloses $94.1B in SpaceX Stock, Marking 6% Stake
Google's return of invested capital has averaged 32% since IPO.

This may be a narrative violation on HN, but the quality of Google's management is exceptional.

andxor··on Kimi K3: Open Frontier Intelligence
I'm not defending Dario. That's not the point I'm making.
andxor··on Kimi K3: Open Frontier Intelligence
Yeah, I think the answer is somewhere in between. Anthropic has been a few months ahead of everybody in terms of internal capabilities.

So it's not 6 months but it's also not a few weeks.

Page 1 of 4Next →