HNHacker News
TopNewBestAskShowJobs

mrbungie

1,466 karma · joined September 25, 2017

submissionscomments
mrbungie··on Claude Haiku 5.5
Yep. I was looking at the prices of lower tier models a few weeks ago for zero/few shot tasks (pre Jev) and Haiku rates just didn't make sense at all. I ended up using 5.6-luna.

Good to know that is going back to being an actual option from perf/price perspective.

mrbungie··on Claude is a Contrarian
Since LLMs are different between each other (and even version of the same family) I'd expect the differences to be explained by each lab idiosincratic biases in the post-training phase rather than by internet corpora characteristics during the pre-training phase.
mrbungie··on Why don't machine learning research agents overfit?
Ah, over-the-top larger-than-life LLM-isms, they are really funny when you see them in a company blog, but they are vomitive when it's your coworker copy-pasting it and insisting you on reading it.
mrbungie··on More questions about whether researchers can trust OpenAI with unpublished math
It is an extraordinary claim, it is a millenium prize problem after all. We don't even know, even if there was no copying, how much human involvement there was in the result.
mrbungie··on OpenAI's Sam Altman says it would be 'ill-advised' to go public in 2026
> Wall street is less forgiving.

We've seen multiple times how "Wall street" is an exit for insiders, with retail investors left holding the bag.

mrbungie··on More questions about whether researchers can trust OpenAI with unpublished math
Extraordinary claims require extraordinary evidence.

An article post that wouldn't even amount to a white paper + the LEAN proof is not evidence of how they got to produce it.

mrbungie··on AI could kill all humans in next decade, warn experts
I mean, a disaster could happen even with properly human-aligned AI. In a "AI makes the shots" world with undecipherable and fast feedback loops, which is the one we are seemingly heading towards, a string of short-sighted decisions made by AI agents may be the only thing separating us from extinction.
mrbungie··on AI could kill all humans in next decade, warn experts
You are not seriously suggesting that nature and AI-accelerated gain-of-function research are similar in regards to risks, are you?
mrbungie··on How An AI math breakthrough ignited a controversy
They represent different AI usage patterns. OpenAI wants everyone to believe that it was done with a practically autonomous network of thousands of agents with little to no human intervention for 88 hours, while Buckmaster/Alpöge were using AI in a more guided way for months. If OpenAI actually used anything from Buckmaster/Alpöge work they would be misleading the public.
mrbungie··on Tao: Open math problems being non-renewably mined by AI
Hopefully. We'll need to wait until mathematicians confirm how easy it is to digest whatever GPT did.
mrbungie··on Tao: Open math problems being non-renewably mined by AI
For sure, but this was supposedly ~18 million dollars of compute, afaik 100 pages paper / lean proof and only god knows how many bytes of chat interactions + thought traces. Scale matters.
mrbungie··on Tao: Open math problems being non-renewably mined by AI
Probably an AI-written Lean proof is very different to how a human would write it, and some may say it's more like mathy neuralese. For sure it works but it is not human-friendly and needs to be transformed into something more readable and digestible to be able to extract insights from it.

Not that different from when trying to read an out-of-control vibe coded codebases, or an sloppy AI long email that someone may send you at 9 AM.

mrbungie··on On the Navier–Stokes Millennium Prize Problem
In what world a tweet and a screenshot of a private convo are evidence of good faith? Plain sociopathic behavior.
mrbungie··on On the Navier–Stokes Millennium Prize Problem
We are not in an everything-has-an-API world yet, and it'll for sure take some time to get there.
mrbungie··on On the Navier–Stokes Millennium Prize Problem
> 1) The article is written in a way that states clearly they threw a lot of compute at the problem. In api cost millions of dollars.

Did I say otherwise?

> 2) Millenium Problems have been the goal every AI company wanted to achieve since their diffusion, all companies have thrown a lot of resource to solve these problems, as they are very famous and scientists spent a lot of time trying to solve them. The first company to solve it will remain in history, despite all of you finding excuses about it.

I know, but I don't know how that relates to my point, which is about the way they are doing it.

mrbungie··on On the Navier–Stokes Millennium Prize Problem
Of course, as any drama, it has been developing into a lot more but the main motivation for OpenAI has been about winning that battle.
mrbungie··on On the Navier–Stokes Millennium Prize Problem
Of course AI was involved, you'd expect most mathematicians and researchers to use AI nowadays. This drama is about AI achieving impressive outcomes with little to no human intervention, as that would be signalling AGI.
mrbungie··on On the Navier–Stokes Millennium Prize Problem
They are highly capable, no doubt about that, but:

1) We don't really know how they arrived to this result except that they had a lead and that they threw millions of compute at the problem. The article is written in a way that makes you believe that it was just an agent loop with little human intervention, but without any evidence.

2) If the threats are to be believed, it is concerning how far they are willing to go to show how capable the model is. One would think their products and credibility would be enough to speak for themselves.

mrbungie··on Claude Opus 5
Long-term, sure. Short-term it is going to cause a lot of suffering.
mrbungie··on Better Models: Worse Tools
Just add a --verbose flag that shows the stacktrace when there is an error. Then add a footer message when an error appears in non-verbose mode that invites the user/agent to use --verbose to get the full picture.

It obviously may end up in thousands of tokens burned through though (you can also fix that adding different levels of verbosity), but hopefully errors are not common.

mrbungie··on Claude Sonnet 5
From what I gather from GPs upper post: Technical debt, skill atrophy, delusions of grandeur about one's own abilities / psychosis.
mrbungie··on U.S. government will decide who gets to use GPT-5.6
You don't need SOTA-level LLMs to create value with AI. Hell, you can build good solutions with a simple small finetuned models.

> When models are good, expectations are adjusted accordingly to deliver things on par with the whole industry, you can't just say, I have built my own Intel Pentium II, now I will try to use it to compile Electron App and run 3DS Max there.

I know you are taking your analogy to its breaking point but it really depends on what you are doing. I know people that use 10+-year old thinkpads and they do just fine.

mrbungie··on Apple raises prices of MacBooks, iPads
That's convenient accounting. The reality is that they can't stop training since they risk losing customers if they do so. So they shouldn't factor it out of profitability analysis.
mrbungie··on Apple announces significant price increases for MacBooks, iPads, more
Ah, come on. I remember the scalping of GPUs due to crypto-mining and then all the things Nvidia did to market segment crypto out of the regular (gaming) consumer space. AI is much worse because the scale is OOM greater, but crypto/blockchain effects on the market weren't harmless either.
mrbungie··on Apple raises prices of MacBooks, iPads
How much money does that revenue cost though? If I had to steel-man GPs argument I'd ask for profits rather than revenues.
mrbungie··on Anthropic/OpenAI may be spending more than $1000 for every $100 you pay them
That + when retail investors are the ones holding the bag.
mrbungie··on Amazon scraps AI leaderboard to stop workers chasing usage scores
MBAs are simply unable to learn this.
mrbungie··on Shunning AI is the human choice
It really depends on their environment. Not every city is a car-first city.
mrbungie··on Shunning AI is the human choice
AI as a tech is fine. But disliking it and the social/economic effects around it is fine too, people should be allowed to feel however they want to feel about certain techs and situations.

To recommend people to suck it up is not the answer I wish in the society I want to live in.

mrbungie··on Gemini 3.5 Flash
> I know artificial analysis quite well as the gold standard in llm evals.

I also know them, but it took me a while to realise you were publishing their data in that table. I don't think it was clear.

> The age is important because new techniques keep being developed and so it is a very rough indicator of the size/cost/efficiency trade-off.

Yes but you are already including the name of the model, your potential public for the table already know about model's release history and therefore each model's age, at least roughly.

Page 1 of 19Next →