HNHacker News
TopNewBestAskShowJobs

kzrdude

12,545 karma · joined May 15, 2011

I try to play music. And I actually believe the free software stuff.
submissionscomments
kzrdude··on LeCun has "zero concerns" about AI wiping out humanity, recent "rogue" incidents
I don't think it's about that. It's not about the step by step plan for the murder. It's about - how does "the AI" do it? Do we give it access to our world, or do we give it a body, so that it can take its own physical actions?
kzrdude··on LeCun has "zero concerns" about AI wiping out humanity, recent "rogue" incidents
I like you, I think very similarly. Humans are "like rats": we can live almost anywhere, we'll find a way to survive.

But that just means we won't all be wiped out. We need to understand when discussing global issues, such as this or like climate change that it's about prosperity and quality of life. We're trying to plan for a good life (for all people?).

kzrdude··on Nicholas Polson has authored 258 academic papers in 2026 so far
That's an interesting point and I wonder how far it has been proven. I guess it is validated by the present. Scandinavian cultures had laws and communities with courts all without writing, they had a storytelling and law-reciting culture instead.
kzrdude··on Nicholas Polson has authored 258 academic papers in 2026 so far
Speaking is a culture that has served us well for 300 000 years or so. To note, it has been used for both good and bad..
kzrdude··on Gravity seems holographic. What does that mean for reality?
Fortunately, it's very simple to navigate between pdf, html, and abstract, just put "abs" in the URL where it says pdf and so on.
kzrdude··on Rising sea destroys homes, erases beaches in California
The topographic map says, the Bay Area is just around the first bay, there's a whole second bay opening up behind it (Sacramento area valley) which could flood with enough sea level rise.
kzrdude··on Claude Opus 5.5 Intelligence, Performance and Price Analysis (Max)
What's more valuable than a good benchmark? IMO a benchmark that has been run against very many competitors and versions. Collecting data has something going for it, and it's up to the readers to interpret and make the best use out of it.
kzrdude··on OpenAI GPT–6 Astra breaks Enigma message that has resisted solution since 2005
Looks like Q is used as an abbreviation for CH
kzrdude··on Pirate Face Rescues LLM Models from Deletion
Typical example to show that people who want to be angry will be angry, even for pointless things.
kzrdude··on English: A vs. An
"a history lesson" and "an historic event" have different stress in the history area (or at least they can have, depending on the phrase context), and I think that explains it.

Saying "a historic event" requires you to pause and stress the start of historic (a + glottal stop + "his" + ...). It's a natural phrasing pattern to have the initial part of historic unstressed, and then the h is effectively dropped even if you're not speaking an h-dropping dialect; this way "an historic event" follows as correct.

We simplify a lot of sounds in unstressed parts of words, but because we know and "see" the spelling, our brains can keep the illusion of the whole word going for us.

kzrdude··on C++26: Trivial infinite loops are no longer undefined behaviour
Well you leave the C++ realm (execution model), as you should with UB and it depends on implementation. The implementation of the compiler was such that the two functions are placed after each other in the machine code; and if the first function doesn't return, then you continue executing into the code for the next function.
kzrdude··on After Math
This is tricky, because we really want language-independent training of skills. We know that self-play type of reinforcement learning is incredibly effective when possible. But at the same time, they are our tools - so we need supervised language training for this reason? It's possible that training just needs to be rebalanced so that RL with rewards is balanced with rounds of language adjustment. And to really make that happen, benchmarks need to score the models on that.
kzrdude··on After Math
We need to recognize this as a failure in training. It did some useful stuff but it can be much better. A training signal is likely missing.
kzrdude··on Eating Fruit Skins
Not only in America
kzrdude··on Navier-Stokes Announcement
But if I remember correctly, he gained recognition for his achievement rather quickly after posting.
kzrdude··on Navier-Stokes Announcement
If we read the link, it has a section called Gold Standard: comparator and external checkers, and comparator is how OpenAI has gone about checking their lean proofs.
kzrdude··on A misalignment of AI in mathematics
I wonder what Demis Hassabis thinks about this. I thought he cared a lot about mathematics.
kzrdude··on A misalignment of AI in mathematics
You should read his comment from wednesday here: https://terrytao.wordpress.com/2026/09/07/finite-time-blowup...

And yes, at one point it gives you right about them taking advantage of him. He sat down for an interview and found later that they just used the "best clips" out of it for an OpenAI ad.

kzrdude··on A misalignment of AI in mathematics
Well OpenAI took one step on the back foot at least, withdrawing from sponsoring this math hackathon event https://xcancel.com/danintheory/status/2098125701782372640
kzrdude··on More questions about whether researchers can trust OpenAI with unpublished math
My university has an agreement with Microsoft copilot. We can log into copilot in many ways, and it's only if you log in the correct way that you get the "Enterprise Data Protection" copilot version, with a green shield symbol. There are many ways to go wrong here!
kzrdude··on Astra for Coding: Why Are We Doing This Again?
It seems like we're not supposed to care about the code quality then? I guess that's the compilers argument. But I'm not ready to give up the code just yet.. These LLMs don't even have a stable interface, they change every few months in how they interpret our prompts and tasks.
kzrdude··on Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra
They have plausible deniability on that one: not making any profit.
kzrdude··on Cognition's SWE-2 achieves 92.8 on Terminal-Bench 2.1
Is that something we have credible evidence for? Do they serve a better model when artificialanalysis (the benchmark site) is making the requests, and so on?
kzrdude··on Stockfish 19
Interesting, and if you don't mind, where do we put humans (and superhumans like Magnus Carlsen)? I think they have a heavy evaluation function and do shallower search.
kzrdude··on Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra
There's this whole discussion going on about agents being more independent now. They don't follow instructions so well, they continue until the problem is done (sometimes too long), they don't ask the user for feedback.

That is a kind of benchmaxing: they are made to complete benchmarks tasks and one-offs well, and no longer work well in tandem with the user.

Regardless what you call it, it's a divergence between what the power user wants and what the model developers want, I think.

kzrdude··on More questions about whether researchers can trust OpenAI with unpublished math
But the very fact that you go to "chatgpt.com" and write to them; "Dear Diary, today I thought.."; there is no reason they would not receive and process your data, unless explicitly promising not to (which also requires us to trust them).

The fundamental rule in this case is that if we offload our data to a cloud provider we can assume they read it, if they can, unless they promised very clearly they will not.

kzrdude··on DeepSeek launching v4.1 flash cheaper and more capable than v4 pro
And just a day later we have the tech report available for V4.1: https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash/blob/...
kzrdude··on Stockfish 19
Isn't one problem that it's hard to determine what equivalent compute is, for CPU search vs a neural net based engine like AZ or Leela?
kzrdude··on DeepSeek v4.1 Flash
V4 Flash was one of the big events of this year, and its already retired and replaced by V4.1 Flash.
kzrdude··on OpenAI have no mathematicians capable of understanding what they put out
The main claim, that "they do not seem to have any mathematicians capable of understanding what they put out", was also corroborated by Sebastien (OpenAI) who explained they don't have any experts on Navier-Stokes.
Page 1 of 34Next →