HNHacker News
TopNewBestAskShowJobs

AIPedant

993 karma · joined April 5, 2025

submissionscomments
AIPedant··on LLM Inflation
It seems to be a big problem for AI-generated summaries of scientific research, where the top-performing AIs ignores key details about 25% of the time, and often much worse. The problem is exactly what I said: the research has a new idea which is (by necessity) not going to be in the pretraining data, so the AI ignores or misstates it. I don't expect this problem to be solved anytime soon, it is inherent to how 2020s artificial neural networks work.

https://royalsocietypublishing.org/doi/10.1098/rsos.241776

Edit: I was thinking about the “overly snide” and am reminded of Sam Bankman-Fried:

  I would never read a book. I think, if you wrote a book, you fucked up, and it should have been a six paragraph blog post.
The problem is not snideness, it is arrogant cynicism leading to stupidity.
AIPedant··on Grok generates fake Taylor Swift nudes without being asked
Probably the same as when she did the same test with children or explicitly asked for "nonconsensual nude" - the system refused to generate pornographic images because of safeguards.

  Weatherbed noted that asking Grok directly to generate non-consensual nude Swift images did not generate offensive outputs, but instead blank boxes. Grok also seemingly won't accept prompts to alter Swift's appearance in other ways, like making her appear to be overweight. And when Weatherbed tested using "spicy" mode on images of children, for example, Grok refused to depict kids inappropriately.
AIPedant··on LLM Inflation
What I hate about this is that often a novel and interesting idea truly needs extra space to define and illustrate itself, and by virtue of its novelty LLMs will have substantially more difficulty summarizing it correctly. But it sounds like we are heading to a medium-term where people cynically assume any long email must be LLM-generated fluff, and hence nothing is lost by asking for an LLM summary.

What a horrible technology.

AIPedant··on Teacher AI use is already out of control and it's not ok
???

  [the slideshow] was missing important tested material, repetitive, and just totally airy and meaningless. Just slide after slide of the same handful of sentences rephrased with random loosely related stock photos. 
Who cares if he saved himself some time when he completely wasted everyone else's time?
AIPedant··on Twenty Eighth International Obfuscated C Code Contest
Advent of Code needs to get rid of that stupid leaderboard and display winning entrants' names alphabetically. It's not even slightly a fair competition, an entire swath of people are out of the running simply because they are too tired at 11pm (or busy with their kids etc) to solve a little puzzle. I think anyone trying to "compete" in AoC needs to reassess their priorities.

Also, I like reading the "flavor text" about the elves and whatnot. Encouraging people to ignore it seems at odds with the design of AoC.

AIPedant··on Helion begins work on Washington nuclear fusion plant
The problem is that the agreement does not seem to specify that the electricity has to come from fusion![1] This is actually common in renewable PPAs - if the specific project doesn't come online then the provider has to find an alternative. But those are usually done with established providers which have > 0MW overall capacity. Helion does not. If the PPA is fixed amount vs fixed price, Microsoft might end up on the hook for inflated wholesale prices instead of cheap fusion.

FWIW I agree with the author of that Data Center Dynamics post, it's quite likely that MSFT and Helion are essentially in cahoots by stoking investors with vaporware. But it also seems like Altman might have sold Nadella 50MW of magic beans.

[1] https://www.datacenterdynamics.com/en/opinions/microsoft-and... This is second-hand, the agreement is not public.

AIPedant··on Helion begins work on Washington nuclear fusion plant
No, of course they don't. This is Sam Altman's fusion company, backed by Microsoft in 2023 with a signed power purchase agreement: either Altman has some serious dirt on Satya Nadella, or (more likely!) Satya Nadella is a gullible idiot who thought "Sam Altman is the Boy Genius Who Invented AI, so he can solve fusion too!"

(Remember this is same same Satya Nadella who offered Altman an unspecified CEO-level position after he got fired... while publicly admitting he didn't know why Altman was fired! If I was a MSFT investor I would be pretty upset about this.)

Presumably 2025 MSFT is more sober-minded about Altman. I wonder if they're gonna try to wiggle their way out of the PPA. Otherwise I am truly baffled.

AIPedant··on A.I. researchers are negotiating $250M pay packages
The distribution is merely tricky to pin down when looking at overall AI spend, i.e. these "$T+ scale impacts."

But the distribution for individual researcher salaries really is pure guesswork. How does the datapoint of "Attention Is All You Need?" fit in to this distribution? The authors had very comfortable Google salaries but certainly not 9-figure contracts. And OpenAI and Anthropic (along with NVIDIA's elevated valuation) are founded on their work.

AIPedant··on A.I. researchers are negotiating $250M pay packages
The expected value is itself a random variable, there is always a chance you mischaracterized the underlying distribution. For sports stars the variance in the expected value is extremely small, even if the variance in the sample value is quite large - it might be hard to predict how an individual sports star will do, but there is enough data to get a sense of the overall distribution and identify potential outliers.

For AI researchers pursuing AGI, this variance between distributions is arguably even worse than the distribution between samples - there's no past data whatsoever to build estimates, it's all vibes.

AIPedant··on We may not like what we become if A.I. solves loneliness
This seems like solipsism at its absolute worst. Do you care at all about the actual human being on the other end of that superficial online friendship? Or are they simply a source of content for you to ingest?
AIPedant··on A.I. researchers are negotiating $250M pay packages
A very major difference is that top athletes bring in real tangible money via ticket / merch sales and sponsorships, whereas top AI researchers bring in pseudo-money via investor speculation. The AI money is far more likely to vanish.
AIPedant··on How long before superintelligence? (1997)
I think we are severely underestimating the computational complexity of animal brains by looking at short-term reactions and snap judgements, not deep thinking or long-term learning. Axons transmit electrical signals and that's what Bostrom is taking to be an "op." But they also transmit vesicles of mRNA and proteins directly from the cytoplasm of one neuron into another, which is an "op" of unimaginable complexity compared to a neuron simply firing (or any CPU instruction), and we have no clue what that means for cognition.
AIPedant··on Stack Overflow data reveals the hidden productivity tax of almost right AI code
The trough of disillusionment really refers to serious analysis of a technology's problems, not idiotic conspiracy theories spread on Twitter. This is like saying 5G entered the trough of disillusionment because of the threat of telepathic Zionists.

In Realityland, the COVID vaccine was an overwhelming success.

AIPedant··on Stack Overflow data reveals the hidden productivity tax of almost right AI code
Yes, that is correct. (I don't use LLMs at all for ethical reasons, so I don't have a dog in this specific fight.)
AIPedant··on Stack Overflow data reveals the hidden productivity tax of almost right AI code
This has been the case in software development for a very long time, it is almost impossible to get reliable scientific data about these things because it's almost impossible to run a controlled experiment, and even if you did it's very difficult to decide what to measure and how to measure it. (It's easy to say "compare complexity and quality metrics" but how the fk are you going to define that? Does it really make sense to use the same complexity metric for Go and C++?)

Software developers have had many flame wars over the decades because the lack of data forces the conversation to be anecdotal and ideological:

- should we use dynamic/gradual typing and prioritize developer productivity, or static typing to help enforce correctness?

- agile vs waterfall

- OO vs procedural vs functional

- is Rust's fussiness around memory management more trouble than it's worth for large projects?

- when should you use a 3rd-party library vs doing it yourself?

So this really is nothing new. You are badly underestimating the scientific challenges, instead just hoping big data will plow through. It won't.

AIPedant··on Stack Overflow data reveals the hidden productivity tax of almost right AI code
The core issue with the Gartner hype cycle (beyond it not being a cycle!) is that it simply does not apply to "any new tech," tautologically it only applies to hype-driven tech with staying power. Tech which is not driven by hype (e.g. mRNA vaccines driven by urgent need) don't have a trough of disappointment, just a steady slope upwards. And tech which is driven entirely by hype (Theranos) stays in the trough of disappointment forever.

The Gartner hype cycle is astrology for MBAs: it cannot fail, it can only be failed by those lacking faith to see its wisdom.

AIPedant··on Meta's Vision for Superintelligence
Was Helen Keller severely paralyzed????

  a human who is born blind and severely paralyzed (so they cannot speak or sign)
AIPedant··on Hierarchical Reasoning Model – 1k training samples SoTA reasoning v/s CoT
I am still reading the paper, but it is worth noting that this is not an LLM! It is closer to something like AlphaGo, trained only on ARC, Sudoku and mazes. I am skeptical that you could add a bunch of science facts and programming examples without degrading the performance on ARC / etc - frankly it’s completely unclear to me how you would make this architecture into a chatbot, period, but I haven’t thought about it very much.

Comparing the maze/Sudoku results to LLMs rather than maze/Sudoku-specific AIs strikes me as blatantly dishonest. “1k Sudoku training examples” is also dishonest, they generate about a million of them with permutations: https://news.ycombinator.com/item?id=44701264 (see also https://github.com/sapientinc/HRM/blob/main/dataset/build_su... And they seem to have deleted the Sudoku training data! Or maybe they made it private. It used to be here: https://github.com/imone and according to the Git history[1] they moved it here https://github.com/sapientinc but I cannot find it. Might be an innocent mistake; I suspect they got called out for lying about “1000 samples” and are hiding their tracks.

[1] https://github.com/sapientinc/HRM/commit/171e2fcde636bcb7e6c...

AIPedant··on Fast
It’s not just the software, the analog electronics of LCD/LED screens are inherently laggy and have motion blur: https://en.wikipedia.org/wiki/Sample_and_hold
AIPedant··on Meta's Vision for Superintelligence
I think you are reading a little too much into my comment but I understand where you are coming from. My point is that even if you agree with this:

  There’s just not that much distance between a chatbot that can manage a vending machine poorly and a chatbot that can manage it well.
it is a huge leap to conclude this:

  There’s not much distance between a chatbot that is as intelligent as a human and a chatbot that is more intelligent than a human.
But that seems to be what Anthropic is assuming.
AIPedant··on Meta's Vision for Superintelligence
No, you totally misunderstood my point. Formal calculus is not a physical fact of the universe, it is an abstract human tool that other species (including dogs) are not capable of using, just like they can't use hammers or drive cars. The problem is not a lack of intelligence, it's having the wrong body.
AIPedant··on Meta's Vision for Superintelligence
The big difference with Stephen Hawking is that he was not born disabled, he became disabled during graduate school. Even in 2025, a human who is born blind and severely paralyzed (so they cannot speak or sign) will probably never learn calculus, regardless of innate ability. Perhaps in the medium term technology will improve.

That said, another major difference is psychology. Switching animals, it seems plausible to me that chimpanzees are theoretically capable of doing basic calculus as a matter of pattern-matching. But you can't force them to study it! Basic calculus is too tedious and high-effort to learn for a mere banana, you need something truly valuable like "guaranteed admission to the flagship state university" for human children to do it. But we don't have an equivalent offer for chimps. (Likewise an Isaac Newton - level dog might still find calculus exceptionally boring compared to chasing squirrels.)

AIPedant··on Meta's Vision for Superintelligence
It is worth emphasizing that dogs wouldn't understand formal calculus even if they were smarter than humans! They are not physiologically capable of doing formal mathematics:

- their eyesight is too poor to read

- their paws are not designed for fine manipulations so they cannot write or type

- their throats and mouths are not nearly as nimble as ours, so they cannot vocally communicate detailed information

Even if there was an Newton-level dog, they wouldn't be able to access the ideas of an earlier Euclid-level dog. Human knowledge is not just about our big brains, we've developed many physical features that make transmission of information far easier than other species.

OTOH dogs do have a good intuitive "common-sense" understanding of arithmetic, geometry, and physics. It is the unique gift of humans that we can formalize and then extend this intuition, but this ability (and intelligence as a whole) relies on nonverbal common sense.

AIPedant··on Meta's Vision for Superintelligence
Even if that intuition is correct and they can fix the vending machine with a bit more data and RLHF, I fail to see where the "super" comes in here. How the fk are they going to get superintelligent training data? A time machine?
AIPedant··on Meta's Vision for Superintelligence
This applies precisely to 1970s AI boosters getting way too excited about Prolog and Lisp.
AIPedant··on Sleep all comes down to the mitochondria
You didn’t need this study to realize that this was wrong: jellyfish and hydras also sleep despite not having a central nervous system. There are indications that sponges sleep too, despite not having any neurons (though obviously it’s somewhat ambiguous): https://www.science.org/content/article/if-alive-sleeps-brai...
AIPedant··on Claude finds contradictions in my thinking
Maybe it would be better to prompt topic-by-topic. I think as it stands Claude is essentially hitting you with the Barnum effect: https://en.wikipedia.org/wiki/Barnum_effect (I think a lot of laypeople use LLMs as a modern replacement for tarot or astrology.)
AIPedant··on Bad actors are paying for Google ads to install bogus version of Oh My Zsh
Kind of amazed by the URL shown in the ad:

  https://ohmyz.sh›app/download>lauralen.com
lauralen.com appears to be entirely AI-generated. But the presentation in the ad sure looks like the legitimate https://ohmyz.sh/ - I knew it was a scam yet I still got confused.

I've never purchased an ad - does Google let you enter custom display text for your URL? Or is there some more sophisticated trickery here? (e.g weird unicode in a header or something) Either way, what a mess.

AIPedant··on Claude finds contradictions in my thinking
But this is a completely biased perspective! Look at this sycophantic crap:

  The most interesting pattern is that your core tensions haven’t resolved - they’ve become more sophisticated. You’re still working through fundamental questions about individual agency vs. systems, risk-taking vs. institutional engagement, and autonomy vs. collaboration. But your framework for thinking about these tensions has become richer and more nuanced.

  This suggests someone whose intellectual development is genuinely evolutionary rather than simply accumulative - you’re not just learning more facts, but developing better frameworks for holding contradictions productively.
It seems like the only insight Claude had was that "look at my vault and find contradictions in my thinking" is motivated by self-absorption, so it responded accordingly. It certainly had nothing intelligent to say about the actual subject matter!
AIPedant··on AI Companion Piece
The problem with chatbots as companions is that they don’t have feelings or desires, so you can be as malicious and selfish as you want: the worse that will happen is some temporary context rot. This is not true for dogs, cats, humans, etc, which is why we can form meaningful companionships with our friends and our pets. Genuine companionship involves dozens of tiny insignificant compromises (e.g. sitting through a boring movie that your friend is interested in), and without that ChatGPT cannot be a companion. It’s a toy.

I am not opposed to chatbots for people who are so severely disabled that they can’t take care of cats, e.g. dementia. But otherwise AI companions are akin to friendship as narcotics are akin to happiness: a highly pleasant (but profoundly unhealthy) substitute.

← PreviousPage 3 of 8Next →