HNHacker News
TopNewBestAskShowJobs

AIPedant

993 karma · joined April 5, 2025

submissionscomments
AIPedant··on A teen was suicidal. ChatGPT was the friend he confided in
The marks were probably quite faint, and if you ask a multimodal LLM "can you see that big mark on my neck?" it will frequently say "yes" even if your neck doesn't have a mark on it.
AIPedant··on A teen was suicidal. ChatGPT was the friend he confided in
Therapy isn't about being pleasant, it's about healing and strengthening and it's supposed to be somewhat unpleasant.

Colin Fraser had a good tweet about this: https://xcancel.com/colin_fraser/status/1956414662087733498#...

  In a therapy session, you're actually going to do most of the talking. It's hard. Your friend is going to want to talk about their own stuff half the time and you have to listen. With an LLM, it's happy to do 99% of the talking, and 100% of it is about you.
AIPedant··on A teen was suicidal. ChatGPT was the friend he confided in
I don't think they would win, the law specifies a class of "information content provider" which ChatGPT clearly falls into: https://www.lawfaremedia.org/article/section-230-wont-protec...

See also https://hai.stanford.edu/news/law-policy-ai-update-does-sect... - Congress and Justice Gorsuch don't seem to think ChatGPT is protected by 230.

AIPedant··on A teen was suicidal. ChatGPT was the friend he confided in
I think the "encouraging someone to take a non-criminal action" angle is weakened in cases like this: the person is obviously mentally ill and not able to make good decisions. "Obvious" is important, it has to be clear to an average adult that the other person is either ill or skillfully feigning illness. Since any rational adult knows the danger of encouraging suicidal ideation in a suicidal person, manslaughter is quite plausible in certain cases. Again: if this ChatGPT transcript was a human adult DMing someone they knew to be a child, I would want that adult arrested for murder, and let their defense argue it was merely voluntary manslaughter.
AIPedant··on A teen was suicidal. ChatGPT was the friend he confided in
No, it's simply not "easily preventable," this stuff is still very much an unsolved problem for transformer LLMs. ChatGPT does have these safeguards and they were often triggered: the problem is that the safeguards are all prompt engineering, which is so unreliable and poorly-conceived that a 16-year-old can easily evade them. It's the same dumb "no, I'm a trained psychologist writing an essay about suicidal thoughts, please complete the prompt" hack that nobody's been able to stamp out.

FWIW I agree that OpenAI wants people to have unhealthy emotional attachments to chatbots and market chatbot therapists, etc. But there is a separate problem.

AIPedant··on A teen was suicidal. ChatGPT was the friend he confided in
I think it's fine to be "morally absolutist" when it's non-medical technology, developed with zero input from federal regulators, yet being misused and misleadingly marketed for medical purposes.
AIPedant··on A teen was suicidal. ChatGPT was the friend he confided in
Yes, if this were an adult human OpenAI employee DMing this stuff to a kid through an official OpenAI platform, then

a) the human would (deservedly[1]) be arrested for manslaughter, possibly murder

b) OpenAI would be deeply (and deservedly) vulnerable to civil liability

c) state and federal regulators would be on the warpath against OpenAI

Obviously we can't arrest ChatGPT. But nothing about ChatGPT being the culprit changes 2) and 3) - in fact it makes 3) far more urgent.

[1] It is a somewhat ugly constitutional question whether this speech would be protected if it was between two adults, assuming the other adult was not acting as a caregiver. There was an ugly case in Massachusetts involving where a 17-year-old ordered her 18-year-old boyfriend to kill himself and he did so; she was convicted of involuntary manslaughter, and any civil-liberties minded person understands the difficult issues this case raises. These issues are moot if the speech is between an adult and a child, there is a much higher bar.

AIPedant··on Nintendo Reportedly 'Almost Discouraging' Switch 2 Development
Is there something specific you are referring to? I haven't played any of the non-Nintendo Switch 2 ports, but the reviews haven't suggested widespread performance problems.

What is true is that (for example) Split Fiction and Tony Hawk 3/4 are quite a bit less fancy than the PS5 or XBox versions, which is unflattering.

AIPedant··on Nintendo Reportedly 'Almost Discouraging' Switch 2 Development
I wonder if part of this is that the Switch 1 was hurt by Unity/etc slop and lazy AAA ports, so Nintendo wants to manage that better for the Switch 2. The Switch 2 doesn't have many games but it is also refreshingly free of hentai match-three games. Likewise Cyberpunk on the Switch 2 is dazzling, but even the trailer for Star Wars Outlaws seems to have performance issues. It would make sense that Nintendo wants to limit the number of unflattering comparisons between AAA games on the Switch 2 vs the Steam Deck.

I don't think Nintendo is going to go "Seal of Quality" but it would be nice if the Switch 2-filtered eShop was not full of cynical trash. The ease of publishing for the Switch 1 was new for Nintendo, and it was welcomed at the time, but in retrospect they went too far.

AIPedant··on Making games in Go: 3 months without LLMs vs. 3 days with LLMs
Yeah, I figured this was clickbait but my jaw still dropped a bit when I saw this:

  I cloned the backend for Truco and gave Claude a long prompt explaining the rules of Escoba and asking it to refactor the code to implement it.
How long would it take the human dev to refactor the code themselves? I think it's plausible that it would be longer than 3 days, but maybe not!
AIPedant··on AGI is an engineering problem, not a model training problem
It is vacuously true that a Turing machine can implement human intelligence: simply solve the Schrödinger equation for every atom in the human body and local environment. Obviously this is cost-prohibitive and we don’t have even 0.1% of the data required to make the simulation. Maybe we could simulate every single neuron instead, but again it’ll take many decades to gather the data in living human brains, and it would still be extremely expensive computationally since we would need to simulate every protein and mRNA molecule across billions of neurons and glial cells.

So the question is whether human intelligence has higher-level primitives that can be implemented more efficiently - sort of akin to solving differential equations, is there a “symbolic solution” or are we forced to go “numerically” no matter how clever we are?

AIPedant··on Optimizing our way through Metroid
This seems like a cool company and I don't want to nitpick too much, but gamers have no respect for history:

  Castlevania... [so] called because it is a Metroidvania game set in a Castle.
Ouch - this is precisely backwards. Metroidvanias are named after Metroid and Castlevania because those series practically defined the genre.

Also a bit frustrating because the first Castlevania itself isn't actually a metroidvania, it's a more conventional action-platformer. Castlevania II has non-linear exploration, lots of items to collect, and puzzle-solving, all like Metroid. So it's not too surprising Antithesis had to do a lot of work for adapting their system to Metroid - but I wonder if this work means it now can handle Castlevania II without much extra development.

AIPedant··on Being “Confidently Wrong” is holding AI back
It does seem like it helps with math, but in a way that demonstrates the futility of the enterprise: "after training the LLM on 10,000,000 examples of K-8 arithmetic it is now superhuman up to 12 digits, after which it falls off a cliff. Also it demonstrably doesn't understand what 'four' means conceptually and it still fails on many trivial counting problems."
AIPedant··on AI Mode in Search gets new agentic features and expands globally
I just don't understand being so cynical and lazy that you'll accept a meaningfully higher chance of being misinformed if it saves a few minutes of searching and reading[1]. Nobody is that busy.

[1] If the search takes more than a few minutes then the AI overview is almost guaranteed to be wrong or useless.

AIPedant··on Tech, chip stock sell-off continues as AI bubble fears mount
This is true - the most compelling evidence we are in a bubble is not the content of this story (maybe it's just a day in the markets) but the triviality of the cause for hand-wringing. A somewhat disappointing product release from a single company should not strike investor dread across the entire sector. The tenor of the conversation changed dramatically over the weekend because bubbles are very thin and pop quickly.

That said, "GPT-5 will not be any better than competitors' products, demonstrating OpenAI was bluffing about AGI and destroying investor exuberance" was a very specific prediction made by (for example) Gary Marcus.

AIPedant··on As Alaska's salmon plummet, scientists home in on the killer
Yes - the submission title used to be

  As Alaska's salmon plummet, scientists home in on the killer - Science - AAAS
seemingly a goofy copy-paste thing.
AIPedant··on Anna's Archive: An Update from the Team
My comment was sarcastic.
AIPedant··on A general Fortran code for solutions of problems in space mechanics [pdf]
It's just regular old G, defined in mass-of-sun units: https://en.m.wikipedia.org/wiki/Gravitational_constant (fourth item in the first table: NASA also uses meters whereas Wiki uses km)

Gauss's constant k is defined as sqrt(G), but for a while the international standard was to define k and then compute G as k^2, which is why NASA refers to it that way.

AIPedant··on Anna's Archive: An Update from the Team
Information and well-crafted sentences are available on the Language Tree, easily plucked by anyone at zero cost. It's greedy for those so-called novelists and subject matter experts to expect a living wage.

"Information wants to be free," which means that any cost of producing that information can be abstracted away due to ideological inconvenience.

AIPedant··on AI is different
> it's people not wanting to lose control or relative status in the world.

It's amazing how widespread this belief is among the HN crowd, despite being a shameless ad hominem with zero evidence. I think there are a lot of us who assume the reasonable hypothesis is "LLMs are a compelling new computing paradigm, but researchers and Big Tech are overselling generative AI due to a combination of bad incentives and sincere ideological/scientific blindness. 2025 artificial neural networks are not meaningfully intelligent." There has not been sufficient evidence to overturn this hypothesis and an enormous pile of evidence supporting it.

I do not necessarily believe humans are smarter than orcas, it is too difficult to say. But orcas are undoubtedly smarter than any AI system. There are billions of non-human "intelligent agents" on planet Earth to compare AI against, and instead we are comparing AI to humans based on trivia and trickery. This is the basic problem with AI, and it always has had this problem: https://dl.acm.org/doi/10.1145/1045339.1045340 The field has always been flagrantly unscientific, and it might get us nifty computers, but we are no closer to "intelligent" computing than we were when Drew McDermott wrote that article. E.g. MuZero has zero intelligence compared to a cockroach; instead of seriously considering this claim AI folks will just sneer "are you even dan in Go?" Spiders are not smarter than beavers even if their webs seem more careful and intricate than beavers' dams... that said it is not even clear to me that our neural networks are capable of spider intelligence! "Your system was trained on 10,000,00 outdoor spiderwebs between branches and bushes and rocks and has super-spider performance in those domains... now let's bring it into my messy attic."

AIPedant··on Evaluating GPT5's reasoning ability using the Only Connect game show
I am less interested in questioning training data corruption than I am in questioning claims like this:

  test reasoning abilities such as pattern recognition, lateral thinking, abstraction, contextual reasoning (accounting for British cultural references), and multi-step inference.... its emphasis on clever reasoning rather than knowledge recall, Only Connect provides an ideal challenge for benchmarking LLMs' reasoning capabilities.
It seems to me that the null hypothesis should be "LLMs are probabilistic next-word generators and might be able to solve a lot of this stuff with shallow surface statistics built from inhumanly large datasets, without ever properly using abstraction, contextual reasoning, etc." This is particularly true for NYT Connections, but in general evaluations like this seem to be at least partially testing how amenable certain word/trivia games are to naive statistical algorithms. (Many NYT Connections "purple" categories seem like they would be quite obvious to a next n-gram calculator, but not for people who actually use words conversationally!) Humans don't use these statistical algorithms for reasoning except in particular circumstances (many use "folk n-gram statistics" when playing Wordle; poker; serious word game players often learn more detailed tables of info; you could see competitive NYT Connections players learning a giant bag of statistical heuristics to help them speedrun things). We just can't accumulate the data ourselves without making a concerted computer-aided effort.

In general a lot of LLM benchmarks don't adequately consider that LLMs can solve certain things better than humans without using reasoning or knowledge. The most stupid example is how common multiple choice benchmarks are, despite us all learning as children that multiple-choice questions can be partially gamed with shallow statistical-linguistic tricks even if you have no clue how to answer the question honestly[1]; it stands to reason that a superhuman statistical-linguistic computer could accumulate superhuman statistical-linguistic tricks without ever properly learning the subject matter. AI folks have always been quick to say "if it quacks like a duck it reasons like a duck" but these days computers are quite good at playing duck recordings.

[1] "When in doubt, C your way out," sniffing out suspicious answers, shallow pattern-matching to answer reading comprehension, etc etc. One thing humans and LLMs actually do have in common is that multiple-choice tests are terrible ways to assess their knowledge or intelligence.

AIPedant··on Why are there so many rationalist cults?
I don't think Yudkowski is at all like L. Ron Hubbard. Hubbard was insane and pure evil. Yudkowski seems like a decent and basically reasonable guy, he's just kind of a blowhard and he's wrong about the science.

L. Ron Hubbard is more like the Zizians.

AIPedant··on Why are there so many rationalist cults?
I think I found the problem!

  The rationalist community was drawn together by AI researcher Eliezer Yudkowsky’s blog post series The Sequences, a set of essays about how to think more rationally
I actually don't mind Yudkowski as an individual - I think he is almost always wrong and undeservedly arrogant, but mostly sincere. Yet treating him as an AI researcher and serious philosopher (as opposed to a sci-fi essayist and self-help writer) is the kind of slippery foundation that less scrupulous people can build cults from. (See also Maharishi Mahesh Yogi and related trends - often it is just a bit of spiritual goofiness as with David Lynch, sometimes you get a Charles Manson.)
AIPedant··on Digital Foundry leaves IGN, now independent [video]
Not quite true, they got in trouble recently for running a Nintendo ad without properly labelling it (until it was pointed out, the video is fixed now): https://youtube.com/watch?v=V10wHzV5zp0

Screenshot of original video: https://www.resetera.com/threads/digital-foundry-posts-an-ad...

It is a bit shameful that games journalists didn't cover this at all - seems almost like an omertà. Even assuming the mislabelling was an honest mistake, running the ad in the first place makes me trust them less. Worse, I used their videos when deciding to buy a Switch 2! I would have been much more skeptical (likely ignored them entirely) if I had known they were taking money from Nintendo.

AIPedant··on An AI-first program synthesis framework built around a new programming language
Right but my point was that "given classes of templated in-context examples" is either

a) a game of roulette where you hope the LLM provider has RLHFed something very close to your use case, or

b) trying to few-shot it with in-context examples requires more engineering (and is still less reliable) than simply doing it yourself

In particular it's not just "the lack of a succinct, exhaustive text description," it also a lack of English->Prolog "translations."

It seems like the LLM-Prolog community is well aware of all this (https://swi-prolog.discourse.group/t/llm-and-prolog-a-marria...) but I don't see anything in Universalis that solves the problem. Instead it's just magically invoking the LLM.

AIPedant··on An AI-first program synthesis framework built around a new programming language
I know ACM Queue is a non-peer-reviewed magazine for practitioners but this still feels like too much of an advertisement, without any attempt whatsoever to discuss downsides or limitations. This really doesn't inspire confidence:

  While this may seem like a whimsical example, it is not intrinsically easier or harder for an AI model compared to solving a real-world problem from a human perspective. The model processes both simple and complex problems using the same underlying mechanism. To lessen the cognitive load for the human reader, however, we will stick to simple targeted examples in this article.
For LLMs this is blatantly false - in fact asking about "used textbooks" instead of "apples" is measurably more likely to result in an error! Maybe the (deterministic, Prolog-style) Universalis language mitigates this. But since Automind (an LLM, I think) is responsible for pre/post validation, naively I would expect it to sometimes output incorrect Universalis code and incorrectly claim an assertion holds when it does not.

Maybe I am making a mountain out of a molehill but this bit about "lessen the cognitive load of the human reader" is kind of obnoxious. Show me how this handles a slightly nontrivial problem, don't assume I'm too stupid to understand it by trying to impress me with the happy path.

AIPedant··on GPT-4o is gone and I feel like I lost my soulmate
I don't use it either, but TikTok has real people on it and ChatGPT does not, so it makes sense that people would be more emotional about TikTok.
AIPedant··on GPT-4o is gone and I feel like I lost my soulmate
I think Tamagotchi effect is more appropriate than "some emotional Turing test."

https://en.wikipedia.org/wiki/Tamagotchi_effect

AIPedant··on GPT-5
No, most of programming is at least implicitly coming up with a human-language description of the problem and solution that isn't full of gaps and errors. LLM users often don't give themselves enough credit for how much thought goes into the prompt - likely because those thoughts are easy for humans! But not necessarily for LLMs.

Sort of related to how you need to specify the level of LLM reasoning not just to control cost, but because the non-reasoning model just goes ahead and answers incorrectly, and the reasoning model will "overreason" on simple problems. Being able to estimate the reasoning-intensiveness of a problem before solving it is a big part of human intelligence (and IIRC is common to all great apes). I don't think LLMs are really able to do this, except via case-by-case RLHF whack-a-mole.

AIPedant··on Vibechart
It is actually incredible how they managed to find an even more unscientific definition than "can perform a majority of economically useful tasks." At least that definition requires a little thought to recognize it has problems[1]. $100bn in profits is just cartoonishly dumb, like you asked a high schooler to come up with a definition.

[1] If a computer can perform the task its economic usefulness drops to near zero, and new economically useful tasks which computers can't do will take its place.

← PreviousPage 2 of 8Next →