993 karma · joined April 5, 2025
Colin Fraser had a good tweet about this: https://xcancel.com/colin_fraser/status/1956414662087733498#...
In a therapy session, you're actually going to do most of the talking. It's hard. Your friend is going to want to talk about their own stuff half the time and you have to listen. With an LLM, it's happy to do 99% of the talking, and 100% of it is about you.See also https://hai.stanford.edu/news/law-policy-ai-update-does-sect... - Congress and Justice Gorsuch don't seem to think ChatGPT is protected by 230.
FWIW I agree that OpenAI wants people to have unhealthy emotional attachments to chatbots and market chatbot therapists, etc. But there is a separate problem.
a) the human would (deservedly[1]) be arrested for manslaughter, possibly murder
b) OpenAI would be deeply (and deservedly) vulnerable to civil liability
c) state and federal regulators would be on the warpath against OpenAI
Obviously we can't arrest ChatGPT. But nothing about ChatGPT being the culprit changes 2) and 3) - in fact it makes 3) far more urgent.
[1] It is a somewhat ugly constitutional question whether this speech would be protected if it was between two adults, assuming the other adult was not acting as a caregiver. There was an ugly case in Massachusetts involving where a 17-year-old ordered her 18-year-old boyfriend to kill himself and he did so; she was convicted of involuntary manslaughter, and any civil-liberties minded person understands the difficult issues this case raises. These issues are moot if the speech is between an adult and a child, there is a much higher bar.
What is true is that (for example) Split Fiction and Tony Hawk 3/4 are quite a bit less fancy than the PS5 or XBox versions, which is unflattering.
I don't think Nintendo is going to go "Seal of Quality" but it would be nice if the Switch 2-filtered eShop was not full of cynical trash. The ease of publishing for the Switch 1 was new for Nintendo, and it was welcomed at the time, but in retrospect they went too far.
I cloned the backend for Truco and gave Claude a long prompt explaining the rules of Escoba and asking it to refactor the code to implement it.
How long would it take the human dev to refactor the code themselves? I think it's plausible that it would be longer than 3 days, but maybe not!So the question is whether human intelligence has higher-level primitives that can be implemented more efficiently - sort of akin to solving differential equations, is there a “symbolic solution” or are we forced to go “numerically” no matter how clever we are?
Castlevania... [so] called because it is a Metroidvania game set in a Castle.
Ouch - this is precisely backwards. Metroidvanias are named after Metroid and Castlevania because those series practically defined the genre.Also a bit frustrating because the first Castlevania itself isn't actually a metroidvania, it's a more conventional action-platformer. Castlevania II has non-linear exploration, lots of items to collect, and puzzle-solving, all like Metroid. So it's not too surprising Antithesis had to do a lot of work for adapting their system to Metroid - but I wonder if this work means it now can handle Castlevania II without much extra development.
[1] If the search takes more than a few minutes then the AI overview is almost guaranteed to be wrong or useless.
That said, "GPT-5 will not be any better than competitors' products, demonstrating OpenAI was bluffing about AGI and destroying investor exuberance" was a very specific prediction made by (for example) Gary Marcus.
As Alaska's salmon plummet, scientists home in on the killer - Science - AAAS
seemingly a goofy copy-paste thing.Gauss's constant k is defined as sqrt(G), but for a while the international standard was to define k and then compute G as k^2, which is why NASA refers to it that way.
"Information wants to be free," which means that any cost of producing that information can be abstracted away due to ideological inconvenience.
It's amazing how widespread this belief is among the HN crowd, despite being a shameless ad hominem with zero evidence. I think there are a lot of us who assume the reasonable hypothesis is "LLMs are a compelling new computing paradigm, but researchers and Big Tech are overselling generative AI due to a combination of bad incentives and sincere ideological/scientific blindness. 2025 artificial neural networks are not meaningfully intelligent." There has not been sufficient evidence to overturn this hypothesis and an enormous pile of evidence supporting it.
I do not necessarily believe humans are smarter than orcas, it is too difficult to say. But orcas are undoubtedly smarter than any AI system. There are billions of non-human "intelligent agents" on planet Earth to compare AI against, and instead we are comparing AI to humans based on trivia and trickery. This is the basic problem with AI, and it always has had this problem: https://dl.acm.org/doi/10.1145/1045339.1045340 The field has always been flagrantly unscientific, and it might get us nifty computers, but we are no closer to "intelligent" computing than we were when Drew McDermott wrote that article. E.g. MuZero has zero intelligence compared to a cockroach; instead of seriously considering this claim AI folks will just sneer "are you even dan in Go?" Spiders are not smarter than beavers even if their webs seem more careful and intricate than beavers' dams... that said it is not even clear to me that our neural networks are capable of spider intelligence! "Your system was trained on 10,000,00 outdoor spiderwebs between branches and bushes and rocks and has super-spider performance in those domains... now let's bring it into my messy attic."
test reasoning abilities such as pattern recognition, lateral thinking, abstraction, contextual reasoning (accounting for British cultural references), and multi-step inference.... its emphasis on clever reasoning rather than knowledge recall, Only Connect provides an ideal challenge for benchmarking LLMs' reasoning capabilities.
It seems to me that the null hypothesis should be "LLMs are probabilistic next-word generators and might be able to solve a lot of this stuff with shallow surface statistics built from inhumanly large datasets, without ever properly using abstraction, contextual reasoning, etc." This is particularly true for NYT Connections, but in general evaluations like this seem to be at least partially testing how amenable certain word/trivia games are to naive statistical algorithms. (Many NYT Connections "purple" categories seem like they would be quite obvious to a next n-gram calculator, but not for people who actually use words conversationally!) Humans don't use these statistical algorithms for reasoning except in particular circumstances (many use "folk n-gram statistics" when playing Wordle; poker; serious word game players often learn more detailed tables
of info; you could see competitive NYT Connections players learning a giant bag of statistical heuristics to help them speedrun things). We just can't accumulate the data ourselves without making a concerted computer-aided effort.In general a lot of LLM benchmarks don't adequately consider that LLMs can solve certain things better than humans without using reasoning or knowledge. The most stupid example is how common multiple choice benchmarks are, despite us all learning as children that multiple-choice questions can be partially gamed with shallow statistical-linguistic tricks even if you have no clue how to answer the question honestly[1]; it stands to reason that a superhuman statistical-linguistic computer could accumulate superhuman statistical-linguistic tricks without ever properly learning the subject matter. AI folks have always been quick to say "if it quacks like a duck it reasons like a duck" but these days computers are quite good at playing duck recordings.
[1] "When in doubt, C your way out," sniffing out suspicious answers, shallow pattern-matching to answer reading comprehension, etc etc. One thing humans and LLMs actually do have in common is that multiple-choice tests are terrible ways to assess their knowledge or intelligence.
L. Ron Hubbard is more like the Zizians.
The rationalist community was drawn together by AI researcher Eliezer Yudkowsky’s blog post series The Sequences, a set of essays about how to think more rationally
I actually don't mind Yudkowski as an individual - I think he is almost always wrong and undeservedly arrogant, but mostly sincere. Yet treating him as an AI researcher and serious philosopher (as opposed to a sci-fi essayist and self-help writer) is the kind of slippery foundation that less scrupulous people can build cults from. (See also Maharishi Mahesh Yogi and related trends - often it is just a bit of spiritual goofiness as with David Lynch, sometimes you get a Charles Manson.)Screenshot of original video: https://www.resetera.com/threads/digital-foundry-posts-an-ad...
It is a bit shameful that games journalists didn't cover this at all - seems almost like an omertà. Even assuming the mislabelling was an honest mistake, running the ad in the first place makes me trust them less. Worse, I used their videos when deciding to buy a Switch 2! I would have been much more skeptical (likely ignored them entirely) if I had known they were taking money from Nintendo.
a) a game of roulette where you hope the LLM provider has RLHFed something very close to your use case, or
b) trying to few-shot it with in-context examples requires more engineering (and is still less reliable) than simply doing it yourself
In particular it's not just "the lack of a succinct, exhaustive text description," it also a lack of English->Prolog "translations."
It seems like the LLM-Prolog community is well aware of all this (https://swi-prolog.discourse.group/t/llm-and-prolog-a-marria...) but I don't see anything in Universalis that solves the problem. Instead it's just magically invoking the LLM.
While this may seem like a whimsical example, it is not intrinsically easier or harder for an AI model compared to solving a real-world problem from a human perspective. The model processes both simple and complex problems using the same underlying mechanism. To lessen the cognitive load for the human reader, however, we will stick to simple targeted examples in this article.
For LLMs this is blatantly false - in fact asking about "used textbooks" instead of "apples" is measurably more likely to result in an error! Maybe the (deterministic, Prolog-style) Universalis language mitigates this. But since Automind (an LLM, I think) is responsible for pre/post validation, naively I would expect it to sometimes output incorrect Universalis code and incorrectly claim an assertion holds when it does not.Maybe I am making a mountain out of a molehill but this bit about "lessen the cognitive load of the human reader" is kind of obnoxious. Show me how this handles a slightly nontrivial problem, don't assume I'm too stupid to understand it by trying to impress me with the happy path.
Sort of related to how you need to specify the level of LLM reasoning not just to control cost, but because the non-reasoning model just goes ahead and answers incorrectly, and the reasoning model will "overreason" on simple problems. Being able to estimate the reasoning-intensiveness of a problem before solving it is a big part of human intelligence (and IIRC is common to all great apes). I don't think LLMs are really able to do this, except via case-by-case RLHF whack-a-mole.
[1] If a computer can perform the task its economic usefulness drops to near zero, and new economically useful tasks which computers can't do will take its place.