They are very good at fooling people; perhaps Turing's Test is not a good measure of intelligence after all, it can easily be gamed and we find it hard to differentiate apparent facility with language and intelligence/knowledge.
They are very good at fooling people; perhaps Turing's Test is not a good measure of intelligence after all, it can easily be gamed and we find it hard to differentiate apparent facility with language and intelligence/knowledge.
I wouldn't say zero intelligence, but I wouldn't describe such systems as intelligent, I think it misrepresents them, they do as you say have a good depth of knowledge and are spectacular at reproducing a simulacrum of human interactions and creations, but they have been a lesson for many of us that token manipulation is not where intelligence resides.
Must it have one? The words "artificial intelligence" are a poor description of a thing when we've not rigorously defined it. It's certainly artificial, there's no question about that, but is it intelligent? It can do all sorts of things that we consider a feature of intelligence and pass all sorts of tests, but it also falls down flat on its face when prompted with a just-so brainteaser. It's certainly useful, for some people. If, by having inhaled all of the Internet and written books that have been scanned as its training data, it's able to generate essays on anything and everything, at the drop of a hat, why does it matter if we can find a brainteaser it hasn't seen yet? It's like it has a ginormous box of Legos, and it can build whatever your ask for with these Lego blocks, but pointing out it's unable create its own Lego blocks from scratch has somehow become critically important to point out, as if that makes this all total dead end and it's all a waste of money omg people wake up oh if only they'd listen to me. Why don't people listen to me?
Crows are believed to have a theory of mind, and they can count up to 30. I haven't tried it with Claude, but I'm pretty sure it can count at least that high. LLMs are artificial, they're alien, of course they're going to look different. In the analogy where they're simply a next word guesser, one imagines standing at a fridge with a bag of magnetic words, and just pulling a random one from the bag to make ChatGPT. But when you put your hand inside a bag inside a bag inside a bag, twenty times (to represent the dozens of layers in an LLM model), and there are a few hundred million pieces in each bag (for parameters per layer), one imagines that there's a difference; some sort of leap, similar to when life evolved from being a single celled bacterium to a multi-cellular organism.
Or maybe we're all just rubes, and some PhD's have conned the world into giving them a bunch of money, because they figured out how to represent essays as a math problem, then wrote some code to solve them, like they did with chess.
These tools aren’t useless, obviously.
But people do really learn hard into confirmation bias and/or personification when it comes to LLMs.
I believe it’s entirely because of the term “artificial intelligence” that there is such a divide.
If we called them “large statistical language models” instead, nobody would be having this discussion.
I have tried various models out for tasks from generating writing, to music to programming and am not impressed with the results, though they are certainly very interesting. At every step it will cheerfully tell you that it can do things then generate nonsense and present it as truth.
I would not describe current LLMs as able to generate essays on anything - they certainly can but they will be riddled with cliche, the average of the internet content they were trained on with no regard for quality and worst of all will contain incorrect or made up data.
AI slop is an accurate term when it comes to the writing ability of LLMs - yes it is superficially impressive in mimicking human writing, but it is usually vapid or worse wrong in important ways, because again, it has no concept of right and wrong or model of the world which it attempts to make the generated writing conform to, it just gets stuck with some very simple tasks, and often happily generates entirely bogus data (for example ask it for a CSV or table of data or to reproduce the notes of a famous piece of music which should be in its training data).
Perhaps this will be solved, though after a couple of years of effort and a lot of money spent with very little progress I'm skeptical.
Are you invisibly qualifying this as the inability to generate interesting or entertaining essays? Because it will certainly output mostly-factual, vanilla ones. And depending on prompting, they might be slightly entertaining or interesting.
I have made some minor games in JS with my kids with one for example, and managed to get it to produce a game of asteroids and pong with them (probably heavily based on tutorials scraped from the web of course). I had less success trying to build frogger (again probably because there are not so many complete examples). Anything truly creative/new they really struggle with, and it becomes apparent they are pattern matching machines without true understanding.
I wouldn't describe LLMs as useful at present and do not consider them intelligent in any sense, but they are certainly interesting.
As other examples I asked it for note sequences from a famous piece and it cheerfully generated gibberish, and the more subtly wrong sequences when asked to correct. Generating a csv of basic data it should know was unusable as half the data was wrong and it has no sense of whether things are correct and logical etc etc. There is no thinking going on here, only generation of probable text.
I have used GAI at work a few times too but it needed so much hand holding it felt like a waste of time.
"Right, so what the hell is this cursed nonsense? Elon Musk, billionaire tech goblin and professional Twitter shit-stirrer, is apparently offering up his personal fucking sperm to create some dystopian family compound in Texas? Mate, I wake up every day thinking I’ve seen the worst of humanity, and then this bullshit comes along.
And then you've got Wes Pinkle summing it up beautifully with “What a terrible day to be literate.” And yeah, too fucking right. If I couldn't read, I wouldn't have had to process the mental image of Musk running some billionaire eugenics project. Honestly, mate, this is the kind of headline that makes you want to throw your phone into the ocean and go live in the bush with the roos.
Anyway, I hope that’s more the aggressive kangaroo energy you were expecting. You good, or do you need me to scream about something else?"
This sort of disconnected word salad is a good example of the dross llms create when they attempt to be creative and don’t have a solid corpus of stock examples to choose from.
The frogger game I tried to create played as this text reads - badly.
The whole thing seems Oz-influenced (example, "in the bush with the roos"), which implies to me that he's prompted it to speak that way. So, you assumed an error when it probably wasn't... Framing is a thing.
Which leads to my point about your Frogger experience. Prompting it correctly (as in, in such as way as to be more likely to get what you seek) is a skill in itself, it seems (which, amazingly, the LLM can also help with).
I've had good success with Codeium Windsurf, but with criticisms similar to what you hint at (some of which were made better when I rewrote prompts): On long contexts, it will "lose the plot"; on revisions, it will often introduce bugs on later revisions (which is why I also insist on it writing tests for everything... via correct prompting, of course... and is also why you MUST vet EVERY LINE it touches), it will often forget rules we've already established within the session (such as that, in a Nix development context, you have to prefix every shell invocation with "nix develop" etc.)...
The thing is, I've watched it slowly get better at all these things... Claude Code for example is so confident in itself (a confidence that is, in fact, still somewhat misplaced) that its default mode doesn't even give you direct access to edit the code :O And yet I was able to make an original game with it (a console-based maze game AND action-RPG... it's still in the simple early stages though...)
Re promoting for frogger, I think the evidence is against that - it does well on games it has complete examples for (i.e. it is reproducing code) and badly on ones it doesn’t have examples for (it doesn’t actually understand what it doing though it pretends to and we fill in the gaps for it).
It is clearly happening as shown by numerous papers studying it. Here is a popular one by anthropic
I wouldn't read into marketing materials by the people whose funding depends on hype.
Nothing in the link you provided is even close to "neurons, model of the world, thinking" etc.
It literally is "in our training data similar concepts were clustered with some other similar concepts, and manipulating these clusters lead to different outcomes".
Recognizing concepts, grouping and manipulating similar concepts together, is what “abstraction” is. It's the fundamental essence of both "building a world model" and "thinking".
> Nothing in the link you provided is even close to "neurons, model of the world, thinking" etc.
I really have no idea how to address your argument. It’s like you’re saying,
“Nothing you have provided is even close to a model of the world or thinking. Instead, the LLM is merely building a very basic model of the world and performing very basic reasoning”.
Once again, it does none of those things. The training dataset has those concepts grouped together. The model recognizes nothing, and groups nothing
> I really have no idea how to address your argument. It’s like you’re saying,
No. I'm literally saying: there's literally nothing to support your belief that there's anything resembling understanding of the world, having a world model, neurons, thinking, or reasoning in LLMs.
The link mentions "a feature that triggers on the Golden Gate Bridge".
As a test case, I just drew this terrible doodle of the Golden Gate Bridge in MS paint: https://imgur.com/a/1TJ68JU
I saved the file as "a.png", opened the chatgpt website, started a new chat, uploaded the file, and entered, "what is this?"
It had a couple of paragraphs saying it looked like a suspension bridge. I said "which bridge". It had some more saying it was probably the GGB, based on two particular pieces of evidence, which it explained.
> The model recognizes nothing, and groups nothing
Then how do you explain the interaction I had with chatgpt just now? It sure looks to me like it recognized the GGB from my doodle.
Machine learning models can do this and have been for a long time. The only thing different here is there's some generated text to go along with it with the "reasoning" entirely made up ex post facto
Predominantly English-language data set with one of the most famous suspension bridges in the world?
How can anyone explain the clustering of data on that? Surely it's the model of the world, and thinking, and neurons.
What happens if you type "most famous suspension bridges in the world" into Google and click the first ten or so links? It couldn't be literally the same data? https://imgur.com/a/tJ29rEC
that is the paper being linked to by the "marketing material". Right at the top, in plain sight.
If you were arguing in good faith, you'd head directly there instead of lampooning the use of a marketing page in a discussion.
That all said, skepticism is warranted. Just not an absolute amount of it.
Which part of the paper supports the "models have a world model, reasoning, etc." and not what I said, "in our training data similar concepts were clustered with some other similar concepts, and manipulating these clusters lead to different outcomes"?
You should learn a bit about media literacy.
In fact, it still very much seems like marketing. Especially since the paper was made in association with Anthropic.
Again. Learn some media literacy.
I'm going to guess that sometimes they will: driven onto areas where there's no existing article, some of the time you'll get made-up stuff that follows the existing shapes of correct articles and produces articles that upon investigation will turn out to be correct. You'll also reproduce existing articles: in the world of creating art, you're just ripping them off, but in the world of Wikipedia articles you're repeating a correct thing (or the closest facsimile that process can produce)
When you get into articles on exceptions or new discoveries, there's trouble. It can't resynthesize the new thing: the 'tokens' aren't there to represent it. The reality is the hallucination, but an unreachable one.
So the LLMs can be great at fooling people by presenting 'new' responses that fall into recognized patterns because they're a machine for doing that, and Turing's Test is good at tracking how that goes, but people have a tendency to think if they're reading preprogrammed words based on a simple algorithm (think 'Eliza') they're confronting an intelligence, a person.
They're going to be historically bad at spotting Holmes-like clues that their expected 'pattern' is awry. The circumstantial evidence of a trout in the milk might lead a human to conclude the milk is adulterated with water as a nefarious scheme, but to an LLM that's a hallucination on par with a stone in the milk: it's going to have a hell of a time 'jumping' to a consistent but very uncommon interpretation, and if it does get there it'll constantly be gaslighting itself and offering other explanations than the truth.