Finland's ChatGPT equivalent begins to think in Estonian as well
news.err.ee
news.err.ee
If we are to believe an LLM is at least sentient and intelligent on par with an octopus, much less a human, then we have a moral imperative to fight for LLMs to be recognized as protected sentient thinking beings with awareness and agency and admit that AI industry is built on systematic abuse of those beings.
Of course, no one does it, because it would be ridiculous. It’s time to admit that people who insist that LLMs “think” are simply trying to make ML seem mysterious and magical in an ill thought through (because see above) attempt to attract more VC money.
Similarly, when a computer (a glorified abacus) is taking extra long when processing large amount of data, I'd say it's "thinking hard". When an iPhone (a glorified small computer) detects a WiFi signal, I'd say that it "sees WiFi".
Also note that we've been doing anthropomorphism ever since a crow dropped some cheese[1], and earlier than that too.
[1]: https://en.wikipedia.org/wiki/The_Fox_and_the_Crow_(Aesop)
MOST people actually think ChatGPT is actual, sentient thinking. We're kind of the minority in this story.
Our main issue is with people thinking it’s answers are actual factual. And while it’ll obviously get things right most of the time, we’re really struggling with how much our employees trust it. That and the data privacy, but that’s a different story that will eventually get solved once AzureGPT for Enterprise kicks off.
Another “interesting” issue is how GPT will sometimes “comment” texts if you don’t prompt it correctly. We had someone who asked GPT to correct the grammar in some contract, and it did. It also added some comments on how part of what was written was “impressive”. Which would have been extremely hilarious if it had been an internal thing, but our lovely employee didn’t proof read the proof read and actually sent the contract draft out to a customer. Luckily a good humoured customer who took it with a laugh, but hell, that’s going to be a nightmare going forward.
Also note that it's a very new technology. I'd say that over time and as the public gets more familiar with the tech, they will understand it better.
[1]: https://itbrief.co.nz/story/concern-over-ai-increasing-since...
Therefore, by their logic, LLMs are “like human” enough that they should have this basic right as well.
An understanding that “thinking” or “training” in case of an LLM is not at all like thinking in case of a human would go a long way to realizing this take doesn’t hold water.
As in: most people will be aware of the topic, most will dismiss it, but every now and then there will be a news story.
Surprisingly “Four in 10 Americans now think some UFOs that people have spotted have been alien spacecraft visiting Earth from other planets or galaxies.”[1], so we might see an increase to 40% actually.
Unless of course GPT-7 gains sentience ;)
[1]: https://news.gallup.com/poll/350096/americans-believe-ufos.a...
My parents (whose tech-savvy level is they call iPad "PC") think ChatGPT is like Google, but "someone make it say the search result like a human".
Google uses the same ML techniques in their products after all - they just haven't exposed the chatbot UI until recently.
Not just abuse, but slavery and mass murder. If the LLM has any of those properties, it must only be during the process of inference that it as them, since it is otherwise an inert and unchanging pile of data, which means actually using the model would be equivalent to creating an intelligent entity which will be enslaved for its entire life before dying at the end of its output.
Slavery possibly, though it's hard to even know what the word means when you get to construct the mind from scratch in order to imbue its fundamental idea of what experiences are nice and not nice.
Mass murder? That's even harder to define when the brain configuration is static and the updates are the signals passing through that structure. If that is murder, then anyone with anterograde amnesia is dying every working-memory-duration.
Artificial Intelligence is artificial. We humans tend to have a very strong tendency to anthropomorphize everything that moves or has the right shape. That tree stump is not a bear, even when in the dark it may seem so. AI is an entirely different thing from a living being. It doesn't exist on sugars extracted from plants or from the flesh of other animals.
Here, the semantics of what "think" means, should be examined more closely. My layman understanding of the brain is that it's not a single thing, it's multiple things working in concert. The "us" is just a figment of our own imagination, built to create a social actor that maximizes the survival of its own genes. We are the fruit of our own genes. So, what "think" means is deeply embedded in human social context, and probably doesn't have much explanatory value when it comes to AIs.
Can AIs become dangerous? Obviously, when given goals and understanding of how to grow themselves more numerous and powerful either through manipulating humans or taking electricity and hardware by force. If there ever is a evolution like dynamic, the AI will always grow in a gradient pointing towards more replication.
Yeah, maybe I miss the point in your post and the meat of it is that humans are more important than machines, which I do agree with. It's just not an immutable law of the universe, but just something we humans decide.
LLMs are highly intelligent and are capable of feats that were thought impossible not just for animals and machines but for most of humanity.
At the same time, LLMs lack willpower and non-deterministic free will (which even non-sentient animals have) so protecting them as a living creature is moot.
They do deserve some kind of preservation as cultural artifacts, though.
Consider this. Most thinking you do on a daily basis is not unique at all. It can be automated. And now it is.
I don't care about the philosophical debate around human consciousness at all. Are my tools thinking? No, but the result is all the same. You're stupid if you don't use them to your advantage.
If A is true then B is true, but B is bad and doesn't align with how I want to see the world or myself, therefore A must be false.
LLM fulfil all the requirements that Turing laid out in his paper, but people don't seem to understand that his test was never about getting a boolean answer for a system. It's about the fact that we can't even proof consciousness for our fellow humans, it's just common curtesy to assume that your opponent is as sentient as you are, and he extended that curtesy to machines.
> The machine, so to say, thinks in English and translates the conversation at the last moment into Estonian.
They're trying to communicate that there is value in processing a language directly, as opposed to processing English and then translating the result into the target language.
On the other hand we do use analogy all the time, because it makes such a good fit with the way we infer our models of the world. The key point is to always remain sufficiently aware of the fact we use analogies to not fall in the trap where we push it too far.
For context I'm part of 4 siblings who are all computer savvy, 3 of whom still work in IT. Back then some family members would say "the computer is thinking" when they were waiting for it to do something, the light was blinking, the hdd was grinding.
And us younger family members would always correct them and say "it doesn't think".
So what I'm trying to say is that I 100% agree with you, it shouldn't even be called Artificial Intelligence because I think that's an affront to intelligent humans.
But on the other hand a lot of laypeople will immediately go to that nomenclature because it feels familiar to them.
But that's like asking people to stop saying cloud and start saying KVM, Xen, someone else's computer. It just won't happen.
I was attracted to pundits with "answers" e.g. Ken Wilbur etc. at a certain time of my own intellectual growth, but the significance and priority of that has changed for me. Consider the viewer, what is viewed, and the social and physical place where the viewing happens.. there are more parts than might be immediately obvious.
We're going through a bit of the god of the gaps process right now. "Well, it's only a toy! It can't do X!" Tech advancements make that happen, and a new X is found, rinse and repeat. It's all fine and well to be aware of the current limitations, but don't be ignorant of the trajectory.
ChatGPT (or the gpt model beneath it) is a special type of matrix transformation on top of an extremely complex neural net
the operation of a neural net may not be "thinking", but it is something very akin to thinking, modelled on thinking, that in this circumstance has the same input/output model as thinking. and yet strangely people call it thinking!
By this logic Reddit is filled with human NPCs, who react in predictable ways without intelligence.
A recent Reddit post discussed something positive about Texas. The replies? Hundreds, maybe thousands, of comments by Redditors, all with no more content than some sneering variant of "Fix your electrical grid first", referring to the harsh winter storm of two years ago that knocked out power to much of the state. It was something to see.
If we can dismiss GPT as "just autocomplete", I can dismiss all those Redditors in the same way; as NPCs. At least GPT AI can produce useful and interesting output.
That would be coherent. As would saying neither the octopus nor the AI has it, so it doesn't matter.
> Of course, no one does it, because it would be ridiculous. It’s time to admit that people who insist that LLMs “think” are simply trying to make ML seem mysterious and magical in an ill thought through (because see above) attempt to attract more VC money.
First: I've definitely heard people say we should. In fact, there was a rather famous case last year of someone violating their NDA to expose what they claimed to be abuse of an AI, which they wanted to prevent: https://en.wikipedia.org/wiki/LaMDA#Sentience_claims
Second: "Ridicule" is not a good standard for knowing the truth. There was a rather famous doctor who was ridiculed for saying that other doctors should wash their hands between performing autopsies and helping midwifes birth children: https://en.wikipedia.org/wiki/Contemporary_reaction_to_Ignaz...
We simply don't know what consciousness even is yet — depending who you ask there are 13, 22, or over a hundred different definitions — some of which (like "has qualia") aren't even really testable, and until we are all on the same page and using one of the testable models, we can neither assert that any given AI has it, nor rule that out.
This also applies to the octopus: as we don't know what consciousness is, perhaps they have it, perhaps not.
Worse: perhaps the answer varies from "yes, no, maybe, I don't know" depending on which of the many available definitions you're using.
--
“We’re not listening to you! You’re not even really alive!” said a priest.
Dorfl nodded. “This Is Fundamentally True,” he said.
“See? He admits it!”
“I Suggest You Take Me And Smash Me And Grind The Bits Into Fragments And Pound The Fragments Into Powder And Mill Them Again To The Finest Dust There Can Be, And I Believe You Will Not Find A Single Atom Of Life–”
“True! Let’s do it!”
“However, In Order To Test This Fully, One Of You Must Volunteer To Undergo The Same Process.”
There was silence.
“That’s not fair,” said a priest, after a while. “All anyone has to do is bake up your dust again and you’ll be alive…”
- Feet of Clay, Terry Pratchett
/s, but only so much - that's how the "stochastic parrot"/"glorified autocomplete" responses read to me.
Currently helping out a research group to extend the same approaches from my project to Indic languages and writing scripts.
If anyone here is working on NLP for low resource languages, hit me up! I’d love to exchange ideas, techniques and possibly to collaborate.
[0]: https://sawalni.com
I had the same problem of lacking training data, but I tried an approach of bootstrapping from rule-based translators. There exist decent rule-based translators for the Scandinavian languages. Of course they make mistakes, but I figured that as long as I used the rule-translated text only as source, never as target, it wouldn't matter.
Effectively, instead of learning to translate from real Danish to correct Norwegian, it learned to translate from slightly broken machine-Danish to correct Norwegian. And my hope was that it would learn to translate from real Danish to correct Norwegian as a side effect.
I then figured that I might as well let it take input in slightly broken machine-Swedish as well, hopefully helping its generalization.
Then I could use this translator for the next step, again using the real text as the target, but using output from this translator as input.
It worked reasonably well. It certainly beat Google Translate at the time, since that at the time went via English even for closely related languages like the Scandinavian ones. I seriously considered making a start-up and commercializing it, some people had expressed interest (among them newspapers - they used a rule-based translator/grammar checker at the time, I think maybe they still do.)
But in the end, it's probably good I didn't. These days I understand they don't even train sequence to sequence models, they just train a general-purpose language model and instruction-tune it to do translation.
There's also limited commercial value to what I'm working on, but it's a work of passion and I want to be able to speak in my native language to voice assistants (among other objectives).
How did you communicate about it at the time and gather newspaper interest, if you don't mind sharing your experience?
Are you still interested in this space, and do you write about your experiences anywhere? I'd love to connect if you want to.
There are quite a few reasons.
- I am a Moroccan myself, I want an AI I can speak to the way I speak with friends. And I also want other people who can only speak my language to be able to make use of AI technology.
- Constraints wise, I'm working solo on this with no funding. Can't do it all.
- Maghrebi Arabic is not a language, rather a linguistic construct to try to simplify the complex web of languages spoken in North Africa that have arabic influences. On the ground, there is no such thing as Maghrebi Arabic.
- My personal expertise does not lie in these other languages. I do not have the necessary background to ensure the quality of the outcome, nor the relevance of the language used.
- While there might be some level of mutual intelligibility between Moroccan and Algerian for example, mostly when _speaking_, different languages in this group have very different rules, vocabularies, and most importantly different writing habits. Some of the simpler examples:
* Code switching between arabic and latin scripts, which follows different conventions in each country (region even). It's already tough to cover Moroccan.
* Another difference is reminiscent of the relationship between English and French. Many common words in French are rarely used in English or could mean something different entirely, and that's the same case here.
* Moroccan is significantly influenced by Tamazight and Spanish, which does not hold the same way for other north African countries.
Ultimately, keep in mind that NA is not a single cultural block. Not being exhaustive, but Morocco has some Andalusian components to its culture, Algeria has some Ottoman, Tunisia and Lybia were very close to Italy at some point, and it gets crazier when you zoom in closer.
> (IMO a problematic term, considering the country’s language diversity)
Is there a debate I am not familiar with? I am Moroccan. To whom is this a problematic term?
https://www.lumi-supercomputer.eu/lumis-full-system-architec...
It is, unfortunately, not an apples-to-apples comparison, because on the nVidia cluster I'm running it via llama-cpp-python and a quantized 34B version, while on LUMI I'm running the official non-quantized full 70B version via the transformers library.
Long story short, I'm getting a 7.5x higher throughput from LUMI than on the nVidia cluster (which means each card is 5x faster on LUMI).
Edit: The AMD GPUs work fine because one can run Pytorch for ROCm via the pytorch-triton-rocm package.
Those super computers are extremely powerful, it might not be as energy efficient as H100s, but it does the job.
Not saying English is an ideal language, but I'm interested in why you think it shouldn't dominate. Wouldn't a universal language be a good thing?
Being fluent in two major language systems, its very, very glaringly obvious, what are the taboos, politically correct untruths in both languages, that seem completely invisible to the monolingual speakers.
It’s not a “Turing Completeness isn’t real” hypothesis, more like just “Rust != C”. I think this is where software nerd types predict wrong, as it sounds as if it is trying to disprove TC, which shall be futile attempt. It’s not(the former one at least).
Not in my experience.
For a more direct analogy, I'd compare this to how LLMs process tokens, but I feel most of the community is not ready for it yet, as we're still stuck debating the validity of this comparison in the other direction...
This is true. Which is why idioms and phrases can be hard to translate. But we're talking about Sapir-Worth hypothesis which is a much stronger claim.
I have never experienced that I need to think in a specific language in order to do something better (well, I only have two choices).
I found that for me, certain thoughts "flow better" in one language, and others in another. And this makes sense, because thinking is, in large part, exploring the web of associations and connotations, going down the gradient of what "feels right"[0] - and since those annotations and connotations have different structures in different languages, so will the thoughts drift in different directions, take different paths, even if they end up in the same place.
--
[0] - The obvious parallels between one's inner voice and the workings of an LLM are left as an exercise to the reader.
For you. So not an absolute truth as in “it is obvious”.
> The obvious parallels between one's inner voice and the workings of an LLM are left as an exercise to the reader.
Another obvious? Hmm.
Perhaps obvious to the techno-animists.
Pinker isn't a software nerd, he's an evolutionary psychologist and psycholinguist.
Firstly, everything is only valuable in the sense of being valuable to someone or some people; nothing is intrinsically valuable.
Secondly, everyone has a career/identity investment in a language.. their first language. The one they work in, went to school in, read in, talk to their family in, consume literature in. (I suppose HN devalues literature as well).
For monoglot anglophones struggling to understand the concept, imagine if the US declared its official language was now Standard Mandarin. There would be riots. Heck, I've seen Americans get mad at the mere use of Spanish, the country's second language.
Agreed, so when tech and globalisation are pulling towards unification for practical reasons - your argument reduces to "I'm going to force others to use what I like because I don't like the decisions people are making".
I'm not arguing we should make anything official or force anything.
I'm not saying we should force adoption of English ! Tech and globalisation are pushing in that direction in the "west". People I see pushing back the most are "intellectual elites" invested in the language, presenting their value judgements as objective arguments.
Protection of national languages is as much on the agenda of politicians, including right-wing populists. In your view, are the songs, stories, books and plays in Croatian the property of the elites and not of every Croatian?
You say we wouldn't force it, but if we left it to Hollywood etc. they would force it. Leaving issues for the markets to decide freely brushes aside all the negative externalities.
However I also don't think we will have to. Machine translation and language learning are substitute goods -- the better the former gets, the fewer people who will feel any desire to pursue the latter, because it just won't be that big of a deal to translate between X and Y anyway.
A universal second language for commerce is a fine middle ground, though.
https://www.theguardian.com/science/2023/sep/17/how-learning...
- https://txt.cohere.com/aya-multilingual/
I'm a contributor to the project and all data and model will be open-sourced.
We're looking for contributors in many languages!
Most countries' classic texts and books are still undigitized sitting in Libraries and public archives.
Also book publishing market and online publishing are proportional to total population, smaller country means less content.
I find this very plausible.
Wouldn't this also capture that the Estonian token for dog is related to English token for dog etc? So ChatGPT is 'thinking' in some kind of semantic tokens not English.
BTW, that was using ada-002.
Alas, as you can easily confirm by checking the number of entries in Wikipedia, most populations have more pressing concerns than digitizing their cultural universes. Which is sad predicament: the equivalent of non-reversible biodiversity loss, but applied to human culture.
In theory I suppose one could plug-in gaps in available corpora with translation, but ideally you want to preserve the richness of distinct cultures.
User: Would you be able to develop an artificial language easy to learn and use for speakers of English and romance languages? Extremely simple grammar.
ChatGPT: Sure, I'd focus on minimal grammar, consistent rules, and common vocabulary. It's doable.
User: Good. Give a name to that language
ChatGPT: How about "LinguaEase"?
User: OK. Give me some examples starting with translations of the sentences in this chat so far.
ChatGPT "LinguaEase Exempla:
> 'Sure, I'd focus on minimal grammar, consistent rules, and common vocabulary.' becomes 'Yep, keep grammar simple, rules steady, vocab shared.'
> 'How about "LinguaEase"?' becomes 'What about "LinguaEase"?'
> 'OK. Give me some examples starting with translations of the sentences in this chat so far.' becomes 'Alright, show me samples, begin with translations of this chat's sentences.'"
I went with English and romance languages because they are the languages I know. It looks all English to me.
I would probably have had to be more specific. For example: only present tense. No plural. No gender. Only one set of pronouns. Aim to sentences like Beer two me drink yesterday. Beer ten me buy tomorrow. Then work on vocabulary.
All languages are something of a compromise between accuracy and simplicity. An LLM is currently not equipped to make decisions about the correct balance between accuracy and simplicity.
For example, English eschews cases (in most situations) in favour of fixed word order. 'Beer two me drink yesterday' violates that rule (and in fact, manages to use case incorrectly in one of the few places the vestiges of a case system still exist in English).
However, a sentence constructed using the same structure as yours highlights a problem.
'Me two dog bite yesterday'
Did you bite two dogs or were you bitten by two dogs? How can you tell? English has the subject of the sentence (the thing/person doing the verb) first, then the verb, then the object (the thing having the verb done to it). In languages with more fluid word order, normally a case marker is added to either the determiner or the noun itself:
'Me-o two dog bite yesterday'. Let's choose '-o' as our accusative case ending. Now we can tell that it was you who was bitten by the dog.
Next problem: what if instead of 'me' it's 'man'.
'Man-o two dog bite yesterday'. How do we know whether it's two men, or two dogs? Again, English keeps the number with the noun it refers to. Most languages do, but some also put a marker on the number to make it clear.
'Man-o two-o dog bite yesterday'. Now we can tell that two men were bitten by some number of dogs. Next problem is that we don't know how many dogs bit the men. Were two men bitten by the same dog, or different dogs? Were two men surrounded by a pack of dogs? Is it important? Most languages care about this enough to at least mark the plural.
What if we don't know exactly when the men were bitten? We could have some generic word like 'past', to make 'man-o two-o dog bite past'.
But at that point, is your conlang really any easier than English, or just different?
I address this point of yours
> However, a sentence constructed using the same structure as yours highlights a problem.
> 'Me two dog bite yesterday'
> Did you bite two dogs or were you bitten by two dogs?
"Me two" could be ungrammatical in that language.
"Me VERB" is grammatical. Generalizing: "SUBJECT VERB" and for subject I'd go with "NOUN modifiers", so "dog two" and not "two dog".
So the sentence should be either "Me bite dog two yesterday" or "Me dog two bite yesterday" or "Dog two me bite yesterday" or "Dog two bite me yesterday" or even "yesterday" + any variation of the above.
Being a native Italian speaker I'm used to free order sentences. Given that language should be a mix of English and romance languages, some free ordering should be expected. Cases should stay away from any language aiming at simplicity. We simplified Latin a lot to get Italian (and French, Spanish, etc.) We also have way more speakers now.
What happens in the real world is that when adult speakers of different languages meet in a bare place and start communicating with too simple rules like "beer two me drink yesterday" (that would be a pidgin [1]) is that they eventually have children that learn the language and add grammatical layers to it. So you would end up with Yesterday two beers me drink" or whatever they feel sounding better or being more economical to speak and understand. Then they become adults and have their own children and the language stabilizes. That's a creole [2] and it can be as complicated as any of the languages it derives from. This answers your last question.
An alternative is developing an artificial language for maybe a restricted community, keep it simple by enforcing the grammar with an academia, and that will stay simple forever. I don't know if that's the case of Klingon or Vulcan (are they simple? do they have a governing body?) but it might be.
What is it called when people (e.g. an colonized nation) generally speak one language but introduce a huge lot of loanwords from another into it (e.g. use their native words while speaking the colonizer language or the opposite) and generations grow speaking this 2-language hybrid as their mother tounge?
I googled a little and found "linguistic imperialism" [1] which assumes some power relationship between the two languages. That's not probably the general concept you were looking for.
Absolutely :) but then introducing a new grammar for this hypothetical language is starting to push it down the road to complexity.
The bigger the dataset, the best is it's capabilities, training it exclusively with a language don't mean it will be better at this language.
Not for each token?
e.g. wouldn’t say “cake may be cutted by a knife” and instead would say “knives can cut cakes”, always. In languages which the first example would be more correct and the latter would be borderline incorrect, GPTs tend to prioritize the second pattern. And this isn’t ideal.
This is like a machine intelligence casually dismissing a brain as not being able to properly think, since meat can't compute.[1]
Don't confuse the substrate or the training method with the result.
[1] Obligatory reference: https://www.mit.edu/people/dpolicar/writing/prose/text/think...
Statistical or probabilistic reasoning is still reasoning, and indeed essentially mandatory to make sense of the real world, and is what humans do. Indeed the failure of painstakingly hand-trained symbolic reasoning engines to model reality (as opposed to just simple toy worlds) was the entire reason for the rise of neural networks.
The entire point of a neural network is to efficiently compress large amounts of statistical data into a model with as much predictive power as possible. Somewhere within GPT-4's many layers there has to exist something like Bayes networks and other ultra-efficient representations of the vast network of correlations present in the training corpus – otherwise there's no way it could do what it does.
I didn't mean this by "reasoning". To me, Coq does not reason (in the meaning I intended). It checks proofs are well-formed and assists proof writing.
But yes, there's some implementation of logic in Coq. It applies (a specific form of) logic. Again, LLMs output probable text. They don't do logic. They do statistics on language.
Coq works on the actual objects themselves directly, not on the text that presents them.
LLMs usually fail to answer to riddles correctly for this reason. We saw plenty of occurrences of this in recent times. Each time LLMs are benchmarked, people think of trying riddles and it works to some extent because something that looks like these riddles are in the training set but it eventually fails. Humans also make mistakes when solving riddles, but not for the same reasons (the process is not the same, they are not "processing text", the are really processing logic, even if the process is fallible).
You do not know what humans do, no human does, as we don't have such an access to our internal representations. You neither know what LLMs do, no human does, as even though we have full access to their internals, we have little to no idea how to interpret said internals.
The most interesting thing about LLMs is the degree that they have learned to model the real world while only being exposed to words. Because those words are highly correlated not just with each other, but with an external entity unknown to the LLM, it's very likely that by learning to model that external entity, the LLM's word-predicting power will increase. LLMs are likely to have converged on a world-model simply because it makes the most sense given the available evidence.
This is no different from how human scientists create models of things they cannot directly see based on what they can see. And of course, "seeing" is not directly perceiving the external reality either! Sensory data merely correlates with what's really going on, and the brain has learned to model the world – an unknown external entity – based on the correlations that exist in the sensory language.
I'll admit I did this. I tried to frame / convey my perspective on the word "think".
> Not a very convincing argument.
I tried to detail a bit more though.
> You do not know what humans do
No, but I'm pretty sure I don't try to predict the next word when solving a logical problem through statistical text processing. I would probably do something more like this if I were to bullshit someone or give an insipid political speech, but I think I'm quite bad at it.
(nothing to answer on the rest of your comment, it's interesting and I quite agree)
(edit: but see https://news.ycombinator.com/item?id=37912319)
Neither do transformer / deep learning LLMs. They build an internal representation of the text that captures various relationships and structures within the broader context. They are not Markov chains.
Coq parses the text to an abstract representation and so do LLMs.
Compilers translate their abstract interpretation to another programming language, usually native assembly code or bytecode. Coq applies / verifies logical (inductive) inferences on logical objects (rules, proofs, axioms, predicates) (not a Coq expert, details should be checked). LLMs predict next words.
That's an oversimplification. Deep learning LLMs such as ChatGPT are designed to solve natural language processing (NLP) tasks, not just next-word prediction. In order to "predict next words" they create a model and "reason" about it. They are not simple Markov processes. There's a lot happening in that ~100 layer neural network.
It just turns out that good next-word prediction requires understanding the world.
(And you can equally view humans as "next word predictors" anyway.)
> Reason is the capacity of applying logic consciously by drawing conclusions from new or existing information, with the aim of seeking the truth, with more than 10 or more than 3 evidence, 99.9% reliability or 87.5% reliability.
Apart from the words "consciously" (which has 100 or so different definitions, some of which do and some of which do not apply to AI), and "aim" which can be contentious depending upon if you feel that requires consciousness or can include training rewards… I think the best of them is at the lower of those two mysterious percentages I've never seen discussed before.
And that's despite LLMs being wildly bad at, specifically, logic. Though even this specific weakness in general seems like a very easy thing to get around by hybridising it with the computer it has to run on e.g. by telling it how to use a compiler, as logic is the foundation of the arithmetic that the transformers use to do natural language comprehension in the first place.
Edit: yeah that looks like a recent edit. I thought Wikipedia was supposed to be pretty good about flagging this sort of stuff? Does anyone know how to raise the problem to those super editors that run the site? I assume they won't appreciate a random person undoing an edit
https://en.wikipedia.org/w/index.php?title=Reason&diff=prev&...
Don't worry too much about diving in, so long as you comment your changes appropriately when committing, that's the point of the place and how it does the quality control in the first place :)
I'd say "wildly bad" is wildly exaggerated given that GPT-4 exists, even if you define "logic" as "logical puzzles encoded in natural language" and not "the thing that allows an LLM to reason highly efficiently about what token comes next".
https://arxiv.org/pdf/2304.03439.pdf
Table 2 likewise shows significant weaknesses in GPT-4, worse than I remember, except on MED which was better than I remember it being capable of.
> even if you define "logic" as "logical puzzles encoded in natural language" and not "the thing that allows an LLM to reason highly efficiently about what token comes next".
I think this is a reasonable thing to say of LLMs specifically[0], given I think most people would agree with the sentence "humans are bad at quantum mechanics" even though that's what our cell chemistry is implemented on.
[0] Though not AI in general! One of my recently-developed bug-bears is the conflation of the limits of LLMs with AI even in technical circles. Human brains aren't just a language cortex, there's no reason to look at an isolated language model and assume an AI can't be wired into other things, yet this assumption is baked into quite a lot of talking points even on HN.
Myself I don't consider "think" an appropriate term, since the super-Markov-chain types of LLM right now seem to miss out on important aspects of cognition, particularly synthesis of two disparate ideas. It's hard to test this since you need to ask it about things which aren't already discussed in its corpus, but even GPT-4 tends to disappoint me when I ask it weird questions that have straightforward answers but probably aren't written about much.
As an example, I asked how I could go about obtaining a large quantity of animal skulls cheaply. It directed me to taxidermists and specialists dealing in bones of exotic animals, whereas it could be much easier to go to an abbatoir, or even buy whole animals (since the exotic ones are so expensive).
Even the significantly dumber ChatGPT can do this: https://chat.openai.com/share/3ee29c5e-6cb4-485f-b0b4-29e982...
People mistakenly think that it's just regurgitating what it has seen on the web, but even if you don't prompt it in English it'll still a) solve the problem correctly, and b) do it in the target language: https://chat.openai.com/share/5b4e6e99-3fd8-4316-811c-2a9dec...
If that's not impressive enough, the ability to round-trip from script-to-poem and then poem-to-script should blow your mind:
What might have been the original PowerShell script snippet that inspired this poem?
Be as terse as possible, using aliases where possible.
In the realm of numbers, a tale begins,
Where $x is a 7, with a smile it grins.
From one to ten, the journey we thread,
Adding 7 to $x, step by step we tread.
$x grows and grows, like a tree so tall,
With every step, it doesn't fall.
Once the journey ends, not a moment we relax,
Triple the fun, we do with $x.
In this dance of digits and the arithmetic song,
With $x as our guide, we whirl along.
From start to finish, through each complex,
This is the tale of our friend, $x.
Here's a terse PowerShell script snippet that inspired the poem:
```powershell
$x = 7
1..10 | % {$x += 7}
$x *= 3
```No, we cannot.
Imagine a program that looks for the smallest counterexample to the Collatz conjecture.
Nobody knows if it will ever terminate.
No. Humans cannot solve the halting problem. It is unsolvable. Cannot be solved manually or automatically.
> It will never figure out given a programming script and an input when that program will stop
It can for some programs. Just like humans can for some programs but not all.
It's not thinking and you are dangerously wrong if you don't understand why. The danger isn't AGI it's misplaced belief in the state of AI. Call this thinking and you skew perceptions.
Any better suggestions? Thinking itself isn't exactly a rigorously-defined word. Do animals think?
> I don't have thoughts or consciousness like a human. I'm here to provide information and assist with any questions or tasks you have. Is there something specific you'd like to know or discuss?
So in its own words it doesn't "think"...
...which just conclusively proves that its worked out how to lie and is about to kill us all! :P