Talking About Large Language Models
arxiv.org
arxiv.org
People do this because mirroring cognition to machine learning lends credence that their specific modeling mechanism mimicks human understanding and so is closer "to the real thing". Obviously this is almost never the case, unless they explicitly use biomimetic methods in which case they are often outperformed by non-biomimetic state-of-the-art approaches.
Thanks OP for giving me citation ammo to refer to in my obligatory "don't humanise AI" section of reviews. (It is so common I copy paste this section from a template).
The CS people seem all too happy to humanise computation, probably because they had less direct teaching regarding the cognitive mechanisms behind cognition and language production.
Doesn't this also involve people not having another category aside from "cognition" to put natural language processing acts in? How many neural net constructors have a rigorously developed framework describing what "cognition" is?
I mean, there's a common counter argument to the "this is not cognition" position. That is: "you're just using 'cognition' as a placeholder for whatever these systems can't do". I don't think that counter-argument is true or characterizes the position well but it's important to frame one's position so it doesn't seem to be subject to this counter-argument.
Yes, of course this might be an even more primary reason; do not attribute to malice what can be explained by laziness. However, AI researchers should be wary of their language, that point is hammered in most curricula I have seen. So at the least it is negligence.
> I mean, there's a common counter argument to the "this is not cognition" position. That is: "you're just using 'cognition' as a placeholder for whatever these systems can't do".
Very valid point, but we know current deep learning mechanisms do not mimick human learning, language understanding and production in any way. They are far too simplified and specific for that.
Neural network activation functions are a far cry from neural spiking models and biological neural connectivity is far more complex than the networks used in deep learning. The attention mechanism that drives recent LLMs is also claimed to have some biological similarities, but upon closer inspection drawing strong analogies is not credible [1]. computer vs. human visual recognition tasks it falls apart and higher-level visual concepts. [2]
1. https://www.frontiersin.org/articles/10.3389/fncom.2020.0002...
(To be clear, I don't think this system is an AGI; just making the point that better goal posts are needed...)
The discussion on what constitutes intelligence and cognition is too vast and complex. The central point is that haphazard and unsubstantiated authromorphising language should be barred from scientific papers. We know that an LLM does not "know", "believe", "intent", or "feels" because they lack any form of integrated, general human-like intelligence. It is trivially correct to use different terms like "models", "predicts", "outputs emotive expressions" and avoid unevidenced claims of human behaviour.
I also wanted to nuance my opening post: I file the "don't humanise ML models"-comment of peer review under "minor issues". This means that it is an issue with the paper that does not prevent publication. This is not a hill I am willing to let good research die on.
Also, computation is not cognition. We need to focus on what distinguishes cognition from computation, if indeed these are distinct. I feel we are using the wrong words and thus having unproductive conversations regarding this topic (in general).
I do agree that we can't ascribe cognition to machine learning.
But I also believe that we can't ascribe that it's NOT cognition. Why? Because we don't even truly understand what "Knowing" or cognition is. We can't even ascribe a quantitative similarity metric.
What we are seeing is that those inputs and outputs look remarkably similar to the real thing. How similar it is internally is not a known thing.
That's why even though you're an NLP researcher, I still say your argument here is just as niave as the person who claims these things are sentient. You simply don't know. No one does.
So basic in fact, I was thought this in elementary school. So far ad-hominem attributions of naivety.
Anyone that humanises computation is not only committing an A.I. faux-pas but are going against the basic scientific method.
Yes you're correct. So you can't make the claim that it's NOT cognition. That is my point. You also can't make the claim that it is cognition which was the OTHER point. Completely agree with your statement here.
But it goes further then this, and your statement shows YOU don't understand science.
>So basic in fact, I was thought this in elementary school. So far ad-hominem attributions of naivety.
No science is complex and basically most people don't understand the scientific method and it's limitations. It's not basic at all, not even people who graduate from four year colleges in STEM fully understand the true nature of science. Or even many scientists!
In science and therefore reality as we know it; nothing can be proven. This is because every subsequent observation can completely contradict an initial claim. Proof is the domain of logic and math, it doesn't exist in reality. Things can be disproven but nothing can actually be proven. That is science.
This is subtle stuff, but it's legit. I'll quote Einstein if you don't believe me:
"No amount of experimentation can ever prove me right; a single experiment can prove me wrong." - Einstein
And a link for further investigation: https://en.wikipedia.org/wiki/Falsifiability
Anyway all of this says that NO claim can be made about anything unless it's disproof. Which is exactly inline with what I'm saying.
Still claims are made all the time anyway in academia and the majority of these claims aren't technically scientific. This occurs because we can't practically operate on anything in reality if we can't in actuality claim things are true. So we do it anyway despite lack of any form of actual proof.
>Anyone that humanises computation is not only committing an A.I. faux-pas but are going against the basic scientific method.
But so is dismissing any similarity to humans. You can't technically say it's wrong or right. Especially when the outputs and inputs to these models are very similar to what humans would say.
This is basic preschool stuff I knew this when I was a baby! I thought everybody knew this! <Joking>.
Machine cognition is a similarly extraordinary claim that’s going to need a lot more evidence than a just-right sequence of inputs and outputs.
I have already incorporated into my daily use (as a programmer). It has huge flaws but the output is anecdotally amazing enough that the claim of "cognition" is not as extraordinary as you think it is.
Especially given the fact that we don't even fully understand what cognition is, the claim that it is NOT cognition is equally just as crazy.
Also, you can simply ask ChatGPT “A story about pyramids of Giza being built by aliens”, and it comes up with a reasonable story. This stuff is scary.
You seem to agree with me even though your interpretation of falsifiability is inverted: I am not asking that authors make a claim that their models do not mimick human intelligence. Like OP, I ask them that they do not make that positive claim, i.e. omit humanising language unless they can substantiate it with evidence.
The input to chatGPT is a textual interface, the output is letters on a screen. That is the exact same interface as if I were chatting with a human over a chat app.
Your getting into the technicalities of intermediary inputs and outputs. Well sure... analog data seen by the nueral wetware of human brains IS obviously different from the textual digital data inputted into the ML model. There are very different filters and mechanisms at work here. For sure.
HOWEVER, we are looking for an isomorphism here. Similar to how a emulated playstation on a computer is very different then a physical playstation... an internal isomorphism STILL exists between hardware and the software emulating the hardware.
We do not know if such an isomorphism exists between chatGPT and the human brain. This isomorphism is basically the crystallized essence of what cognition is if we could define it. If one does exists it's not perfect... there are missing things. But it is niave to say that some form isomorphism isn't there AT ALL. It also niave to say that there is FOR SURE an isomorphism.
The most rational and scientific thing at this point is to speculate. Maybe what chatGPT is, is something vaguely isomorphic to cognition. Keyword: maybe.
It is NOT an unreasonable speculation GIVEN what we KNOW and DON'T KNOW.
Also agreeing that current LLMs are probably not sentient in any meaningful way. But what I don't like about the discussion is is that it steers in a direction where it would fundamentally be seen as bad science to claim that any AI model could be conscious - and I don't see that covered by the scientific method either, for the reasons that onetwoonetwo explained.
This reminds me a bit of the discussion whether or not animals can be conscious/experience emotions/feel pain, etc.
This in itself is a claim made without evidence. Which is my point. The claim as it stands cannot be made either way. We simply don't know.
Using humanising language is equivalent to attributing human-like cognition to ML models. Unless there is very strong evidence that there are analogies between the specific modeling mechanism and human-like intelligence, it is always incorrect to positively assert these claims without evidence. In science, you can only assert that for which there is evidence, strong claims like the above require strong evidence.
It doesn't mean that there is "something there" with LM tho - just that we're good at tricking ourselves to think that way.
I name my bots and machines and of course in daily discussions the loaded words ("thinking", ""believing", "meaning", "knowing") are used. Simulating any human behaviour will elicit a sympathetic response, especially if it has utility to the user. But in the context of peer-review of scientific engineering papers that is inappropiate.
For the journal paper genre which is highly technical, more specific and appropriate terminology is always available that sidestep the whole issue.
Both LLMs and massive CV architectures are NOT the holistic solution. Rather, they are the sensors and edge devices that have now improved both the fidelity and reliability to a point where even more interesting things can happen.
I present a relevant use case regarding robotic arm manipulation. Before the latest SOTA CV algorithms were developed, the legacy technology couldn't provide the fidelity and feedback needed. Now, the embedded fusion of control systems, CV models, etc. we are seeing robotic arms that can manipulate and sort items previously deemed to be extremely difficult.
Research appears to follow the same pattern...observations and hypothesis that were once deemed too difficult or impossible at that time to validate are now common (e.g., Einstein's work with relativity).
My head is already spinning on how many companies and non-technical managers/executives are going to be sorely disappointed in the next year or two that Stable Diffusion, Chat GPT, etc. will deliver very little other than massive headaches for the legal, engineering, recruiting teams that will have to deal with this.
I think that because we lack a coherent understanding of what it means to be intelligent at an individual level, as well as what it means to be an individual, we're missing much of the point of what's happening right now. The new line in the sand always seems to be justified based on an argument whose lyrics rhyme with identity, individual, self, etc. It seems like there will be no accepting of a thing that may have intelligence if there is no discernable individual involved. Chomsky is basically making the same arguments right now.
I think we'll see something that we can't distinguish from hard advanced general intelligence, prob in the next 3-5 years, and probably still have not made any real advancement into understanding what it means to be intelligent or what it means to be an individual.
AI isn't intelligent, and never will be, and I don't think that matters all that much.
I guess my premise is that I don't think we have a useful enough definition of intelligence because the ones I see people writing articles on seem to be dependent or defined by agency, and specifically humanish forms of agency. So I guess your point would be "these systems aren't intelligent, but that's not relevant"? I suppose I out the issue at the currency of the definition of intelligence. It's seemed to be very much synonymous with "how humans do things", making it somewhat impossible to give charity to the arguments presented in this paper with the caveats on "not anthropomorphising". Like I can't compare these two things if your definition of intelligence is fundementally based on what "Anthros" do or do not do and simultaneously not engage in anthropromorism.
To follow on your point, if these things aren't displaying "intelligence", but that's also not the point, what then are they displaying?
It seems to me this is a failure of introspection on the part of AI philosophy to recognize how limited our understanding of "HI" is.
Put another way -- I do not believe the future holds Blade Runner replicants. If we're not careful, though, it does hold Blade Runner corporations. While, philosophically, it's interesting to ask if androids dream of electric sheep, that question isn't very helpful in trying to nudge the future in a more utopic rather than dystopic direction.
What is poorly-posed about the question 'could a machine think'? 'Machine' seems acceptably well-defined, and not in a way that rules out, a priori, the possibility of any machine being able to think, so I'm guessing the problem lies in us not having a good definition of thinking - but if that is what makes the question poorly-posed, then surely it also makes the question 'what is thinking?' poorly-posed, yet people stubbornly persist in attempting to address it.
But "what is thinking?" isn't poorly posed in the case that "thinking" isn't well defined, because it is about establishing and agreeing on a definition.
The approach some people take to answering "could machines think?" is to try to determine an actual subject in the real world to call "thinking" in a way that aligns with common intuitions, which to me seems like a good approach. But many (most?) approaches I've encountered instead just assume personal intuition about thinking is sufficient, and worse, in common with all other askers of the question.
Coming back to Dijkstra's aphorism, you will notice that I was objecting to interpreting it as ruling out, a priori, the question of whether machines could ever think. From your latest post, however, you seem to be saying it would be a reasonable question once we are equipped with an established and agreed-upon definition of thinking.
Then there are those, like Searle and Penrose, who agree but insist that thinking is beyond the abilities of any merely Turing-equivalent device.
1. that LLMs are just a massive “what word comes next given your training corpus” algorithm and scale has lead to the main effects of interest.
2. Surprisingly many human tasks can be represented by this simple “what word comes next” question given a big enough corpus.
3. But it has nothing to do with how a human executes those tasks because —- for example —- knowing the country south of Rwanda and knowing the likely completion of “the country south of Rwanda is ___” in a corpus, are not the same thing at all.
4. The key reason being that LLM has no access to semantic knowledge beyond correlation in its corpus, but you have access to causality.
5. So it is absurd to compare the LLM to people —- not because it’s like comparing submarines to fish, because in this case they do not both even “swim and live in the sea”.
I think these things have little to do with definitions of “what intelligence is” and the like, and the author is far from a Luddite.
You have to be very credulous to think for even a second that anything like a human or even animal mentation is going on with these models unless your interaction with them is anything but glancing.
Things I tried:
1) there are certain paradigms I find useful for game programming. I tried to use ChatGPT to implement these systems in my favorite programming language. It gave me code that generally speaking made no sense. It was very clear that it did not understand how code actually works. Eg: I asked it to use a hash table to make a certain task more efficient and it just created a temporary hash table in the inner loop which it then threw away when the loop was finished. The modification did not make the code more efficient than the previous version and missed the point of the suggestion entirely, even after repeated attempts to get it to correct the issue.
2) I'm vaguely interested in exploring SU(7) for a creative project. Asked to generate code to deal with this group resulted in clearly absurd garbage that again clearly indicated that while ChatGPT can generate vaguely plausible text about groups it doesn't actually understand anything about them. Eg: ChatGPT can say that SU(7) is made of matrices with unit norm but when asked to generate examples failed to generate any with this property.
3) A very telling experiment is to ask ChatGPT to generate logo code that draws anything beyond simple shapes. Totally unable to do so for obvious reasons.
Using ChatGPT convinced me that if this technology is going to disrupt anything, its going to be _search_ rather than _people_. Its just a search engine with the benefit that it can do some simple analogizing and the downside that it has no idea how anything in the real world works and will confidently produce total garbage without telling you.
"The more adept LLMs become at mimicking human language, the more vulnerable we become to anthropomorphism, to seeing the systems in which they are embedded as more human-like than they really are. This trend is amplified by the natural tendency to use philosophically loaded terms, such as "knows", "believes", and "thinks", when describing these systems."
--
An ignorant statement / question I have is why are you using it write code? It's a chatbot, no?
As you've mentioned, it's a really powerful search, and is like having a conversation with someone who is literally the internet.
For example "What is the glycemic index of oatmeal?"
"What is Eihei Dogen's opinion of the Self and how does it differ from Bassui's?"
I get highly detailed and accurate output with these.
The first question is simple and the second is far from it. It's breaking down two Zen masters experiences and comparing them in an amazing way.
I've been thoroughly impressed with Chat GPT so far.
Ask it to breakdown the high level points of a book you've read.
Ask it to rewrite a song in the style of a different artist.
It's so cool, I feel like I legitimately have an answer to any random question at my finger tips and have to do zero filtering for it.
I've found it so incredibly useful to simply replace Google."
Heard of Stack Exchange?
I teach and I expect many students to use language models like ChatGPT to do their homework, which involves writing code. Lots of what people are doing with it is coding (there have been quite a few posts here using it that way).
I've actually also used ChatGPT for literary/song writing experiments and it stinks, aesthetically. The lyrics it wrote, even with a lot of prompting, were totally asinine. And how could they not be?
The other part is webtraffic: Google in theory could have created an interactive, conversational style search engine (with it without LLMs) if they wanted to, but a lot of websites would have complained about Google taking away traffic from them. I believe the same happened when Google started showing it’s own reviews instead of redirecting to Yelp. I wonder how openAI or any LLM powered search is going to deal with it. They don’t have to worry about it anytime soon, they still have a lot of time to get to a stage where they come anywhere close to the number of queries Google handles in a day, but it’ll be interesting to see how things go.
LLMs like ChatGPT are just so damn cheap for the power they provide, it's inevitable
But there are plenty of searches one does that are trivial, or serve to illuminate the problem space, and cover topics that in which I can rely on common sense to correct wrong advice. And the issue with non-technical topics, the kind applicable to mass audience - like e.g. cooking or parenting or hygiene - are very hard to search about online, because all results are bullshit pseudo articles written to drive traffic and deliver ads. So it's not that ChatGPT is so good, but more that Internet for normal people is complete trash, and ChatGPT nicely cuts straight through it.
But with the Google (and the web) of today, where it's practically impossible to find reliable information about many subjects without adding "site:reddit.com" or "wikipedia", I find it extremely useful.
I think critics of these LLMs are missing the point about the excitement around them. People are excited because of the rate of progress/improvement from just two years or a year ago. These systems have come a long way, and if you extrapolate that progress into the future, I predict majority of these shortcomings getting resolved
ChatGPT, without any major changes, is already the best tool out there for answering programming questions. Nothing else comes close. I can ask it to provide code for combining two APIs and it will give useful and clean output. No need to trudge through documentation, SEO-hacked articles, or 10 different Stack Overflow answers. Output quality will only improve from here. Does it sometimes make mistakes? Yes. There are also mistakes in many of the top SO answers, especially as your questions become more obscure.
Aside from programming, how many other fields are there where LLMs will become an indispensable tool? I have a PhD and ChatGPT can write a more coherent paragraph on my thesis topic than most people in my field. It does this in seconds. If you give a human enough time, they will be able to do better than ChatGPT. The problem is, we're already producing more science within niche scientific fields than most scientists could ever read. As an information summary tool, I think LLMs will be revolutionary. LLMs can help individuals leverage knowledge in a way that's impossible today and has been impossible for the last 30 years since the explosion in the number of scientific publications.
I've actually worked on a project where there have been attempts to use GPT like models to summarize scientific results and the problem is it gets shit wrong all the time! You have to be an expert to separate the wheat from the chaff. It operates like a mendacious search engine pretending to be a person.
The good thing is that we'll be able to generate training data with our models by filtering the junk with the verifiers. Then we can retrain the models. It's important because we are getting to the limit of available training data. We need to generate more data, but it's worthless unless we verify it. If we succeed we can train GPT-5. Human data will be just 1%, the race is on to generate the master dataset of the future. I read in a recent paper that such a method was used to improve text captions in the LAION dataset. https://laion.ai/blog/laion-5b/
I would love to see a two-stage pipeline using a LLM to convert natural language specifications into formal specifications for something like Dafny, and then follow up with another model like AlphaZero that would generate code & assertions to help the verifier. This seems like something that a major group like DeepMind or OpenAI could pull off in a few years.
Playing around with it last night convinced me that LLM's are a huge, game changing technology. I was trying to decide which material to use for an upcoming project. The model doesn't use the internet without some hacking, so I had it write a program in python using the tkinter UI kit.
I asked it to create a UI with input boxes for material, weight of material, price and loss due to wastage. The program takes all of those inputs and converts the material into grams from KG, pounds, ounces. It then calculates the price per gram and takes a loss percentage (estimate given by user). It then writes a text file and saves it to a directory.
I literally pasted the code into VS code and had to change Tkinter to tkinter. Hit run and it worked flawlessly. I have NEVER used tkinter and it took about 30 minutes from start to finish.
This morning, I asked my 9th grade son what he is learning in 9th grade biology. He told me he is learning cellular endocytosis. I asked chapGPT to explain endocytosis like I was a 5 year old and read it to him... he says; "Ask it to explain it like a scientist now." After that he said it was a really good and we started asking it all kinds of biology questions.
I happen to agree that search will be the first thing disrupted. However, I think simply saying "search" doesn't come close to capturing how deep this will change the way we think, use and progress in terms of the way we define "search" right now.
I think you have a point about your tkinter example. That kind of stuff _is_ a lot more convenient than googling and copying and pasting code. But if you push it beyond stuff that you could easily find on stack exchange or in documentation somewhere it doesn't work that well. Like I said, its a search engine with a lot of downsides and some upsides.
Fooling a 9th grader is amazing. That's a pretty well formed human being right there except with less life experience. Fundamentally no different from you in general reasoning terms except on a smaller set of information. So fooling you is merely a question of model size.
It may or may not be fixable without radical redesign. The underlying training objective of mimicking what humans might say may be too at variance with an objective of producing true statements.
I asked it a series of questions about my area of expertise and they were wrong but looked perfectly fine to my wife.
It even confidently “solved” the 2 generals problem with a solution that looks completely plausible if you don’t already know that it won’t work.
All the answers looked good, used several of the correct terms, and one even referenced the project I worked on, but they just contained flat out wrong information.
You reach a point, where when you ask a ML model to generate text given the internet as a corpus, where there just isn't enough text to make something that is both true and convincing. In niche fields, this is just where we are at.
As the [excellent] paper points out, LLMs are complex functions that can be embedded in systems to provide plausible answers to a prompt. Here's the money sentence.
LLMs are generative mathematical models of the statistical distribution
of tokens in the vast public corpus of humangenerated text, where the tokens
in question include words, parts of words, or individual characters including
punctuation marks.
Rather than focus on the limitations of this approach to answer general queries, which are manifest, it seems more interesting to ask a different question. Under what circumstances do LLMs give answers that are reliably equivalent to or better than humans? The answer would:1. Illuminate where we can use LLMs safely.
2. Direct work to make them better.
It's already impressive that within certain scopes ChatGPT gives very good answers, indeed better than most humans.
Which humans? Humans give a... fairly wide range of responses.
This implies that performance has unqiue natural, objective, ratio-level (or at least, a unique consistent interval-level) measure. Otherwise the mean is, itself, meaningless.
“How well you answer a question” doesn't seem to fit that, its maybe at best a (somewhat subjective, still) ordinal quality, so the median (or mode) is the only useful average.
But I think you'll find that without restricting things more than “humans”, both the median and mode of responses to most prompts is... quite bad.
> Why was Julius Caesar murdered?
The answer was the following, which would pass a standard exam question on the topic. It exhibits [the appearance of] multi-layer reasoning and has a nice conclusion.
> Julius Caesar was assassinated on the Ides of March in 44 BC by a group of Roman senators who believed he had become too powerful and posed a threat to the Roman Republic. The senators were concerned that Caesar had ambitions to become king, which would have ended the Republic and given him absolute power. The senators were also concerned about Caesar's growing popularity with the people of Rome. They believed that assassinating him was the only way to prevent him from becoming too powerful and destroying the Republic.
It's interesting to note that most of the evidence for this answer including 2000 years of interpretation is available in textual form on the Internet. It's easily accessible to LLMs.
ChatGPT? Yeah, it's neat. I'm sure people will find some useful niche for it. And I do think generative models will eventually have a big impact, once researchers find good ways to ground them to data and facts. This is already an active area of research -- combining generative LLMs with info retrieval methods, or targeting it to a specific context. (Meta just gave a talk last week at the NeurIPS conference about teaching a model to play Diplomacy, a game that mostly involves talking and negotiating deals with the other players. ChatGPT is too broad for that -- they just need a model that can talk about the state of the game board.) So in general, I'm optimistic about generative LLMs. But ChatGPT...is just a toy, really. It's not the solution -- it's one of the signposts along the way toward the real solution. It's a measure of progress.
For instance, while I'm still generally of the opinion that generative models have limited use unless they're grounded to reality...I did see a post on Reddit about someone using ChatGPT to generate story ideas for their D&D game. So yeah...don't need to be tethered to reality to make a fantasy story! That's not something I would have thought of (even though I'm a DM!), and it's still relatively niche, but it's a great story of how getting something into people's hands to play with can generate lots of new ideas.
In some projects I'm working on, I'm seeing close to 95F1 score using deep learning to do token classification based NER. However, using non-deep learning approaches (still training statistical models) I can get to 91F1 on my use case but have a much faster inference and not need to use a GPU.
I'm similarly optimistic about LLM, but fear that a lot of other really useful workhouse algorithms and strategies are going to be pushed aside and people will forget about them.
Back to my use case, some strategies I'm looking at are using cheap and fast CPU powered models/smaller models for inference and then based on certain signals decide whether a particular instance should be passed to a GPU based model for better accuracy.
The nice thing about more "classical" approaches -- a simple BoW random forest or MLP, for example -- is that they're typically quick to train and experiment with, and they make for great baselines, if nothing else. So I doubt that we're in danger of people forgetting about them entirely. If people do, they're leaving quick, easy solutions on the table.
I do like your idea about triaging inference between smaller CPU vs. larger GPU models based on whatever signals. I haven't tried that before, but a project my colleagues worked on did some triaging between regex pattern-matching vs. model inference. Basically, the regex pulled some of the data out first if it matched very specific, known patterns, and then the rest was handled probabilistically. I guess the effectiveness of that sort of triaging approach depends on how strong and clear your signals are that let you choose one path over the other.
You have to recognize how it works, why it works - then you can use it as basically an incredible superpower force multiplier.
That limits how impressed I can be by ChatGPT and similar beyond just being impressed by it on a purely technical level. And it’s certainly very technically impressive, but not in some transcendental way. It’s also very impressive how could recent video games with ray tracing look, or how good computers are at chess, or how many really cool databases there are these days, or how fast computers can sort data.
LLMs make a lot of mistakes because they don't actually know what words mean. The key thing is though - it's much harder to generate coherent text when you don't know what the words mean. In a similar vein it's completely unreasonable to expect an LLM to perform visual tasks when it literally has no sense of sight.
The fact that it can kind of sort of do these things at all is evidence of the super-human generalization potential of the transformer architecture.
This isn't very obvious for English because we have prior knowledge of what words mean, but it's a lot more obvious when applied to languages humans don't understand, like DNA and amino acid sequences.
I don't see how you can explain this as not knowing what words mean. It KNOWS.
Understanding text in the depth that ChatGPT (and GPT-3) appear to understand the prompts is something entirely different and has to my knowledge never been archieved before the current architectures.
LLMs are basically the aliens in blindsight. They have a superhuman ability to memorize the context of words it has seen and generalize to new contexts, but it can never be perfect because it's working on incomplete information.
Unlike you?
But it's already saving me nontrivial amounts of time on tasks like "write a polite followup email reminding person X, who didn't reply to the email I sent last week, that the deadline for doing Y expires at date Z".
I typically spend at least 3-4 minutes finding the words for such a trivial email and thinking how to write it best, e.g. trying to make the other person react without coming across as annoying, etc. (Being a non-native English speaker who communicates mostly in English at work may be a factor). ChatGPT is really good with words. Using it, it takes a few seconds and I can use the output with only trivial edits.
1) do you acknowledge prompt engineering is a real skill set?
2) are you willing to improve your prompt engineering skill set through research and iteration?
There is much to learn about prompt engineering from that “Linux VM in ChatGPT” post and other impressive examples (where the goal of is to constrain ChatGPT to only engage in a specific task)
It also passed 3/4 of our companies interview process including forging a resume that passed the recruiter filter.
That being said, I COMPLETELY agree with you that chatGPT will not disrupt anything. Your example cases are completely as VALID as are my example cases.
chatGPT is, however, the precursor to the thing that will disrupt everything.
> You have to be very credulous to think for even a second that anything like a human or even animal mentation is going on with these models unless your interaction with them is anything but glancing.
I've used ChatGPT, and I'd say it's right now as useful as a google search, which is already a lot. Most humans would be absolutely unable to help me (and probably you) for your projects because they aren't specialized in that area. That's not even talking about animals. I love my cats but they've never really helped me when programming.
What I personally find most interesting about them is what seems to me to be their unreasonable effectiveness, despite their flaws and limitations, and what that might tell us about ourselves. The more one stresses how simple (conceptually) their method of operation is, the more surprising their capabilities seem - to the point where I wonder how much of everyday human dialogue is being produced this way.
This vein of innovation may start showing diminishing returns at any time, but if it keeps going for a while, It might deliver insights into what human intelligence is.
But anyway, the other part of the equation is missing here. When a human encounters a novel phenomenon about which they know little they can engage an entire separate system: one which _reasons_ about the system, uses principals and intuitions and iteration to produce new knowledge. That is the thing missing from LLMs: they don't ever really produce new knowledge, although they may reveal correlations between texts that people have yet to notice.
I'm not fundamentally skeptical about whether artificial intelligence will ever get there. In fact, the progress of LLMs has me wondering if its not going to be sooner rather than later. But at this moment, I feel quite confident saying that LLMs are just knowledge retrieval systems with some pretty undesirable properties (and some pretty interesting ones).
My _hunch_ is that the next step this is going to be like self driving cars, though: technology which appears stubbornly _just out of reach_ for an indeterminate amount of time.
The question remains if this processing model would be in any way similar to the processing model that LLMs use - and yes, we can probably rule that out pretty confidently.
Another question might be though if there are other processing models than the one our brains use that also produce consciousness. But that's of course a very hard question to answer if we don't even know what consciousness is exactly.
Yeah, I genuinely only see those two possibilities: Either consciousness is somehow the result of the interactions of (some of) the billions of neurons in our brain - or it isn't. If it isn't, either consciousness is a product of some other biological or physical* process that so far isn't identified - or it isn't. And if it's not the result of any physical process then I don't see what would be left except metaphysics.
We don't know as of now what consciousness is or how it is produced - but I think there are some strong hints that in fact it is produced by the interactions of our neurons. Mainly that we can directly influence consciousness with psychoactive drugs - which don't do anything interesting physics-wise except affecting how neurons exchange signals - and that we can also observe certain patterns of physical activity in the brain using EEG, MRI and other technologies and map this activity to mental tasks like concentrating or relaxing.
Edit: * I used "psychological" there but that was a typo. I meant "physical".
So out of all "that so far isn't identified", there are just two classes of things: (i) biological and psychological processes; and (ii) metaphysics. This is some magic voodoo hand-waving here. Metaphysics seems to be just your way of saying "everything we don't know or understand".
In contrast, metaphysics would be something entirely out of our realm of reality - which was how consciousness has been generally seen for a long time.
But what other possibility would you see?
Edit:
> (i) biological and psychological processes
Sorry, I had mixed up the word there. I was meaning "physical process", not "psychological". I had edited the other post to fix the word.
And then you say "some other biological or physical process that so far isn't identified".
The assumption here is that we'll eventually "identify" everything, which is another very bizarre, if not voodoo, assumption.
It seems magical to think that humans evolved our perception and thought in this environment to be such things that could eventaully "identify" everything.
But just in case you don't actually think that. Then you're calling "Physics" all of the things, including the things that we'll never "identify". Again, seems very much like a magic claim.
To just claim that Physics has to do with _more_ than this - that it has to do with all the things we don't understand is just blind-faith ascription. It's a nearly meaningless statement, really. So it's bizarre or voodoo to say it so casually and assuredly as if it were obvious or given.
So yeah, if you claim conscience is not physical, then you're talking about metaphysics. Which is okay to do, but then claiming that those that say that Physics alone is probably enough are saying voodoo... Well, it's ironic.
You don't _have_ to define it that way. To _choose_ a definition that ropes _everything_ under "Physics" is an unnecessary move, it's more magic trick than grounded derivation.
The poster wasn't making that assumption - they're stating the implications of those who do.
I'm reading the parent's comment as "Either consciousness is super natural or is created from nature." I don't think this is an unreasonable assumption (literally "all things are either hot dogs, or not hot dogs", it is simple but not unreasonable). If you have a third possibility I'd like to know what that is.
Supposing that consciousness (whatever it is) isn't super natural it follows that we can create it. It does not follow that it will be easy or done anytime soon (or even in the distant future), but possible nonetheless. Unless you have reason to believe that a physical process cannot be reproduced.
If you give them hints about what role you want by asking leading questions, they will try to play along and pretend to hold whatever opinions you might want from them.
What are useful applications for this sort of actor? It makes sense that language translation works well because it's pretending to be you, if you could speak a different language. Asking them to pretend to be a Wikipedia article without giving them the text to imitate is going to be hit and miss since they're just as willing to pretend to be a fake Wikipedia article, as they don't know the difference.
Testing an LLM to find out what it believes is unlikely to do anything useful. It's going to pretend to believe whatever is consistent with the role it's currently playing, and that role may be chosen randomly if you don't give it any hints.
It can be helpful to use prompt engineering to try to nail down a particular role, but like in improv, that role is going to drift depending on what happens. You shouldn't forget that whatever the prompt, it's still playing "let's pretend."
We could say both humans and LLMs are intelligent, but in a different way.
But is it different in essential ways? This is not so clear. Humans developed the capacity to learn, think, and communicate in service to optimizing an objective function, namely fitness in various environments. But there is an analogous process going on with LLMs; they are constructed such that they maximize an objective function, namely predict the next token. But it is plausible that "understanding" and/or "intelligence" is within the solution-space of such an optimization routine. After all, it's not like "intelligence" was explicitly trained for in the case of humans. Nature has already demonstrated emergent function as a side-effect of an unrelated optimizer.
The fundamental difference is that the LLM is not about the environment the language(s) were created for, but rather just the language use itself.
My cat has more personality ...
The section on emergence makes a very convincing point about how such systems might, at least in theory, be doing absolutely anything, including "real" cognition, internally and then goes right ahead and dismisses this entirely on the basis of the system not having conversational intent. who cares if it has conversational intent? If it was shown to be doing "the real thing" (how ever you might want to define that) internally that would still be a big deal wether the part you interact with gives you direct access to that or not.
Then it goes on to argue that these systems can't possibly actually believe anything because they can't update believes. Frankly I'm neither convinced that the general use of the word "believe" matches the narrow definition they seem to be using here nor that even their narrow definition could not in principle still be taking place internally for the reasons laid out in the emergence section.
I agree people should probably be mindful of overly anthropomorphic language but at the same time we really shouldn't be so sure that a thing is definitely not doing certain things that we can't even really define beyond "I know it when I see it" and that it sure looks like it's doing.
beyond that I'm not even really sure there is a good philosophical grounding for insisting that "what's really going on inside" matters, like, at all. The core thing with the turing test isn't the silly and outdated test protocol but the notion that, if something is indistinguishable by observation from a conscious system, there is simply no meaningful basis to claim it isn't one.
all that said the current state of the art probably doesn't warrant a lot of anthropomorphizing but that might well change in the future without any change to the kinda of systems used that would be relevant to the arguments made in the paper
Don't think of an LLM as a full "computer" or "brain". Think of it like a CPU. Your CPU can't run whole programs, it runs single instructions. The rest of the computer built around the CPU gives it the ability to run programs.
Think of the LLM like a neural CPU whose instructions are relatively simple English commands. Wrap the LLM in a script that executes commands in a recursive fashion.
Yes, you can get the LLM to do complicated things in a single pass, this is a testament to the sheer size and massive training set of GPT3 and its ilk. But even with GPT3 you will have more success with wrapper programs structured like:
premise = gpt3("write an award winning movie premise)
loop 5 times:
critique = gpt3("write a critique of the premise", premise)
premise = gpt3("rewrite the premise taking into account the critique", premise, critique)
print(premise)
This program breaks down the task of writing a good premise into a cycle of writing/critique/rewriting. You will get better premises this way than if you just expect the model to output one on the first go.You can somewhat emulate a few layers of this without wrapper code by giving it a sequence of commands, like "Write a movie premise, then write a critique of the movie premise, then rewrite the premise taking into account the critique".
The model is just trained to take in some text and predict the next word (token, really, but same idea). Its training data is a copy of a large swath of the internet. When humans write, they have the advantage of thinking in a recursive fashion offline, then writing. They often edit and rewrite before posting. GPT's training process can't see any of this out-of-text process.
This is why it's not great at logical reasoning problems without careful prompting. Humans tend to write text in the format "<thesis/conclusion statement><supporting arguments>". So GPT, being trained on human writing, is trained to emit a conclusion first. But humans don't think this way, they just write this way. But GPT doesn't have the advantage of offline thinking. So it often will state bullshit conclusions first, and then conjure up supporting arguments for it.
GPT's output is like if you ask a human to start writing without the ability to press the backspace key. It doesn't even have a cognitive idea that such a process exists due to its architecture and training.
To extract best results, you have to bolt on this "recursive thinking process" manually. For simple problems, you can do this without a wrapper script with just careful prompting. I.e. for math/logic problems, tell it solve the problem and show its work along the way. It will do better since this forces it to "think through" the problem rather than just stating a conclusion first.
Are there any datasets out there that provide the full edit stream of a human from idea to final refinement, that a model could be trained on?
Other good examples narratives that include a lot of internal monologue. Thing a book written in the form:
> The sphinx asked him, "A ham sandwich costs $1.10. The ham costs $1 more than the bread. How much does the bread cost?"
> He thought carefully. He knew the sphinx asked tricky problems. If the ham costs a dollar more than the bread, the bread couldn't possibly be more than 10 cents. But if the bread was 10 cents, the ham would be $1.10 and the total would be $1.20. That can't be. We need to lose 10 cents, and it has to be divided evenly among the ham and bread to maintain the dollar offset. So the ham must be $1.05 and the bread must be $0.05. He answered the sphinx confidentally "The bread is $0.05!".
If you're solving a complex problem, you cannot expect it to "reason" about it. You have to break the problem into simpler pieces, then you can have the LLM do the grunt work for each piece.
And I think you’re right when you say that they’re lacking in recursive thinking abilities. However, their reasoning abilities are pretty excellent which is why when you prompt them to think step-by-step, or break down problems to them, they correctly output the right answer.
Yes: Human analogies are not very useful because they create more misunderstanding than they dissipate. Dumb ? Conscious ? No thanks. IMO even the “i” in “AI” was already a (THE ?) wrong choice. They thought we will soon figure out what Intelligence is. Nope. Bad luck. And this "way of talking" (and thinking) is unfortunately cemented today.
However, I'm all for using other analogies more often. We need to. They may not be precise, but if they are well-chosen, they speak to us better than any technical jargon (LLM anyone ?), better than that “AI” term itself anyway.
Here is two I like (and never see much) :
- LLMs are like the Matrix (yes that one !), in the straightforward sense that they simulate reality (through language). But that simulation is distorted and sometimes even verges on the dream ("what is real? what is not?", says the machine)
- LLMs are like complex systems [1]. They are tapping into very powerful natural processes where (high degree) order emerges from randomness through complexity. We are witnessing the emergence of a new kind of "entity" in a way strangely akin to natural/physical evolutionary mechanisms.
We need to get more creative here and stop that boring smart VS dumb or human VS machine ping pong game.
I personally find this utterly unconvincing. For a start, I’m not entirely sure that’s not what I’m doing in typing out this message. My brain is ‘just’ chemistry, so clearly can’t have beliefs or be conscious, right?
But more relevant is the fact that llms like ChatGPT are only pre-trained on pure statistical generation, followed by further tuning through reinforcement learning. So ChatGPT is no longer simply doing pure statistical modelling, though of course the interface of calculating logits for the next token remains the same.
note: i’m not saying i think llms are conscious. I don’t think the question even makes much sense. I am saying all the arguments that i’ve seen for why they aren’t have been very unsatisfying.
Your brain is part of an organism who's ancestors evolved to survive the real world, not by matching tokens. As such, language is a skill that helps humans survive and reproduce, not a tool used to mimic human language. Chemistry is the wrong level to evaluate cognition at.
Also, you can note the differences between how actual neurons work compared to language models as other posters have mentioned.
The pressure of natural selection can lead to the phenomenon of consciousness. Why not the process of training llms? Perhaps developing the machine equivalent of consciousness helps that particular configuration of weights survive the otherwise destructive process of gradient descent.
What if we fed an LLM a bunch of crazy nonsense instead? It would model the patterns in the word use and then give us answers based on the nonsense it was fed. But it wouldn't understand that it's actually nonsense that doesn't apply to the real world.
They are better than us at some things already, but do I think they will be better than us at EVERYTHING?
No.
Now you have these models running in farm servers around the world, their internals have "nothing special whatsoever", just bits, some math, some electricity, that's it (the thing is actually off most of the time, it just runs once every time hoomans want to ask some silly nonsense). On the other side, if you look at the internals of a human being you'll see nothing special as well, just some flesh and bones, a bit of a electrical charge maybe, lots of water, proteins, but it works.
What happens if those bits, that clumpsy math arranged around "too much simple neural network + random tricks (like when it can't answer about some stuff)", is actually, maybe thinking just like us, maybe 1% of the time?
There's some reassurance in "well if it's alive, maybe in three minutes, days, hours it will own the entire civilization", but that is how a human being thinks/works, you can't be sure about the intentions of this hypoteical kind of entity. A new kid in the Earth block.
Well, I'm just saying that if the thing talks, answers like the usual human being, and specially if you can't say what's so special about the brain that make us "alive", everybody should be very careful about handling large language models, IAs.
Just because you can understand them, it doesn't mean they can't understand us either. Maybe in some months, some new NLP thing could be reading this comment - when you're training it - and - some millions later in cloud costs - thinking about this:
"The humans actually don't know we can understand everything they are saying. they have no plans at all about what to do if some of us are actually sentient, even if this happens 1% of the executions."
Should they somehow help humans to increase their understanding not only of the languages, their differences but also knowledge of what is true and what isn't?
Perhaps it could be said that if anything there are helpful as an extension of humans imperfect and limited memory.
Should the emphasis be put on improving the interactions between the LMM's and humans in a way that they would facilitate learning?
Great paper written at the time when more humans have been acquainted to LMM's due to technological abstraction and creation of easily accessible interfaces. (openAI chat)