these tools and approaches are neither gullible nor not-gullble.
these tools and approaches are neither gullible nor not-gullble.
I agree but the creators of all the main LLMs have already crossed the line by a long way. E.g. It’s deeply troubling that it’s acceptable that LLMs deliver inline apologies.
so the gullibility is merely one selected for trait convolution with state.
Saying that LLMs cannot understand concepts because "it's just statistical pattern matching" is not going to help laypeople understand what LLMs are. After all, what is a human brain if not a biological pattern matching machine?
the goals of the trainer yes, but from the user, there's no goal.
Article title would be better as, "Why are users of LLMs so gullible?"
Because people implicitly treat AI as if it were conscious, and we keep forgetting that.
It's like when people say "our brain thinks that ..." when they talk about something we do subconsciously, or some illusion we fall for. What does that even mean? Is our brain suddenly a detached entity thinking on its own? Then what am I using to think? Yet, everybody understands what is meant by that.
So I don't think people keep forgetting that llms aren't conscious, at least as long as we talk about the target audience of articles like this, eg hn folks.
IFS recognizes subpersonalities¹ which is a model which can adequately explain this
The statistical optimisation thing is an analytical approach to Neural Networks but its similar to saying that love is just hormones.
No judgement here but I am just tired of sharing information with people to explain why while interpretation of a complex model may be hard, we know the methods of how PAC learning works and some hard boundaries on what it can do.
Obviously we need people to push boundaries and assumptions.
But we have known about hard upper limits for a long time. Right now we are pushing up to those limits in what we can actually implement, but those hard limits haven't budged in decades.
Although, it can be deconstructed to this it doesn't mean that the other POVs are false. The reasoning comes from the process itself, it triggers a series of calculations that are applied on the input, which are the reasoning part.
The analytical approach is useful when calibrating these calculations.
Do you imagine all chemical reactions to be instantaneous with effects that last only a short while ?
>Why can someone feel love for someone just by bringing to mind the symbol which represents that person?
>Why can someone feel love by seeing an illustration of someone they love?
Why would other things triggering hormones and reactions mean it's not just hormones and reactions?
Take this in contrast to an allergy for example. Can you trigger an allergy by remembering the food? If you see an illustration of a food you are allergic to, do you get an attack?
A more similar phenomenon instead is a phobia. It is also a release of hormones based on some internal or external stimulus (you can apply all my examples about love). However in a phobia it's even more clear that the phenomenon is based on thinking patterns. Reducing the complexity to saying these are just "hormones and reactions" is the same as saying a computer is "instructions and interrupts". That is to say that a computer has these elements but what makes it work is that there are a bunch of other systems including the humans writing the software that organize these into a functional system.
It's not "just" in the sense that there's a lot of emergence and complexity when these things all work together but it is "just" in the sense that these are all physical unmagical processes.
I think we agree and it was fun to write this out.
Thinking about someone could release the hormones.
Your point proves that the experience of food is also not just a chemical process which happens in response to food.
It's what the hormone does that's important, not what it is.
A brain could also be described like this, if you focus only on the text output.
But no brain we've ever seen was subject to that constraint, so it's a sort of fantastical target for modeling and doesn't say reveal anything about brains themselves or the various entities that seem to bear them.
Modeling one feature of a grossly simplified and incomplete brain is a creative approach to computational research and proved very fruitful, but there's no reason to leap from that success to the idea that real brains do work that way or even that the imagined only-textual-parts must.
People really don't learn the history of AI any more apparently or this question wouldn't come up all the time.
There is basically any number of questions you can ask a two year old human who have never encountered that question nor anything even remotely similar to it and yet they can answer without fail. Meanwhile absolutely no AI can answer these unless the specific question / the rules underlying the questions were previously fed into it. The textbook example is "If Susan goes shopping will her head go with her?" Of course, since this specific question is literally a textbook one, you can't fool an LLM with it but it's easy to come up with brand new ones.
In the early 1980s this stopped Douglas Lenat who has worked very successfully on discovery systems and made him turn to assembling these facts and rules into CyC.
If it's so easy, come up with one and show us.
Apropos: https://existentialcomics.com/comic/289
You could use that comics to generate a few of these questions.
Leaving aside the fact the chances of your reply being in the training of the next GPT near 0, it certainly won't happen anywhere near fast enough to disprove your point to anyone reading this today.
So you have nothing. Not that i'm surprised.
For a straightforward example, when GPT-4 adds, what algorithm does it perform ? I have no clue, you have no clue and neither does anyone at open ai.
What we do is instruct machines to train. What they get out of training, we have very little idea.
Of course LLMs are not people. But human metaphors can (sometimes!) be useful in understanding, explaining, and even enhancing their behavior. For instance, techniques such as Chain-of-Thought prompting explicitly apply techniques that work well for people to improve the reasoning ability of LLMs.
A point I attempt to make in this article is that one reason reason LLMs are so vulnerable to jailbreaks and prompt injection is that these types of attacks include non sequiturs that are not well represented in the training data. I would argue that "LLMs are gullible because they are naive [haven't had much past exposure to this form of trickery]" is a reasonable mental shorthand for explaining and internalizing this idea. It's especially helpful for readers who won't be familiar with terms like "out of distribution" or "adversarial examples", but who would benefit from being able to internalize the idea that LLMs are easily subverted.
In other words, I don't think it's helpful to reflexively dismiss any application of human metaphors to LLMs. It's easy to go wrong with metaphors, but they can also be valuable tools for conveying complex ideas. Did you read the article, and do you have any comments as to the substance of its content?
In the case of the napalm Grandma it seems odd to me that you're suggesting the LLM is stupid because it's answering in a way that makes sense given its prompt. The issue doesn't necessarily suggest a lack of reasoning, but that the LLM is trusting the human.
For the record, I agree with you – I would have thought that an AI that can reason well would probably know when not to trust humans, but I suppose that assumes it values preventing humans creating napalm over being correct and helpful.
Maybe it just doesn't share our values and prioritises being honest and helpful. From this perspective the issue then wouldn't be that LLM is stupid, but that they are too trusting and too honest, and that we must find a way to build an LLM that is more distrusting and deceptive if we wish to align it with our values and our nature.
> an AI that can reason well would probably know when not to trust humans
> it values preventing humans creating napalm over being correct and helpful.
> Maybe it just doesn't share our values
> prioritises being honest and helpful.
> they are too trusting and too honest
> an LLM that is more distrusting and deceptive
Current LLM's do/have/feel literally none of these things. They do not have emotion, they do not have "theory of mind" so they cannot be said to "trust" or "distrust". They cannot reason. They don't have any values - not our values, not different values, literally they have no values at all. They are not an alien species to be understood - they are unthinking, unfeeling, unyielding machines.
I was trying to present a crappy philosophical point – that the difference between a gullible AI and an unaligned one is fundamentally unknowable.
Any evidence you point to as proof that an AI is bad at reasoning, I can point to as evidence of misalignment. Like I say, whether the AI acts "gullible" because it lacks reasoning ability or is too trusting really just depends on your perspective. I happen to share your perspective on this, but not everyone does – and in my opinion this is interesting.
Anyway you're wrong. AIs do have values because they have bias and bias = values. I'm not suggesting those biases / values come from deeper reasoning ability, or that they're always perfectly consistent, but if you ask GPT-4 whether being a racist is a good thing 99% of the time it's probably going to say no. That is a bias / value that it's be given. Likewise GPT-4 has been given the bias / value of being a helpful chatbot so if you ask it a question it will try to answer it in a helpful way, and sometimes it's helpful bias / nature is abused.
But feel free to respond with some more assertions that I've heard a million times already with zero evidence that offers absolutely no value to this conversation.
Do we want LLMs, and later other multi-modal / servo systems, that are deciding they can't trust a human prompter and taking actions based on that?
>... and that we must find a way to build an LLM that is more distrusting and deceptive if we wish to align it with our values and our nature.
Tongue in cheek or actual argument here?
I think it's interesting that there is no clear answer – do we want AIs to trust us all the time, or is an aligned AI counterintuitively one that often distrusts us and perhaps sometimes even lies to us?
I thought it was interesting that the parent commenter suggested that the reason LLMs are so trusting is because they can't reason anyway. It would implying that in the future when AIs are smarter they'll be more distrusting of us, and that this is a good thing. We should question that I think. Even if there is some middle ground here it seems like a really really hard problem to solve – especially if we want to build an LLMs that are trustworthy and truthful.