Also, many people seem to be mistaking "artificial intelligence" for "actual intelligence".
You know that these LLMs are not actually intelligent, right? Right???
Also, many people seem to be mistaking "artificial intelligence" for "actual intelligence".
You know that these LLMs are not actually intelligent, right? Right???
They've had considerable success getting traction for these views with politicians in Europe, and that's the reason you see so much of this in the media currently. I'm hoping more sane voices will soon get organized, because significant parts of this are reminiscent of a well-funded doomsday cult.
I think this article is an early attempt at creating a politically viable counterpoint.
Maybe L2 systems. But those are just a driver assistance, the driver must supervise and intervene.
Actual L4/L5 systems haven't killed anyone as far as I know, except for the Uber accident, and even then the driver was supposed to be supervising it, I'm not sure if it counts as L4/L5
Got to hang out with some "hotshot" Facebook developers after hours. I remember Lee Byron (of React, GraphQL, etc) stating very confidently that the roads would be teeming with self driving cars by 2020.
Everyone wants to pose as the front(wo)man of that innovation, thought-leaders. Whatever.
It's all marketing, for their own careers, their own personal agenda.
The whole text looks like written by ChatGPT in a prompt like "Write about AI safety and relates it to the Age of Reason".
This is a good reason why AI should replace us, so we no longer have human beings writing that stuff to self-promote themselves.
The funniest part is in the end of the article:
"Eric Ries has been my close collaborator throughout the development of this article"
So the guy who wrote "The lean startup" is helping him with that article. drops mic
Can a submarine swim?
Does it matter if an LLM — which is far from the only kind of AI, and trivial to include in another more complex AI — does or doesn't match your definition of "intelligent" when it can write fluently in any language (human and code), and also getting near top-of-class results in the Bar exam and the biology olympiad?
Corporations — sometimes used as another example of an unaligned optimising agent — can't drive cars, they have to hire humans to do that for them, but corporations can still be prosecuted for dangerous chemicals (3M); and AI, just like other computer programs, can have dangerous output that you should not rely on (e.g. Thule early warning radar incident where someone forgot to tell it that it's OK for the moon to not respond to IFF pings and it nearly started WW3).
I think once you give AI working limbs and the ability to manipulate matter like humans, that's concerning, because if the AI becomes a runaway AI with no upper limits on the types of behavior allowed, we get into Paperclip Maximizer[0] territory very fast.
https://arxiv.org/abs/2303.11366
GPT-4's emotional intelligence seems to be very high by the looks of things
https://arxiv.org/abs/2304.11490
Empathy =/ Intelligence
Isn’t that a metric we use to determine the intelligence of animals and such?
Does GPT love its creators the way we love our parents? Does it mourn a loss the way we do?
Elephants are known to be of the more intelligent animals and apparently they mourn loss as well. So apparently there is something there linked with intelligence.
Maybe it's not directly intelligence related, but what makes a human have a spontaneous emotional response like laughter for instance?
The self-reflection link you sent isn't what I was thinking.
I was thinking about things like self reflection on my existence. Like am I happy doing what I'm doing? Do I feel I am making a positive difference in the world? If not what should I start doing to work toward becoming what I want to be?
The fact that I want to be something seems different. Nobody tells me what I want to be. But we still need to tell computers an awful lot about what they should be and what they should want to be. When will AI be able to decide for itself what it wants to be and what it wants to do? And to do all that sort of self reflection on it's own without us specifying a reward system that "makes" it want to do that? I don't strive to become a better in the world person for some reward.
> Isn’t that a metric we use to determine the intelligence of animals and such?
Depends what you mean by those words.
"Do they pass a mirror test" is, as far as I know, the best idea anyone's had so far for self awareness, and that's trivial to hard-code as a layer on top of whatever AI you have, so it's probably useless for this.
> Does GPT love its creators the way we love our parents? Does it mourn a loss the way we do?
I'd be very surprised if it did, given it wasn't meant to, but then again it wasn't "meant" to be able to create unit tests for a Swift JSON parser when instructed in German either, and yet it can.
Trouble is… how would you test for it? Do we have more than the vaguest idea how emotions work in our own heads to compare the AI against, or are we mysteries unto ourselves?
Or can we only guess the inner workings from the observable behaviours, like we do with elephants? But "observable behaviour" is something that a VHS can pass if it's playing back on a TV, and we shouldn't want to count that. Is an LLM more like a VHS tape or like an elephant? I don't know.
LLMs are certainly alien to how we think; as it's been explained to me, I think it should be thought of as a ferret that was made immortal, then made to spend a few tens of thousands of years experiencing random webpages where each token is represented as the direct stimulation of an olfactory nerve, and it is given rewards and punishments based on if it correctly imagines the next scent.
> I don't strive to become a better in the world person for some reward.
Feeling good about yourself is a reward, for most of us.
But I think this is a big difference between us and them: as I understand it, most AI have only one reward in training, while humans have many which vary over our lives — as an infant it may be pain va. food or smiles vs. loud noises, as a child it may be parental feedback, as an adolescent it may be peer pressure and sex, as a parent it may be the laughter vs. the cries of your child… but that's pop psychology on my part.
There’s an argument that the vast majority of human-generated research isn’t really “novel” but just derivative of other ideas. I’m not so sure ChatGPT couldn’t combine existing ideas to come up with something “novel” just like humans. I think there’s a case that it already comes up with creative, novel solutions in the drug space.
If an LLM can compose a poem found in no existing book, paint an attractive picture that no one has ever seen, or write a program that has never run before, it's intelligent enough to be called "intelligent."
Conversely, if you can do these things in the absence of any training or outside influence, you're intelligent enough to be called a god. Assuming that this is not the case, don't hold machines to standards you're not willing to hold humans to.
This has to be a candidate for the most unhelpful statement possible. If you are sincere about it, then why not apply it to every word in your reply and see where that leaves you?
Humans certainly have flaws when it comes to this. I've heard some discussions about the success and failures of AI in this regard. Can someone in this domain elaborate on the current state-of-the-art performance in this regard?
2. LLMs see positive transfer in multi lingual capability. For example, an LLM trained on 500B tokens of English and 10B tokens of French will not speak in french like a model trained on only 10B tokens of French. What will happen is that the model will be nearly as competent in French as it is in English https://arxiv.org/abs/2108.13349
3. Language models of code reason better even if the benchmarks have nothing to do with code.
(Apologies to any linguists. Please correct anything above if I'm off).
As for the other post, degraded performances are highly non trivial still. Some aren't actually poor, just worse.
Even the authors admit humans would see degraded performance on counterfactuals unless given "enough time to reason and revise", something they don't try to do with GPT-4.
Think about it. Do you genuinely belief you would score as accurately on a multiplication arithmetic test taken in base 8 ?
No, but I believe this is a different question. I think the more relevant question is whether a human can (even with the caveat of needing more time to reason about it). The larger question for a LLM is whether it can answer it at all and interpret why, without additional training data.
The paper seems to point that the ability of LLM to transfer is related to proximity to the default case. E.g., if default is base 10, is better at base 9 than base 2. I would interpret that as indicating more simple pattern recognition than deductive reasoning. The implication being that real transference is more dependent on the latter.
arithmetic but in a different base is one of the counterfactual examples in the paper. That's why i mentioned that. and yes it can answer them with worse performance.
You can juice arithmetic performance as is with algorithmic instructions. https://arxiv.org/abs/2211.09066. I see no reason the same for other bases wouldn't work.
Even if you gave a human substantial time for it (say a week of study), i believe he/she almost certainly reach the same accuracy unless he had access to specific instructions for working in base 8 he/she could call upon when taking the test.
I know, that's why I referenced the proximity of bases seemingly being important to the LLM. I think this is what differentiates it.
>and yes it can answer them with worse performance.
It's accuracy is dependent on proximity to it's training set (going back to my original point). I think that points to a different mechanism than humans and that's what my last post was focusing on.
I think we agree that humans would do less well in most other bases than base-10. But that side-steps the point I was making. Will humans do worse in base-3 than base-9? I doubt it, but according the the article, it's reasonable to assume the LLM would be progressively worse. That, IMO, is an indicator that something different is going on. I.e., humans are deriving principles to work from rather than just pattern recognition. Those principles can be modified on-the-fly to adjust to novel circumstances without needing additional training data. Humans are using reasoning in addition to pattern recognition.
This is probably a clunky example, but I'll try. Suppose an autonomous vehicle is trained to recognize that when a ball rolls into the street, it needs to slow down or stop because a child may not be far behind. A human can infer that seeing a kite blow into the street may signal the same response, even though they've never witnessed a kite blow into the street. The question is: can the autonomous vehicle infer the same? (This shouldn't be conflated with the general case of "see object obstructing the street and slow down/stop." The case I'm drawing here specifically adjusts the risk by the nature of the object being a child's toy. So, can the AV not only recognize the object as a kite but also adjust the risk accordingly?) I think one of the possible pitfalls is that we solve a more simple problem like image/pattern recognition and conflate it to a more difficult problem set being solved.
Circling back to the original point, one guess is that it's not understanding context as much as merely matching patterns really, really well. That can be incredibly useful but it may be something different than what's going on in our heads and maybe would should be careful not to conflate the two. Or, it's possible that all we're doing is also matching patterns in context, and eventually LLM will get there too.
I genuinely don't see how that would be a reasonable assumption.
>Will humans do worse in base-3 than base-9?
Why not? If you haven't learnt base 3 but you have base 9 you'll do poorer on it.
>That, IMO, is an indicator that something different is going on.
Whether something different is going on is about as relevant as the question of whether submarines swim or plans fly or cars run.
>I.e., humans are deriving principles to work from rather than just pattern recognition.
Not really. Nearly all your brain does with sense data is predict what it should be and adjust your perception to fit. You can mold these predictions implicitly with your experiences but you're not deriving anything from first principles.
>This is probably a clunky example, but I'll try. Suppose an autonomous vehicle is trained to recognize that when a ball rolls into the street, it needs to slow down or stop because a child may not be far behind. A human can infer that seeing a kite blow into the street may signal the same response, even though they've never witnessed a kite blow into the street. The question is: can the autonomous vehicle infer the same? (This shouldn't be conflated with the general case of "see object obstructing the street and slow down/stop." The case I'm drawing here specifically adjusts the risk by the nature of the object being a child's toy. So, can the AV not only recognize the object as a kite but also adjust the risk accordingly?) I think one of the possible pitfalls is that we solve a more simple problem like image/pattern recognition and conflate it to a more difficult problem set being solved.
Casual reasoning ? all evidence points to LLMs being more than capable of that https://arxiv.org/abs/2305.00050
It's not an assumption. It's literally based on the results of your own reference:
>LM performance generally decreases monotonically with the distance
If you can't be bothered to read your own reference, I don't think additional conversation is worthwhile because it becomes apparent that it's more dogmatic than reasoned.
Your newest link is not really supportive of your "all evidence" claim. It goes into further detail about how LLM can have high accuracy while also making simple, unpredictable mistakes. That's not good evidence of a robust causal model that can extrapolate knowledge to other contexts. If I didn't know better, I'd assume you could just as well be a chat bot who only reads abstracts and replies in an overconfident manner.
Human performance generally decreases with level of exposure so I figured you were talking about something else. Guess not.
>I don't think additional conversation is worthwhile because it becomes apparent that it's more dogmatic than reasoned
By all means, end the conversation whenever you wish.
>It goes into further detail about how LLM can have high accuracy while also making simple, unpredictable mistakes.
I'm well aware. So? Weird failure modes are expected. Humans make simple, unpredictable mistakes that don't make any sense without the lens of evolutionary biology. LLMs will have odd failure modes regardless of whether it's the "real deak" or not, either adopted from the data or from the training scheme itself.
>If I didn't know better, I'd assume you could just as well be a chat bot who only reads abstracts and replies in an overconfident manner.
Now you're getting it. Think on that.
Are you saying that as humans get more experience, they perform worse? I disagree, but irrespective of that point it’s wild that you can have this many responses while still completely bypassing the entire point I was making.
I don’t think most would argue that performance increases with experience. The point is how well can the performance be maintained when there is little or no exposure. Because that implies principled reasoning rather than simple pattern mapping. That is the entire through line behind my comments regarding context dependent language, novel driving scenarios, etc.
>Think on that
In the context of the above, I don’t think this is nearly as strong of a point as you seem to think it is. There nothing novel about a text-based discussion.
If that child had the basic teaching most children do (little exposure) then a quiz will result in much worse performance than a base 10 equivalent test. This is very simple. I don't know what else to tell you here.
2. You must understand that a human driver that stops because a kite suddenly comes across the road doesn't do so because of any kite>child>must not hurt reasoning. Your brain doesn't even process information that quickly. The human driver stops (or perhaps he/she doesn't) and then rationalizes a reason for the decision after the fact. Humans are very good at doing this sort of thing. Except that this rationalization might not have anything at all to do with what you believe to be "truth". Just because you think or believe it is so doesn't actually mean it is so. Choices shape preferences just as much as the inverse. For all anyone knows and indeed most likely, "child" didn't even enter the equation untill well after the fact.
Now if you're asking whether LLMs a matter of principle can infere/grok these sort of casual relationships between different "objects" then yes as far as anyone is able to test.
Regardless, it still misses the point. I’ve never been explicitly exposed to base-72, yet I can reason my way through it. I would argue my performance wouldn’t be any different than base-82. So I can transfer basic principles. What the LLM result you referenced shows is that it is not learning basic principles. It sure seems like you just read the abstract and ran with it.
As far as the psychology of decision making, again, I think you're speaking with greater confidence than is warranted. In time critical examples, I’m inclined to agree. And there’s certainly some notable psychologists who would expand it beyond snap judgments. But there are also some notable psychologists who tend to disagree. It’s not a settled science, despite your confidence. But again, that’s getting stuck in the limitations of the example and missing the forest for the trees. The point is not in whether decisions are made consciously or subconsciously, but rather how learning can be inferred from previous experience and transferred to novel experiences. Whether this happens consciously or not is besides the point. And you are further going down what I was explicitly taking against: confusing image/pattern recognition for contextual reason. You can see this in the relatively recent Go issue; any human could see what the issue was because they understand the contextual reasoning of the game but the AI could not and was fooled by a novel strategy. The points I’ve been making have completely flew over your head to the point where you’re shoehorning in a completely different conversation.
I guess so. I've never meant to imply greater exposure leads to worse outcomes.
>I would argue my performance wouldn’t be any different than base-82.
Even if that were true and i don't know that i agree, the authors of that paper make no attempt to test in circumstances that might make this true for LLMs as it might for people. So the paper is not evidence of the claim (no basic principles) either way. For example, i reckon your performance on the proceeding 82 test will be better if taken a immediately after than if taken weeks or months later. So surrounding context is important even if you're right.
>What the LLM result you referenced shows is that it is not learning basic principles.
I disagree here and i've explained why.
>You can see this in the relatively recent Go issue; any human could see what the issue was because they understand the contextual reasoning of the game but the AI could not and was fooled by a novel strategy.
You're talking about this ? https://www.zmescience.com/future/a-human-just-defeated-an-a...
KataGo taught itself to play go by explicitly deprioritizing “losing” strategies. This means it didn’t play many amateur strategies because they were lost early in the training. This is hard for a human to understand because humans all generally share a learning curve going from beginning to amateur to expert. So all humans have more experience with “losing” techniques. Basically what I’m saying is, it might be that the training scheme of this AI explicitly prioritized having little understanding of these specific tactics, which is different than not having any understanding.
This circles back to the point I made earlier. Having failure modes humans don't or won't understand or have is not the same as a lack of "true understanding".
We have no clue what "basic principles" actually are on the low level. The less inductive bias we try to shoehorn into models, the better performing they become. Models literally tend to perform worse the more we try to bake "basic principles" in. So presence of an odd failure mode we *think* belies a lack of "basic principles" is not necessarily evidence of a lack of it.
>The points I’ve been making have completely flew over your head to the point where you’re shoehorning in a completely different conversation.
You're convinced it's just "very good pattern matching", whatever that means. I disagree.
E.g., racism/sexism/...most -'isms' appear to be a general heuristics that help us make quick judgements. But we can also our decision-making process by reverting to basic principles, like the idea that humans have equal moral worth regardless of skin tone or gender. AI can even mimic these mitigations, but you haven't convinced me that it can fundamentally change away from it's training set based on an understanding of basic principles.
As for the Go example, a novice would be able to identify that somebody is drawing a circle around it's pieces; your link even states this. But you recharacterizing this as a specific strategy is weird when that strategy causes you to lose the game. It misses the entire meaning of strategy. We see the limitations of AI in it's reliance to training data from autonomous vehicles to healthcare. They range from the serious (cancer detection) to the humorous (Marines overtaking robots by hiding in boxes like in Metal Gear). The paper you referenced similarly shows it is reliant on proximity to the training set, rather than actually understanding the underlying principles.
Humans don’t have a grasp of the “principles of reasoning” and as such are incapable of distinguishing “true”, "different" or “heuristic” assuming such a distinction is even meaningful. Where you are convinced of “faulty shortcut”, I simply think “different”. Multiple ways to skin a cat. a plane's flight is as "true" as any bird. There's no "faulty shortcut" even when it fails in ways a bird will not.
You say humans are "true" and LLMs are not but you base it on factors that can be probed in humans as well so to me, your argument simply falls apart. This is where our divide stems from.
>I don't think you've provided evidence that AI can have a principled understanding while we can show that humans can.
What would be evidence to you? Let’s leave conjecture and assumptions. What evaluation exist that demonstrate this “principled understanding” in humans? and how would we create an equitable test in LLMs?
>a novice would be able to identify that somebody is drawing a circle around it's pieces; your link even states this. But you recharacterizing this as a specific strategy is weird when that strategy causes you to lose the game.
You misunderstand. I did not characterize this as a specific “strategy”. Not only do modern Go systems not learn like humans, but they also don’t learn from human data at all. KataGo didn’t create a heuristic to play like a human because it didn’t even see humans play.
>The paper you referenced similarly shows it is reliant on proximity to the training set, rather than actually understanding the underlying principles.
Even the authors make it clear this isn’t necessarily the bridge to take so it’s odd to see you die on this hill.
The counterfactual of syntax is Finding the main subject and verb of something like “Think are the best LMs they.” in verb-obj-subj order (they, think) instead of “They think LMs are the best.” in subj-verb-obj order (they, think). LLMs are not being trained on text like the former to any significant degree if at all yet the performance is fairly close. So what, it doesn’t “underlying principles of syntax” but still manages that ?
The problem is that you take a fairly reasonable conclusion from these experiments. I.e LLMs can/often also rely on narrow, non-transferable procedures for task-solving and proceed to jump the shark from there.
>but you haven't convinced me that it can fundamentally change away from it's training set based on an understanding of basic principles.
We see language models create novel functioning protein structure after training, no folding necessary.
https://www.researchgate.net/publication/367453911_Large_lan... So does it still not understand the “basic principles of protein structures”?
"we do not expect our language model to generate proteins that belong to a completely different distribution or domain"
So, no, I do not think it displays a fundamental understanding.
>What would be evidence to you?
We've already discussed this ad nauseum. Like all science, there is no definitive answer. However, when the data shows evidence that something like proximity to training data is predictive of performance, it's seems more like evidence of learning heuristics and not underlying principles.
Now, I'm open to the idea that humans just have a deeper level of heuristics rather than principled understanding. If that's the case, it's just a difference of degree rather than type. But I don't think that's a fruitful discussion because it may not be testable/provable so I would classify it as philosophy more than anything else and certainly not worthy of the confidence that you're speaking with.
Good thing they don't make sweeping declarations or say anything about that meaning narrow learning without transfer. Jumping the shark yet again.
https://www.pnas.org/doi/full/10.1073/pnas.2016239118
>We find that without prior knowledge, information emerges in the learned representations on fundamental properties of proteins such as secondary structure, contacts, and biological activity. We show the learned representations are useful across benchmarks for remote homology detection, prediction of secondary structure, long-range residue–residue contacts, and mutational effect.
From the sequences of just the proteins alone, Language Models learn underlying properties that transfer to a wide variety of use cases. So yes, they understand proteins in any definition that has any meaning.
>Good thing they don't make sweeping declarations or say anything about that meaning narrow learning without transfer.
That's exactly what that previous quote means. Did you read the methodology? They train on a universal training set and then have to tune it using a closely related training set for it to work. In other words, the first step is not good enough to be transferrable and needs to be fine tuned. In that context, the quote implies the fine tuning pushes the model away from a generalizable one into a narrow model that no longer works outside that specific application. Apropos to this entire discussion, it means it doesn't perform well in novel domains. If it could truly "understand proteins in any definition", it wouldn't need to be retrained for each application. The word you used ('any') literally means "without specification"; the model needs to be specifically tuned to the protein family of interest.
You are quoting an entirely different publication in your response. You should use the paper from which I quoted to refute my statement, otherwise this is the definition of cherry picking. Can you explain why the two studies came to different conclusions? It sure seems like you're not reading the work to learn and instead just grasping at straws to be "right." I have zero interest in having a conversation where someone just jumps from one abstract to another just to argue rather than adding anything of substance.
This is why it's mostly meaningless to for a LLM to pass the bar, but not meaningless for a human to do so. We (rightly, for the most part) assume that a human who passes the bar can transfer those skills into unique and novel situations. We can't make that assumption for LLMs, because they are lacking adaptability that is needed for true intelligence.
If you took arithmetic tests in base 8, you wouldn't reach the same accuracy either.
The word was made up to cover a range of cognitive abilities that humans and animals (to varying degrees) possess. And we're gradually figuring out how to design machines with similar abilities. The general idea of intelligence being you can figure out how to do things that aren't just instinctual. a generalized intelligence can do this across any number of domains without some fixed limit.
So you're saying humans, collectively, are gods? After all, as a species we started with nothing - no "clean" training data, no paintings to replicate, not even language itself. And here we are - arguing about whether one creation of ours, trained on a bunch of other stuff we created, is as intelligent as us.
So is all human knowledge. They're all things we "made up" to describe the world around us. That's a really cheap way to reduce something. There's a mountain of psychological literature that intends to describe and measure intelligence. Saying it doesn't mean much of anything is a bit of a postmodern hot take.
> If an LLM can compose a poem found in no existing book, paint an attractive picture that no one has ever seen, or write a program that has never run before, it's intelligent enough to be called "intelligent."
Painting pictures and composing poems aren't good ways to measure general intelligence. That's creativity. They're correlated but not the same.
There are specific tests for measuring general intelligence, and there's absolutely no way current gen LLM would score highly on them because they don't even accept visual input. Yes I've seen that they can score highly on verbal-only tests. That's obvious anyway because they have seen, written down, and have access to the answers during the test.
Even if they did accept visual input they would only be able to solve problems for which they've previously been given the answers (trained on the dataset). Even then I'd have my doubts given how terribly wrong ChatGPT4 is when I ask it to produce a very simple Kubernetes manifest which has a strictly defined spec.
One problem is that ChatGPT4 has been degrading over time -- and no, I don't GAF about any opinions or assurances to the contrary. So we're seeing more and more people look at it for the first time and ask what the big deal is. Maybe the Code Interpreter feature will reverse that trend, we'll see.
Those of us who were around (read: paying $20/month) when they first enabled GPT4 support are left waving our hands fecklessly, muttering "Yeah, but you should've seen it back in the old days, back in March of '23. Now you kids get off my lawn."
An LLM can write a poem that fools me into thinking it’s good, but not anyone who reads poetry as a hobby. Poetry is inspired in a way that’s hard to replicate with a prompt.
Artists also take inspiration from other artists, but the “distance” between their influences and their work is so much greater than the distance between an LLM’s that I don’t fault anyone who think it’ll take time for LLMs to catch up.
And you don't grasp the fundamentally-incremental nature of this point?
What about version 5? 6? 10?
quality is required.