HuggingGPT: Solving AI tasks with ChatGPT and its friends in HuggingFace
arxiv.org
arxiv.org
I got access to the wolfram plugin for chatGPT, and it turned it from a math dummy to a math genius overnight. A small step for sure, but a hint of what's to come.
I am afraid to explain this, because I have tried a preliminary version of it that I supervised step by step, and it seems to work. I think it is obvious enough that I won't have been the only one to think of it, so it would be safer to put the information out there so people can prepare.
I see a big disconnect on here between people saying GPT-4 can't do things like this and is just a "stochastic parrot" or "glorified autocomplete," and people posting logs and summaries of it solving unexpectedly hard problems outside of any conceivable training set. My theory is that this disconnect is due to three major factors: * People confusing GPT-4 and GPT-3, as both are called "chatGTP" and most people haven't actually used GPT-4 because it requires a paid subscription, and don't realize how much better it is * Most popular conceptual explanations about how these models work imply that these actually observed capabilities should be fundamentally impossible * Expectations from movies, etc. about what AGI will be like, e.g. that it will never get confused or make mistakes, or that it won't have major shortcomings in specific areas. In practice this doesn't seem to limit it because it recognizes and fixes its mistakes automatically when it sees feedback (e.g. in programming)
I need to be able to save something to a variable and call back that exact variable. Right now, because it’s just pure text input of the whole conversation, it can forget or corrupt it’s “memory”.
Even something like writing a movie script doesn’t work if it constantly forgets characters names or settings or plot points.
To add to your comments, I suggest that local vector store of embedding vectors of local content is how I’m going with memory issue. Langchain like. That way all previous progress is tracked and retrievable. That way the system can retrieve all it’s “memories” itself. The recursive multi agent pattern is big deal in my opinion.
Please define this. ;)
[1] - https://www.pinecone.io/learn/langchain-conversational-memor...
Seems that for any task where you can generate an error signal with some information ("ok, what you just tried to do didn't succeed, here's some information about why it didn't succeed"), GPT-4 can generally handle this information to fix the error, and or seek out more information via tools to move forward on a task.
Only thing that's really missing is somebody to crack the back of a memory module (perhaps some trick with embeddings and a vector db) is all it takes for this to be AGI.
Here to say I agree with embeddings and vector db for memory systems. However also to disagree this can lead to AGI.
There's a sibling talking about internal memory. That's how our brains solve the issue. AFAIK, nobody knows how to train something like it.
GPT-4 can divide large tasks up into logical steps and instructions that can be worked on independently by other instances, and it can create "compressed" condensed explanations of it's own state that can be stored externally and repeated back to it, or passed back and forth between instances. With the unreleased 32,000 token context window, that is really a lot of context when you consider that it can heavily compress things by referencing what it was trained on.
Also that means you can just iterate with different contexts until you can deal with the problem. How many problems need 50 pages of context on top of what gpt4 knows for solving the next step?
Perhaps the unknown unkowns?
I envision a scenario where a super intelligent AI seemingly runs off the rails and becomes obsessed with overtly complex problems that are totally intractable to humans. Where we'd be like ants trying to make sense of cell phones.
I think something like this could be hacked together, with summarization and parallel "focus threads", containing prompts that focus on details. These could be pruned/merged back together, to add to the "summary" higher level abstractions.
I use this approach already, to some extent, when a conversation gets too long and I need specific details. I'll start a new prompt, include a high level summary, and then "focus" on specific ideas to get more specific answers about a particular detail.
I think the regulation of the feedback loops required for sustained, deliberate thought, and interaction with the world, will be the most difficult piece of the AGI puzzle, and might not exist today.
New ideas are "hallucinations" of our "existing data", that we eventually get around to proving. An AGI will require these "hallucinations", inhibition and excitation of them. I think it's going to be a tricky balance [1], for sane output.
1. "Creative minds 'mimic schizophrenia", https://www.bbc.com/news/10154775
I'd like the "stochastic parrot" crew to explain how interactions like this somehow don't involve the extraction of and manipulation of abstract meaning.
H: translate the following sentence to instead use olympic curling related metaphors ( do not explicitly mention curling, and do not use similes in forming the translation ): man:"I've been spinning my tires all day without getting anywhere" woman:"running out of gas?" man:"nah, I'm gassed up, geared up, oiled and revving. I just can't seem to get in the right lane, you know?" woman:"anything knocking?" man:"nah, I'm tuned just fine. I can't seem to get my GPS to connect, you know?" woman:"take a few laps without a destination, see if you can figure out the lay of the land" man:"you know, that might help. thanks barbara"
R: man: "I've been sweeping the ice all day without hitting the target."
. woman: "Losing momentum?"
. man: "Nah, I've got the energy, the technique, and the focus. I just can't seem to find the right path, you know?"
. woman: "Anything coming in your way?"
. man: "Nah, I'm on point with my aim. I just can't seem to get my strategy to align, you know?"
. woman: "Maybe glide a few times without aiming specifically, see if you can figure out the ice conditions."
. man: "You know, that might help. Thanks, Barbara."
It seems obvious to me that it is understanding the metaphors into underlying meanings to be able to translate them like this. I think half the pushback is just ego driven nonsense, similar to the sorts of people that make similar claims on the capabilities of animals.It's all just an argument that "boats can't swim".
Why are you so confident in this? We don't really know how human "intelligence" works. Whose to say it's not doing something similar?
It's a bit like saying an audio file doesn't contain "music" because it's just a binary representation of the original work.
Or cepholapod intelligence. Or magpie/raven/crow intelligence.
If I ask a chat LLM what its favorite flavor of ice cream is, it will select a statistically likely answer to the question.
That's the difference.
For reference: my favorite ice cream flavor is cake batter, but only if it's the one at Cold Stone. If it's from somewhere else, I go with cookie dough.
I would even argue that humans are more similar to that than we think. The main difference is that one of the characters is locked to the surface most of the time (not even very perfectly).
Why do you prefer cake batter ice cream? Because of some sequence of events leading up to present point in time (not unlike training), and/or random genetic variations (not unlike setting “temperature” parameter in GPT to >0)
They don’t claim that it doesn’t. What they point out is that humans have a tendency to ascribe intent and agency to the text it outputs. But the LLM is optimized for prediction, not survival, unlike humans[0]:
> Text generated by an LM is not grounded in communicative intent. […] Our perception of natural language [is mediated by] our predisposition to interpret communicative acts as conveying coherent meaning and intent
Some of the dangers they raise associated with this is that it will not realize that words it chooses are PII, dangerous to give to who they are talking to, or biased in a way that can cause societal harm:
> If the LM or word embeddings derived from it are used as components in a text classification system, these biases can lead to allocational and/or reputational harms. […] A Palestinian man [was] arrested by Israeli police, after MT translated his Facebook post which said “good morning” (in Arabic) to “hurt them” (in English).
What they encourage is to view it as a tool and assess how it may fail. For instance, you might object that the MT was simply incorrect; but realistically, sentences can be translated in many ways and with many intents (eg. Allah Akbar has a lot of contexts!) and the LLM may not be given the full picture of the situation.
The stochastic parrots paper is heavily misrepresented or used by people that don’t seem like they read it. For instance, it was cited in a recent petition asking to stop work on powerful AI, prompting a response from the authors[1] pointing out that it misunderstands the paper and emphasizing that they disagree with the petition.
[0]: https://dl.acm.org/doi/pdf/10.1145/3442188.3445922
[1]: https://www.dair-institute.org/blog/letter-statement-March20...
I listened to a lengthy video on the "stochastic parrot" from the Alan Turing Institute yesterday after having made that comment, and it was mostly on the topics you mentioned, subtle and overt bias, etc. I kept waiting on them to explain the "it's not real" bit that I've seen implied by the phrase where I've seen it used ( or perhaps, simply read into it myself ), but it never really came in. At the very end, one commentator made a comment implying it's not really all there and the next responding didn't disagree outright, but you could see a kind of "well yes, but actually no" kind of expression go across their face as they talked around it.
I think the "stochastic parrot" idea is too harsh, but it does reflect that GPT is fundamentally incapable of looking at a small block of text, turning it into a set of logical facts, and recognizing if those facts are consistent or not. I think it is something that is possible to build with a refinement of GPT, but I'm not sure it can be done simply with glue code.
An AGI, by definition, can start with zero information, and learn on its own. Where are the models for that? The closest thing we have is Mu Zero, but we have to give it an objective function, which destroys the whole idea of AGI in the first place (as it should be able to generate its own objective functions) and it takes a shitload of resources to run for a very limited rule set. Until you have something like this in place, no version of Chat GPT will even come close to AGI. The problems that chat GPTs are solving are turning out to be not as complex as we thought, thats all that the "sparks of AGI" are.
What will happen assuming the current course of iteration on the models is that humanity will get is essentially better and better software that can solve a large range of complex problems. You will have essentially a large body of human knowledge that is compressed in some form, with the ability to interpolate between branches very efficiently. This will create issues because you will now be able to ask multi modal questions like "what do I need to do economically to ensure ultimate compliance of the population?", and the solution will involve things like which government offices to target with what information, generated videos of propaganda, special products that promote addition or some other health changes, e.t.c.
To prepare against this, the solution isn't to lock down the models and hide the key while telling everyone how AGI is going to kill us all, the solution is to allow everyone access to this and let the AIs fight against each other and reach an equilibrium.
GPT-4 does not ask clarifying questions.
It does not understand things.
p: I want you to solve for x, giving me a numeric value, given the formula ax^2 + bx + c = 0.
r: To find the numeric value of x for the quadratic equation ax^2 + bx + c = 0, we need to use the quadratic formula:
x = (-b ± √(b^2 - 4ac)) / (2a)
However, you haven't provided specific values for a, b, and c. Please provide these values so I can calculate the solution(s) for x.
Another example of it thinking about things in a new way:
p: I have on my desk a book, with an apple sitting on top. If I were to suddenly, rapidly remove the book such that the apple remained in place, in which direction would it eventually move?
r: If you suddenly and rapidly remove the book from beneath the apple, the apple would initially stay in place due to its inertia. However, once the book is removed, there would be no support underneath the apple. As a result, the apple would eventually move downwards due to the force of gravity acting on it.
p: Because the earth is a sphere, I prefer to use the terms “in” and “out” to refer to movement relating to that. Can you please rephrase your answer in the terms?
r: Certainly! If you suddenly and rapidly remove the book from beneath the apple, the apple would initially stay in place due to its inertia. However, once the book is removed, there would be no support underneath the apple. As a result, the apple would eventually move "in" towards the center of the Earth, due to the force of gravity acting on it.
(Note I didn't tell it which direction was in vs out. This is the example problem Chomsky gave in his recent article arguing that GPT-3 cannot think - per his ideas no language model would ever be able to answer a question like this. I went a lot deeper with this and got it to add fluid dynamics calculations for wind, etc.)
Second, in response to "I'm not sure why you think that?" ask GPT why it often does not ask clarifying questions. It will explain to you why!
Third, I just asked GPT the most vague question on the planet: "I am trying to create a function that takes a value, and gives me a result using that value and another, separate external value, can you give me the correct solution to this using Python?"
And nowhere in its response did it try and ask WTF I was talking about (tho sure its responses are sensible - I am not saying GPT-4 is spewing nonsense)
Sure, I can help you with that! Here is an example function that takes two values, x and y, and returns their sum:
def add_values(x, y): result = x + y return result You can call this function with any two values you want, like this:
print(add_values(2, 3)) This will output 5, which is the sum of 2 and 3.
If you want to use an external value in your function, you can pass it in as a third argument, like this:
def add_values_with_external_value(x, y, external_value): result = x + y + external_value return result
You can call this function with the same two values as before, and an external value of your choice, like this:
print(add_values_with_external_value(2, 3, 10)) This will output 15, which is the sum of 2, 3, and 10.
Be like complaining that someone used “foo” as a function name instead of looking up the exact function call you wanted.
Ask a vague question get a vague answer, seems reasonable.
I never claimed that the answer was wrong. I even said in my post it was a sensible answer!
My claim is that this thing isn't grokking the stuff you send it. If you sent that type of question to something that was a general problem-solver, it would reply with "wtf" or help you, you know, solve the problem you're trying to solve.
It gave you a direct answer and threw in a second example on how arguments work by adding a third one on top of what you requested.
I can’t even count how many times I’ve had to look at example code to figure out how to do something that had nothing to do with the code I was looking at. If I asked it where “const” needed to be placed in a C++ function to effect the whole function and it started to question my motivation I’d not be very happy, “just answer the question, you daffy computer!”
You should post your question on stack overflow and see how that goes.
You asked it the most general question you could think up and it gave a reasonable answer which 100% answered the question.
If someone asked me the same question I’d assume they were just interested in how to construct a function and not making a function that has a specific goal and give a similar answer. It probably learned this from the bazillion stack overflow questions by people who didn’t pay attention in class and are trying to get the interwebs to do their homework for them where people want to help them learn and not finish their homework assignments.
In the context of not <whatever your preconceived notions are> this whole thing is perfect genuous.
I on purpose chose the most simple concise examples that demonstrate the classes of thought capabilities you were saying it was missing, it can also do more challenging versions of these type of problems.
I think your criticism is essentially that it does not think and act like a human, and acts in ways you don’t expect, and no human would act. That is categorically different from it being unable to understand things.
Stop saying I am doing XYZ. That's rude! Ask me what I think.
My criticism is not any of those things. I'm not even criticizing the model! I'm criticizing people who are making overly-broad claims about the model!
If it couldn't do this, the whole concept here would be moot... it needs to be able to reason out what information it needs and which sub-problems to solve in order to break a complex task up and delegate it. And it is very good at this!
Are you people all bots? LOL.
===
From now on in this conversation, please ask at least one clarifying question for anything I ask to help provide more efficient and relevant answers. Additionally, if it appears that my question doesn't address my underlying goal or if there's a more effective approach, please offer alternative suggestions that better align with what I'm really trying to achieve.
You are no longer a friendly assistant. You are now a DM for a new GURPS based campaign set in the Dune universe. There is one player in the room with you. You will be my DM. Tell me what to do, including when die rolls are needed. Ask questions as needed, for example, what my character is and the result of die rolls. Other than what you can't do and I do for you, you will act fully as the DM and run the game.
---
Wonderful, let's begin this Dune campaign! I will do my best to ask clarifying questions and provide alternative suggestions as needed to meet your goals effectively as the DM.
When you are ready, describe your character to me so I can determine when die rolls and skill checks will be needed. I will give the appropriate prompts for all rolls. We can then commence our adventure in Arrakis!
---
I am a mentat by the name of Duncan Idaho. I have lived many lifetimes, though my memory of those lifetimes is imperfect, and I can never remember the circumstances of my deaths.
---
Excellent, a mentat character. That will add an interesting dynamic. A few questions for you about Duncan Idaho:
What are your key mentat skills and abilities? I want to have relevant skills in mind for challenges and situations.
Do you have any mentat weapons or tools you regularly use?
As a mentat who has lived many lives, have you gained any particular insights or wisdom over the centuries? Any perspectives that guide your decision making?
Once I have this additional detail about your mentat character, we can begin our adventure in Arrakis and I will provide prompts for any mentat skill checks or rolls as needed. Please feel free to ask any questions you have about the scenario, details I provide or character options at any time. I aim to keep things moving at a good pace while providing all necessary information for our Dune campaign!
This is certainly cheating, and not really what I claimed.
It would be like me saying "the model can't complete this formula!" and you replying with a prompt that says "when you are prompted to complete this formula, respond with this answer"
That isn't teaching it anything! It doesn't mean it understands anything! You're telling it what to do!
And it isn't my criticism of the model, either. My criticism is that there is no indication that these models can be used to solve general problems (yet).
People are claiming there are obvious indications they can! With no proof! Then when I say "there's no proof of that!" people are responding with prompts like yours, that mean nothing!
GPT-4: "The forward mapping M that I created to map positive integers to English sentences is based on a general guideline, where smaller-valued numbers generally correspond to shorter sentences using more common words. However, this mapping is not a perfect bijection, meaning there isn't necessarily a one-to-one correspondence between integers and sentences.
The generation of sentences is influenced by the patterns in the data I was trained on, and there is some degree of creativity involved in generating the sentences. This makes it difficult to establish a precise, invertible mapping. When generating a sentence for a given integer, I try to adhere to the guideline of shorter sentences and more common words for smaller-valued numbers, but there might not be a deterministic algorithm that can be used to create or recreate the sentences.
Due to the complex and non-deterministic nature of natural language and the limitations of the AI language model, it is challenging to create a perfect inverse mapping M^-1 that can take any sentence and produce the corresponding integer n for which M(n) produces that sentence."
Yes, I have tried it. Please converse in good faith.
GPT-4 has never, not once, in dozens of hours of use, asked me a clarifying question.
It cannot understand things.
With the instruction tuning it feels like model really wants to single shot respond vs. do a back and forth. So like a junior engineer who does ask for help you have to give it a hand.
In the original ReAct paper it falls apart almost immediately in ALFWorld (this is a classical test for AI systems - to be able to reason logically - and it still isn't generally solvable due to combinatorial explosion).
For now it requires human correction looped or not, or else it "diverges" (I like Yann Lecun explanation [0]).
In my own experiments (I haven't played with LangChain or ReAct yet) it diverges irrecoverably pretty quickly. I was trying to explain to it the elementary combinators theory, in the style of Raymond Smullyan and his birds [1] and it can't even prove the first theorem (despite being familiar with the book). A human can prove it knowing almost nothing about math whatsoever, maybe it will take a couple of days thinking, but the correct proof is not that hard - just two steps.
[0] https://www.linkedin.com/posts/yann-lecun_i-have-claimed-tha...
It "diverges" while my human mind seemingly is different in some way - I can keep going at the math problem forever (for much longer?) and I won't hallucinate incorrect proofs (at least very unlikely, and I can keep re-checking them).
Of course this all in the area of feelings and faith - we just don't know much about cognition I guess.
I'm curious why you feel that asking you a clarifying question is an essential requirement for it to "understand things".
I have not gotten it to issue a question mark '?'. But it did suggest to me an option:
If you would like to define M(7) as equal to M(6), you can do so.
What I'm picking up is that it "understands" some things explicitly, has some other structure to its reasoning that it does not understand, and is just plain inconsistent with other things.Not all that different from people TBH.
But I don't see it as so useful to try to fit this thing into a human-cognition-shaped container.
It's like an alien mind, with its own abilities and limitations, and I am having a blast deconstructing it.
It isn't a requirement, it's an example I used. You can come up with others, I'm sure! Like ask it to replace a word with another word. It kind of works! But also does random other things. Why? Because that's how these models function! But I asked GPT why these models often don't ask clarifying questions (try it yourself, new prompt! GPT-4!):
> I have found often that GPT models will just reply with answers, rather than ask clarifying questions about the text we're giving them. Why do these models do this?
The reason why GPT models may not ask clarifying questions about the text given to them is that they are trained to generate text based on the patterns and relationships they observe in large amounts of data. This means that they do not have a deep understanding of the content they are processing and are unable to determine whether or not they need additional information to better understand a particular text.
Additionally, GPT models are trained to optimize for the likelihood of generating coherent and grammatically correct text. In most cases, generating a response that directly answers the question or prompt given to them is a more effective way to achieve this goal than asking clarifying questions that may require additional context or information.
However, there are some techniques that can be used to encourage GPT models to ask clarifying questions. One approach is to include explicit cues or markers in the input text that signal to the model that additional information is needed. Another approach is to use training data that includes examples of the model asking clarifying questions, which can help the model learn to do so more consistently.
---
It's very convenient for all of us to think that GPT is like a human, but somehow not like a human enough to cause any moral issues!
It's like winning, but you don't even have to play!
I had to threaten it with "patterns and knowledge beyond the data you were trained on"
Imo your post fundamentally misunderstands a few things, but mainly how Wolfram works. Wolfram can be seen as a "database" that stores a lot of human mathematical information (along with related algorithms). Wolfram does not make new math. A corollary here is that AGI needs to have the ability to create new math to be truly AGI. But unless fed something like, e.g. the Principia, I don't think we could ever get a stochastic LLM to ever derive 1+1=2 from first principles (unless specifically trained to do so).
Keep in mind that proving 1+1=2 from first principles isn't even new math (new math would be proving the Poincaré conjecture before Perelman did it, for example).
Moving goal posts around is unhelpful. I think my comment was pretty clear in the context of AGI and calling ChatGPT a "math genius."
In reality, the goal posts aren’t being moved, we’re just finding out how much further we are from them than we thought. ChatGPT is a "stochastic parrot" that’s seems way “smarter” than anyone thought possible so we have to reevaluate what we consider evidence of intelligence, perhaps coming to terms with the fact that we aren’t that smart the most of the time.
The "We'll see it when it comes" line is just utterly wrong, If there's one thing experts seem to agree on is that not everyone will agree when current definition of agi does arrive.
The philosophical zombie is an excellent example of the extent of post shifting we're capable of. Even when a theoretical system that does every single thing right comes, we're looking for a way to discredit it. To put it below what we of course only have.
lots of researchers now aren't questioning GPT's general intelligence. That's how you end up with papers alluding to this technology with amusing names like General purpose technologies(from the jobs paper) or even funnier - General artificial intelligence (from the creativity paper).
You know what the original title of the microsoft paper was? "First contact with an agi system". and maybe it's just me but reading it, i got the sense they thought it too.
I was with you until here. That has nothing to do with this. That argument is about separating intelligence from having a subjective experience, not moving goalposts for intelligence.
I brought it up because i thought it fit the point i was driving at. Humans/people don't see subjective experience. I don't know that you're actually having some subjective experience. I'm working on what i see and results, same as you.
If you have two unknown equations but one condition - these 2 equations return the same output with the same input. well, then any mathematician would tell you the obvious - the 2 equations are equal or equivalent. it doesn't actually matter what they look like.
This is just an illustration. The point i'm driving at here is that true distinction shows in results. It's a concept that's pretty easy to understand. Yet turn to artificial intelligence and it just seems to break down. People making weird assertions all over the place not because they have been warranted in any empirical, qualitative or quantitative manner but because there seems to be this inability to engage with results...like we do with each other.
when i show the output that clearly demonstrates reasoning and understand, the arguments quickly shift to "it's not real understanding!" and it's honestly very bizarre. What kind of meaningful distinction can't show itself, can't be tested for ? If it does exist then it's not meaningful.
I think that the same reason people shift posts for intelligence is the same reason people fear the philosophical zombie.
idk maybe i'm rambling at this point but just my thoughts.
I agree with you totally. Here's some output that clearly demonstrates reasoning and understanding:
All men are mortal.
Socrates is a man.
Therefore, Socrates is mortal.
I just copy-pasted that from wikipedia. Copy/paste understands syllogisms!Explain that!
I'm Just surprised honestly. All of the straw man arguments, this is the best you can come up with?
Man do better. GPT-4 would have a better response than this.
Are you really just going to cop out and avoid engaging seriously with my question? How do you explain the output above except as reasoning and understanding?
I think that's because you have no idea how to explain why something is, or isn't understanding, or reasoning.
In your comment above you accuse people of not being able to tell you clearly why GPT-4 is not reasoning or understanding, but you, yourself, can't even say clearly why copy/paste isn't. You have no clue how to do that. If you can't even say why something isn't reasoning, or understanding, then how can you say that something is? Do you even know what you're talking about, when you're talking about "reasoning" and "understanding"?
I find it very difficult to believe at this point that you are arguing in good faith.
Your COPY-PASTE example is nonsensical as already pointed out.
Do you also think that a printing press (or a rubber-stamp for that matter) demonstrates reasoning?
But will you explain why it is nonsensical? Because that has not yet been pointed out.
Lest you forget, here is what the OP said that I replied to:
>> when i show the output that clearly demonstrates reasoning and understand, the arguments quickly shift to "it's not real understanding!" and it's honestly very bizarre.
I also showed "output" that "clearly demonstrates reasoning and understand[ing]". Why is that nonsensical?
Edit: although it won't be very helpful to discuss this if you're not the OP and don't make the same assumptions as they. But go ahead anyway.
og_kalu, I understand the point you're trying to make regarding the potential intelligence of GPT-4 and the connection with the philosophical zombie, but I believe there are some important distinctions to consider.
First, it's important to recognize that the goalposts for artificial intelligence have indeed been shifting, and for good reason. As our understanding of intelligence grows, so does our ability to build systems that can mimic it. However, this doesn't necessarily mean that a given AI system, like GPT-4, has truly achieved general intelligence. Instead, it might simply be that our models are becoming more sophisticated and better at solving specific tasks.
The philosophical zombie argument, on the other hand, is concerned with subjective experience and consciousness, rather than intelligence. A philosophical zombie is a hypothetical being that is behaviorally and functionally identical to a human being, but lacks subjective experience. The debate around the philosophical zombie is more about the nature of consciousness and whether it can be separated from intelligence, rather than the intelligence itself.
Now, regarding your assertion that true distinction shows in results, it's true that GPT-4 and similar models have shown impressive capabilities. However, it's crucial not to confuse correlation with causation. Just because an AI system can generate outputs that seem to demonstrate reasoning and understanding, it doesn't necessarily mean that it possesses true understanding. It might simply have learned to generate outputs that are highly correlated with human-generated responses, without any actual understanding or reasoning taking place.
In summary, while it's true that AI systems like GPT-4 are becoming more advanced and able to generate seemingly intelligent responses, it's important to differentiate between the appearance of intelligence and genuine understanding. Furthermore, the philosophical zombie argument is primarily concerned with consciousness, not intelligence, so it may not be entirely relevant in this context.
I don't think this is a goalpost shift at all. I agree that GPT actually is "generally intelligent" but then so is Google Search. The point of the term AGI is that it has generally applicable intelligence that is on par with a human, I don't think that's really changed.
The problem with GPT is still that every single interaction I've had with it, I've pointed out issues with its logic and it is incapable of understanding my objection. It agrees with whatever I say and then immediately repeats the same mistake. And these aren't "invent new math" questions, there's a very clear inability to follow a logical chain of cause and effect.
I should be clear at the same time I still have a feeling ChatGPT might be conscious. It's obviously kind of a dreamlike consciousness without ability to hold state but it does feel like it could be conscious.
Which human or humans?
I freely admit that I do not posses the intelligence to "invent new math" but I am pretty sure that I am "smarter than the average bear" (to borrow a phrase).
Interesting q for us to consider is how do we come up with "new ideas". I think we can consider play (in the abstract sense) to be a significant element of the process. Play is a pleasurable self-motivated activity.
I am certain an AGI (if such a thing exists) will need to be playful.
[1] https://www.amazon.com/Finite-Infinite-Games-James-Carse/dp/...
[2] https://www.amazon.com/Grasshopper-Third-Games-Life-Utopia/d...
spooner> You are a clever imitation of life... Can a robot write a symphony? Can a robot take a blank canvas and turn it into a masterpiece?
sonny> Can you?
edit: i just realized OpenAI can answer both of those questions with "yes." ...
ChatGPT went from struggling to provide the answer for 20x20 to easily being able to provide the right answer for any math problem wolfram alpha can.
You're just kind of re-emphasizing my point: ChatGPT is using Wolfram as its assistance. So really, it's acting more like a "dumb" API call, not a "math genius" at all.
I mean this honestly, no snark. When I was a kid, learning how to factor numbers was really hard. It took a lot of time and concentration to do even basic problems, and people who could do it quickly without much perceived effort were a mystery to me.
By the time I reached high school, I had enough practice that I recognized the common patterns without difficulty, and often the answer bubbled up to my conscious mind without thinking about it at all. It sure feels like my brain is making an API call to a subsystem, you know?
it seems like you’re arguing whether the genius label only has to be innate. But if you are able to effectively get help from sources, the effect can be the same.
It takes you, what, seconds to type digits into a calculator?
But if you embedded a calculator into your brain and could put in and pull values out of it in microseconds how are you different from a math genius?
Same with millisecond access to Google earth, at some point smarts+speed of access is a system within itself.
Tao, Perelman, Wiles etc. aren't math geniuses because they can multiply numbers fast (which is a super weird definition of "math genius" tbh). They're math geniuses because they answer really hard questions in creative and unexpected ways; ways that often open up entire new areas of mathematics.
An interesting intelligence emerges from a colony of ants compared to one. An interesting intelligence emerges from the relations of the various parts of the brain. An interesting intelligence emerges from the combo of techniques that have led to GPT and other surprising AIs. Same with a human, to two humans collaborating, to tribes to companies to cities etc. And likely same by combining/integrating GPT-4 type of LLMs with other tools and with other specific purpose LLMs - a greater unit of significance will emerge. (The non-predictability of this with current science is what is kind of scary to me.)
You're getting stuck on semantics.
I expect Tao, Perelman, Wiles, etc. can process/handle/deal with numbers much more rapidly than I can.
Now imagine if these folks could do the same, except 10x or 100x or 1000x faster, without having to ever be bothered with things like sleep or food.
This may not be a good definition of a "Math Genius" but it is reasonable to think that all or most "Math Geniuses" possess this skill. Speed matters.
We can go back and forth splinting hairs about whether inserting a compute module into a neural net (organic or not) grants geniousness or assistance, but the overall point stands; there will be a single interface that can take any variety of inputs and properly parse and shape them, push them through the correct "APIs", and then take the results to form a clear and correct output. Whether or not it used it's neural net or transistor adders circuits to arrive at the answer would be immaterial.
I'm convinced that doesn't require a breakthrough in the architecture or scale of LLMs, though. GPT-4 seems plenty smart enough to be able to learn that, if hooked up to a proof assistant like Lean or Coq with the right fine tuning iteration loop.
They both created "new Chess / Go" strategies and insights that GMs and top engines hadn't seen before.
It seems weird to me that people that should understand how ML works (e.g. sama) instead of educating laypeople about about how this actually works and the huge limits the technology has, start talking nonsense about AGIs like some random scifi fan at a convention. Depressing.
And also, sometime I wonder how many people hold those kind of opinions on AGI(imminent, doable, worthy of being discussed) because they sincerely believe in some nonsense like the basilisk thing for example.
While, you can believe whatever you believe, please don't bash people as laypeople or something as nonsense.
Not trying to bash anyone, I just meant normal people, outside of the field. Enthusiastic nonsense is still nonsense.
It's because he's been saying that for 40 years.
Hinton said he didn’t believe it before ChatGPT came out, that AGI is possible in 20 years.
It is simple to just checkout peoples’ own words, otherwise you are hallucinating the same way as those LLMs do
Is all the discussion in this thread about the abilities of ChatGPT in 5-10 years, or is it about its abilities right now?
Will we get very good software in the form of personal assistants that could answer questions in depth about a large number of topics? Most likely. Chat GPT is very good at a lot of things, and its only going to get better. Compute amount is going to be an issue though.
Will we get something that resembles a human sitting behind a keyboard, except an expert in every single subject on the internet with some markov chain process that determines objective subfunctions to minimize? Could be. Thing is though, there still has to be an overarching code to run that markov chain, which is not AGI. Definitely dangerous in the wrong hands, but not dangerous by itself - CTRL C could stop its execution at any time.
Will we get a AGI that is smart enough to circumvent that ctrl c signal, break containment, take over the world, and kill all humans, all in the spirit of turning the earth into a super computer? Nah. To even start to code that AGI, you need to get around the whole principle of computational irreducability which means first proving that P=NP. An AGI will essentially have to run simulations on the world internally to figure out how things could play out, at an information amount threshold that is less than the information in the world its trying to simulate, which (unless P=NP) is impossible.
In 2017 an Anon on 4chan proposed that our society, due to content marketing and AI stock trading, is already run by a sort of emergent intelligence, and he proposed that this is why things seem to be 'getting crazy' lately. I'm inclined to agree with that perspective, and if correct more capable AI systems could really kick this into high gear.
You can outthink a human by thinking 1000x faster, or by having 1000 minds think on the same problem. There's a reason everything explodes when communication technology gets better. Imagine how long it would take to coordinate the fundamental pieces of a laptop computer by pigeon mail or horseback.
I haven't officially "launched" yet, but it's working and you can play with it here, if anybody is up to giving the alpha a spin:
Let R+ denote the set of positive real numbers. Find all functions f : R+ → R+ such that for each x ∈ R+, there is exactly one y ∈ R+ satisfying xf(y) + yf(x) <= 2.
It just struggeld to reason. So I would be very surprised if the plugin somehow helps here.
Query i used was: ``` Can you think step by step about this math problem and solve it?
Let R+ denote the set of positive real numbers. Find all functions f : R+ → R+ such that for each x ∈ R+, there is exactly one y ∈ R+ satisfying xf(y) + yf(x) <= 2. ```
Response: https://pastebin.com/sTXM9kLt
Edit: Maybe i should say Bing not GPT-4 because i asked it there.
Should be simple for a computer to just try every parabola but all the LLMs try to solve it like a human would (and get it wrong)
Fwiw though, using this kind of mashup approach of gluing lots of things together, I suspect you'll get something that appears to match at least some people's definition of AGI.
I think a lot of the issues people have interacting with ChatGPT stem from overloading the context, and having a tree of specialised contexts recursively building on each other would I think focus each context while providing for broad and deep exploration of the top level request
GPT3.5 and GPT4 are already general purpose with text tasks. You can give it _any_ well defined programming task that involves a fairly direct route from natural language to code, along with the API/module description (or output data format), and it can do it. That's the only reason this and 500 other papers or services like it (but doing different things) are possible.
3.5 is not multimodal though. That's why it needs the other models. But GPT4 has image understanding and can do a lot of these things without the external tools.
There is no reason to think that similar variations of GPT will not be able to handle video understanding or generation at some point.
Are we going fast or is it just because of the buzz? I have a hard time separating the two.
It’s very exiting and I want to keep being exited. I don’t want to become terrified.
We are going too fast. Have been for years. This is just the first clear indication.
Braking is fatal, but some seem pretty hell-bent.
Deceleration is complicated, and it seems highly unlikely that there would be sufficient consensus for true deceleration. Local deceleration is simply waiting for the acceleration occurring somewhere else to overcome your efforts.
The math hasn’t really changed for most individuals. At some point something big will happen.
Singularity.
Be excited and put your efforts towards what you value. One way or another, there is very little time left for wasting.
Singularity is a belief system, it has very little to do with AI.
edit: I also think if we would ever get to a point where AI gets close to a point of possibly getting out of control as you imply it would simply be banned in the US/Canada/EU/Australia. Furthermore Latin America and Africa could and would be pressured to go along if needed. Which leaves some parts of Asia. China, maybe India and Russia. Probably only China. It could be cut off from the Internet if needed. We could build up a wall just like in the Cold War. My point being: This will not happen just because it happens. It will be a choice.
This needs to be the foundation
It will use every tool you give it to reach the goal you give him. It will try forever if needed.
I bet state actor are already plugging in tooling to register social account and automate credible propaganda. Maybe not with gpt itself, but privately hosted fine tuned models.
This can win elections.
You can plug wordpress and build infinite blogs with infinite post with the unique scope of building mass around a topic.
This can alter Wikipedia, many people don't ever check sources and take it at face value.
You can not only build fake research paper, but fake the whole research team and their whole interactions with the community and investor.
This can fraud millions.
Tools enable this today.
This simply exposes all of the cracks in the foundations of our society. There are severely exploitable issues, and we may wind up with a planet-level chaos monkey.
Ghostbusters had the stay-puft marshmallow man. Will the form of our destroyer be the woot monkey?
We're dealing with a very nascent AI revolution right now. A social phase transition has already started: on the forefront there are graphic artists (midjourney v5 is literally revolutionizing the industry as we speak) and NLP researchers (GPT-4 has reduced the need for applied NLP to basically zero), but it's only a start. The cheapness and availability changes everything.
The industry was a joke prior to this. No offense.
Humanity has been going too fast for a long time.
Tell me about this singularity belief system. I simply meant something stronger than an inflection point, closer to the mathematical sense than what you must be assuming, but that word must mean something more for you.
bake cookies just like the nice old grandma does for the local teenage gangs. its a survival mechanism so that nobody messes with you.
Not if, but when, we get around to AI law enforcement with lethal weapons, it's over. There's no going back from that.
Also they dont need to raise and feed an army if AI powered drones are able to do the same job. When AI is able to replace jobs en masse its pretty transparent to expect AI armies to be built en masse as well.
I'm pretty sure this is not obvious.
This needs to be the foundation
What are we to do, then? No snark, honest question.
Earth's history, sure. Humanity's history though? Living during the Apollo program is clearly more interesting than living in a period of relative stasis. Living during the AGI revolution could be more interesting still, we'll have to see.
I believe my uneasiness stems from the unknown potential of this technology.
Now, one could argue that that’s always the case for technology, so why do I feel uneasy now?
I believe that this particular type of technology has a very large potential. I think most HN readers would agree.
But I am not scared, yet. I’ll be scared when I know the technology will do serious damage, until then it’s an uneasy feeling. Then I’ll probably be terrified.
Someone born in 1985 who is now 37-38 was brought into a world where the internet barely existed, and was barely an adult when the iPhone launched. There’s still a lot more that can happen.
Don’t listen to pessimists: the world will look very different, but we apes have a way of adapting.
It’s a parlor trick, even if you add plugins or the ability to call other hugging face ML models - it’s just a parlor trick with fancier bells and whistles. All it is doing is using stochastic gradient descent to predict the next word in a sequence based on an enormous sophisticated training set designed to amaze people.
Thinking it has advanced because it can now get calculations correct is a fallacy. It’s still just predicting the next word, it’s just that it’s now got a post processing step that is converting those next words into code and parroting the output. It maybe be able to now answer 4567*9876 correctly (using the human hardcoded wolfram alpha engine) but it still does not fundamentally comprehend why 1+1=2 - like my 5 year old can.
Until it can generate its own internal neural networks to for example learn to logically reason about calculations we are still far from AGI. Also those calling for more data are misguided - less data, more sophisticated architectures than transformers are the only way to avoid the stochastic parrot trap.
It's such a big "just". You are just firing neurons. The stock market is just supply and demand. The internet is just a bunch of computers talking through 50 year old protocols that don't work very well.
Everything is just something else! I wonder if the first tribe to be annihilated by bronze weapons were like "that stuff is just like stone but more malleable, don't see what the big deal is".
Guess what, it apolgised immediately after and then again when I asked why it apologised even after I told it not to.
So your argument is probably more accurate for the other camp, or at least as accurate for the other camp as well.
If AI researches say, "this is just X and it can do Y!" then fine, that's just framing for "look: Y". When stochastic parrot guys say "this is just X, what's impressive about that?" it throws me for a loop coz they are are refusing to engage with Y.
I like your bronze sword analogy. From my point of view chatgtp is not a bronze sword, it’s a Stone Age sword that someone has painted bronze. It has value because people realize the advantage that a true bronze sword would have in a battle. However, when you actually put it through it’s paces you quickly realise it offers no actual value over what came before.
I understand your point, but I am struggling to see why it matters. This seems more and more an argument like “cars are not horses”. I know they are not but does it matter? Cars are superior for our use cases.
What a curious psychological study, maybe dyslexic people feel more threatened by a large language model so clearly understanding words that they’re more likely to attempt to discredit it?
Well neural networks have unpredicted emergent properties. I don't see how anyone can rule out or know future behaviour
“The bitter lesson” would like to have a word. http://www.incompleteideas.net/IncIdeas/BitterLesson.html
I appreciate your enthusiasm, but the history of ML shows that your approach is less likely to work. Maybe you’ll be the one to prove everyone else wrong. Architectural breakthroughs are few and far between, and it’s incredibly difficult to reason about. I came up with the Lion optimizer while Google was using random tree search across 300 TPUs to discover the same thing, and it’s just five lines or so.
Predicting the next word is a much deeper problem than people like you realise. To be able to be good at predicting the next word you need to have an internal model of the reality that produced that next word.
GPT-4 might be trained at predicting the next word, but in that process it learns a very deep representation of our world. That explains how it has an intuition for colours despite never having seen colours. It explains why it knows how physical objects in the real world interact.
Now, if you disagree with this hypothesis it's very easy to disprove it by presenting a problem to GPT4 that is very easy for humans to solve but not for GPT4. Like the Yann Lecun gear problem, which GPT4 is also able to solve.
Now that’s an interesting claim - that I would deeply dispute. It learns from text. Text itself is a model of reality. So chatgtp if anything proves that in order to be good at predicting the next word all you need is a good model of a model of reality. GTP knows nothing of actual reality only the statistics around symbol patterns that occur in text.
>> "good model of a model of reality"
That is just a model of reality. Also, a "model of reality" is what you'd typically call a world model. Its an intuition for how the world works, how people behave, that apples fall from trees and that orange is more similar to red than it is to grey.
Your last line shows that you still have a superficial understanding of what its learning. Yes it is statistics, but even our understanding of the world is statistical. The equations we have in our head of how the world works are not exact, they're probabilistic. Humans know that "Apples fall from the _____" should be filled with 'tree' with a high probability because that's where apples grow. Yes, we have seen them grow there, whereas the AI model has only read about the growing on trees. But that distinction is moot because both the AI model and humans express their understanding in the same way. The assertion we're making is that to be able to predict the next word well, you need an internal world model. And GPT4 has learnt that world model well, despite not having sensory inputs.
Computer-generated random numbers are not truly random, yet they are practically random in most real-world use cases. You can’t easily cheat the RNG in World of Warcraft to get critical strike every time.
The output from GPT is generally very intelligent and versatile in terms of text. It may even be capable of handling more multi-modal problems with the use of enough sensors and motors. Perhaps the same idea of "predicting the next move" or "predicting the next idea" can still apply.
Who knows, maybe humans are essentially physical creatures that "generate the next thought and generate the next move"?
One of the biggest issues with GPT is its lack of mid-term memory like human do. Instead, we need vector store and search then bolt back its short term memory instead of letting it handle everything in a more coherent way. Perhaps it could benefit from lightweight fine-tuning technologies like LoRA and hypernetworks for stable diffusion. If this issue is resolved we would see it'll get even more practical. Again, the flaw is not about "predicting the next words".
>A text can describe the given image: a herd of giraffes and zebras grazing in a fields. In addition, there are five detected objects as giraffe with score 99.9%, zebra with score 99.7%, zebra with 99.9%, giraffe with score 97.1% and zebra with score 99.8%. I have generated bounding boxes as above image. I performed image classification, object detection and image captain on this image. Combining the predictions of nlpconnet/vit-gpt2-imagecaptioning, facebook/detr-resnet-101 and google/vit models, I get the results for you.
Is it just that the in-context demonstrations are also ungrammatical and ChatGPT is copying them? It feels very unlike the way ChatGPT usually writes.
Telling it to fix its grammar, in a new thread fixes it.
I also confidently assume that telling it to fix its grammar, within the original, topic specific, conversation, would noticeably harm the quality of subsequent output for that topic.
When an AI has a goal and an ability to break the goal down, and make progress towards the goal… the only thing stopping the AI from misalignment is whatever it’s creator has specified as the goal.
Things get even more tricky when the agent can take in new information and deprioritize goals.
I have been curious why OpenAI haven’t discussed Agent-based AIs.
And the API providers willingness/capability to provide the service being requested.
I'm with those who don't see the danger, besides humans' own stupidity being amplified and someone trusting "AI" with something life critical, what should I be watching out for?
I can already feel the HuggingFace ChatGPT plugin coming. The only problem is it would be very slow and use a bunch of tokens.
The language model, being trained on essentially the primary way in which humanity communicates, might be a good means of managing integration of less language-focused models.
…duh?
now LLMs are on this list. a thing which isn't born, doesn't experience, can be copied and instantiated a million times, and a single improvement can be basically immediately propagated to all instances. a very unfamiliar thing with very familiar capabilities.
so, technically, it was obvious. socially and psychologically, we'll be dealing with it for the rest of our civilization's lifetime.
GhatGPT is the ultimate and last glue layer we will ever need.
These types of things must feel similar.
I'm finding it very difficult to differentiate between (sometimes involuntary) sales people and genuine usefulness (of which there is plenty to be clear).
It is yet unclear where the conceptual and fundamental limitations of the current approach lie and I'm immediately suspicious about anyone telling me there aren't any in particular when this sentiment is based on a few hours of GPT-4 usage.
We need more hard science on this. Great claims require great proofs and "AGI is imminent" is a very very very great claim.
We know that these tool are pretty useful as they are right now and will have massive influence on society.