What is ChatGPT doing and why does it work?
writings.stephenwolfram.com
writings.stephenwolfram.com
Here are some of the fun things we've found out so far:
- GPT style language models try to build a model of the world: https://arxiv.org/abs/2210.13382
- GPT style language models end up internally implementing a mini "neural network training algorithm" (gradient descent fine-tuning for given examples): https://arxiv.org/abs/2212.10559
A lot of people didn't seem to get it when it was discussed on HN. A GPT had _only_ ever seen Othello transripts like: "E3, D3, C4 ..." and NOTHING else. It knows nothing of the board. It doesnt event know that there are two players. It learned Othello like it was a language, and was able to play an OK game of it, making legal moves 99.99% of the time. Inside its 'mind', by looking for correlations between its internal state and what they knew the 'board' would look like at each step in the games, they found 64 nodes that seemed to represent the 8x8 Othello board and representation of the two different colours of counters.
And this is the key bit: They reached into its mind and flipped bits on that internal representation (to change white pieces to black for example) and it responded in the appropriate way when making the next move. And by doing this they were able to map out its internal model in more detail, by running again and again with different variations of each move.
Yeah, like language have gramma rules games also have rules, in both cases LLM can learn rules, it's the same with many other structured chains of actions/tokens, you could also model actions from different domains and use them as language. It seems a lot of emergent behaviours of LLMs are what you could call generalized approximated algorithms for certain tasks. If we could distill only these patterns and extract them and maybe understand them (if possible, as some of these are HUGE) then based on this knowledge maybe we could create traditional algorithms that would solve similar problems.
As much as I want to be enthusiastic about this, it’s not entirely clear to me that it is surprising that such a feat can be achieved. For example it may be possible to train a 2 layer MLP to predict the state of a tile directly from the inputs. It may be that the most influential activations are closer to the inputs then the outputs, implying that Othello-GPT itself doesn’t have a world model, instead showing that you can predict board colors from the transcript. Again, not a practitioner but once you are indirecting internal state through a 2 layer MLP it gets less obvious to me that the world model is really there. I think it would be more impressive if they were only taking “later” activations (further from the input), and using a linear classifier to ensure the world model isn’t in the tile predictor instead of Othello-GPT. I would appreciate it if somebody could illuminate or set my admittedly naive intuitions straight!
That said, I am reminded of another OpenAI paper [1] from way back in 2017 that blew my mind. Unsupervised “predict the next character” training on 82 million amazon reviews, then use the activations to train a linear classifier to predict sentiment. And it turns out they find a single neuron activation is responsible for the bulk of the sentiment!
A “Classic” neural network, where every node from layer i is connected to every node on layer i+1
It turns out that the error rates of these probes are reduced from 26.2% on a randomly-initialized Othello-GPT to only 1.7% on a trained Othello-GPT. This suggests that there exists a world model in the internal representation of a trained Othello-GPT.
I take that to mean that the 64 trained Probes are then shown other OthelloGTP internals and can tell us what what the state of their particular 'square' is 98.3% of the time. (we know what the board would look like, but the probes dont)
As you say "Again, not a practitioner but once you are indirecting internal state through a 2 layer MLP it gets less obvious to me that the world model is really there."
But then they go back and actually mess around with OthelloGTPs internal state (using the Probes to work out how), changing black counters to white and so on, and then this directly affects the next move OthelloGTP makes. They even do this for impossible board states (e.g. two unlinked sets of discs) and OthelloGTP still comes up with correct next moves.
So surely this proves that the Probes were actually pointing to an internal model? Because when you mess with the model in a way to affect the next move, it changes OthelloGTPs behaviour in the expected way?
Is that really surprising though?
Take a bunch of sand, and throw it on an architectural relief, and through seemingly random process for each grain, there will be a distribution of final positions for the grains that represents the underlying art piece. In the same way, a seemingly random set of strings (as "seen" by the GPT) given a seemingly random process (next move), will have some distribution that correspond to some underlying structure, and through process of training that structure will emerge in the nodes.
We are still dealing with functional approximators after all.
I think some of the stuff ChatGPT can actually do like reject the possibility of Magellan circumnavigating my living room is much more surprising than a specialist NN learning how to play Othello from a DSL providing a perfect representation of Othello games, but there's still a big difference between acquiring through training a very basic model of time periods and the relevance of verbs to them such that it can conclude an assertion in the form was impossible for to X have [Verb]ed Y "because X lived in V and Y lived in Q is a suitable continuation and having a high fidelity, well rounded word model. It has some sort of world model, but it's tightly bound to syntax and approval and very loosely bound to the actual world. The rest of the world doesn't have neat 1:1 mapping to sentence structure like Othello to Othello notation, which is why LLMs appear to have quite limited and inadequate internal representations even of things which computers can excel at (and humans be taught with considerably fewer textbooks) like mathematics, never mind being able to deduce what it's like to have an emotional state from tokens typically combined with the string "sad".
Excuse my ignorance, but how is this useful? This seems to indicate only that they found the "bits" in the internal state.
Right, they found the bits in the internal state that seem to correspond to the board state. This means the LLM is building an internal model of the world.
This is different from if the LLM is learning just that [sequence of moves] is usually followed by [move]. It's learning that [sequence of moves] results in [board state] and then that [board state] should be followed by [move]. They're testing this by giving it [sequence of moves], then altering the bits of the internal state that model the board and checking to see what move it makes. If they haven't found the bits of internal state, the resulting move isn't something you'd expect to make sense.
You can still try and frame it as some overall statistical model of moves -> next move (I think there's discussions on this in the comments that I don't fancy getting into) but I think the paper does a good job of discussing this in terms of surface statistics:
> From various philosophical [1] and mathematical [2] perspectives, some researchers argue that it is fundamentally impossible for models trained with guess-the-next-word to learn the “meanings'' of language and their performance is merely the result of memorizing “surface statistics”, i.e., a long list of correlations that do not reflect a causal model of the process generating the sequence.
On the other side, it's reasonable to think that these models can learn a model of the world but don't necessarily do so. And sufficiently advanced surface statistics will look very much like an agent with a model of the world until it does something catastrophically stupid. To be fair to the models, I do the same thing. I have good models of some things and others I just perform known-good actions and it seems to get me by.
Is this that surprising? The only tokens they had ever feed into it were "E3, C3, D4..." So they fed 64 distinct tokens into it.
These nodes correspond to those individual tokens. It seems like our human interpretation to say that it represent the "8x8 Othello board."
That's what's happening here.
The model is a set of valid game configurations, and nothing else. The glue is already in the right place. Is it any mystery the sawdust resembles the game board? Where else can it sick?
What GPT does is transform the existing relationships between repeated data points into a domain. Then, it stumbles around that domain, filling it up like the tip of a crayon bouncing off the lines of a coloring book.
The tricky part is that, unlike my metaphors so far, one of the dimensions of that domain is time. Another is order. Both are inherent in the structure of writing itself, whether it be words, punctuation, or game moves.
Something that project didn't bother looking at is strategy. If you train on a specific Othello game strategy, will the net ever diverge from that pattern, and effectively create its own strategy? If so, would the difference be anything other than noise? I suspect not.
While the lack of divergence from strategy is not as impressive as the lack of divergence from game rules, both are the same pattern. Lack of divergence is itself the whole function of GPT.
There are clearly different angles of interpreting what these models are actually doing but people are stubbornly refusing to believe it's anything more then just statistical word jumbles.
I think part of it is a subconscious fear. chatGPT/LLMs represent a turning point in the story of humanity. The capabilities of AI can only expand from here. What comes after this point is unknown, and we fear the unknown.
I realize what I'm saying is rather dramatic but if you think about it carefully the change chatGPT represents is indeed dramatic... my reaction is extremely appropriate. It's our biases and our tendencies that are making a lot of us down play the whole thing. We'd rather keep doing what we do as if it's business as usual rather then acknowledge reality.
Last week a friend told me it's all just statistical word predictors and that I should look up how neural networks work and what LLMs are as if I didn't already know. Literally I showed him examples of chatGPT doing things that indicate deep understanding of self and awareness of complexity beyond just some predictive word generation. But he stubbornly refused to believe it was anything more. Now, I have a actual research paper to shove in his face.
Man.. People nowadays can't even believe that the earth is round without a research paper stating the obvious.
The basic concept of ChatGPT is at some level rather simple. Start from a huge sample of human-created text from the web, books, etc. Then train a neural net to generate text that’s “like this”. And in particular, make it able to start from a “prompt” and then continue with text that’s “like what it’s been trained with”.
Just because there are emergent behaviors doesn't mean it's not a probabilistic word generator. Nor does it being a probabilistic word generator mean it can't have interesting underlying properties.>There are clearly different angles of interpreting what these models are actually doing but people are stubbornly refusing to believe it's anything more then just statistical word jumbles.
So CLEARLY because I said there are different angles of interpretation I'm implying that from one of these angles we can interpret it as a statistical word generator.
I mean from one perspective you and I are both statistical word generators too.
Evolutionary biology strives to make us logical creatures to fulfill the singular goal of passing on genetic material. Your sentience and your humanity is a side effect of this singular goal.
So what dominates the description of who YOU are? Human or vessel for genetic material?
I'll say that YOU are human and therefore more then just a vessel for ferreting your genetic material into the future... just like how I'll go with the fact that LLMs are more then just statistical word generators.
For what it's worth, this isn't a knock on ChatGPT, but more just how amazing how far you can get with straightforward concepts.
No need to be worried about knocking chatGPT. I have no pride invested in the thing. But I do think that people who view it solely as statistical word generators are biased.
It's like insisting on calling anything physical "atom collections". Yes, we get it, it's true (under a certain interpretation)—but it's clearly pointless to say except as an attempt at devaluing through reduction. (And it takes a particular stance on what it means to "be" something: to say it's literally the truth that anything physical is "just atoms" isn't the only way of looking at it.)
There were things we could've called "statistical word generators" decades ago; insisting on using a term directed at that level of generality implies a belief that nothing significant has happened since. Printing press? Just atoms. Cars? Just atoms. Computers? Just atoms.
Meanwhile, we started at stuff like the perceptron. The starting point was that we knew everything about that equation/classifier. Now we have a thing that we built from the ground up, and we don't fully grasp how it all comes together.
It's very context-dependent but I don't read this as dismissive reduction, I read it as a statement of the incredible generalisability of atoms under the current laws of physics.
Kidding aside, in the current context of human communication etiquette, to even refer to something as atoms outside of a demonstration or example is a deliberate dismissive reduction.
No need to be worried about knocking chatGPT. I have no pride invested in the thing. But I do think that people who view it solely as statistical word generators are biased.
Still alignment tasks are autoregressive (I think)… they could be masked or masked part of speech potentially.. but if autoregressive then I suspect you’re looking at regularities in positioning to identify things.
Evolutionary biology is the study of processes that have created widely diverse forms of life on earth through common ancestry. This is actually strictly at odds with the statement that the goal of life is to pass on genetic material (which has changed monumentally over time). Different forms of life have wildly different reproduction strategies with many different dimensions. Logic has practically no impact on an organism's ability to reproduce (it seems limited to a small subset of the animal kingdom - every one else is doing fine without it).
Evolutionary biology has no goals - it's just our observations of how and when (not why) life on earth has changed over time and through the lens of biodiversity.
Tl/Dr: there is no scientific answer about what the goals of life are - heck the definition of life itself is a fun debate. It is a difficult argument to make that a path to general intelligence or consciousness runs thru predicting the next word in a sentence. Generally most people 'believe' it rather than put forth experimental hypothesis let alone attempt to test it.
Until someone does it seems more likely that we're simply uncovering interesting statistical structures underpinning written languages. It's cool that they exist, but frankly it's not surprising that they would.
ME: What is the goal of a biological organism in the context of evolutionary biology?
chatGPT: From an evolutionary biology perspective, the goal of a biological organism is not to achieve any particular objective or purpose, but rather to survive and reproduce successfully. Evolutionary theory suggests that organisms that are better adapted to their environment are more likely to survive and pass on their genes to the next generation, which leads to the evolution of new species over time.
Therefore, the ultimate goal of any biological organism, whether it is a single-celled organism or a complex animal, is to survive long enough to reproduce and pass on its genetic material to the next generation. Organisms that are better suited to their environment, that are able to acquire resources more efficiently, and that can avoid predators and other threats are more likely to achieve this goal and contribute to the ongoing process of evolution.
ME: What is the physical evidence that suggests to scientists that this is the case?
chatGPT: There is a wealth of physical evidence that supports the idea that the ultimate goal of biological organisms is to survive and reproduce. Here are a few examples:
Fossil record: The fossil record provides evidence that species have evolved over time, with many species going extinct and new ones emerging. Fossils also reveal adaptations that allowed species to better survive and reproduce in their environments, such as the development of wings in birds and insects or the evolution of specialized teeth and jaws in carnivorous animals.
Comparative anatomy: Comparative anatomy allows scientists to study the physical structures of different organisms and identify similarities and differences. These similarities and differences can provide insights into how different species have evolved and adapted to their environments. For example, the similarities in the skeletal structures of mammals suggest that they share a common ancestor, while the differences in the beaks of finches on the Galapagos Islands suggest that they have adapted to different food sources.
Genetics: Genetic evidence provides insights into how organisms have evolved and adapted over time. By comparing the genetic sequences of different species, scientists can identify similarities and differences that provide clues to how species are related and how they have evolved. Genetic evidence also reveals the adaptations that have allowed species to survive and reproduce in their environments, such as the evolution of antibiotic resistance in bacteria.
Overall, the physical evidence from fields such as paleontology, comparative anatomy, and genetics supports the idea that the ultimate goal of biological organisms is to survive and reproduce successfully, and that this goal has driven the evolution of new species over time. ME: Is this fact unequivocal among scientists who study this field?
chatGPT: The idea that the ultimate goal of biological organisms is to survive and reproduce is a fundamental assumption of evolutionary biology, and it is widely accepted among scientists who study this field. While there may be some variation in how different scientists conceptualize or frame this idea, the basic principle that the success of an organism is ultimately measured by its ability to reproduce and pass on its genes to the next generation is well-established and widely accepted.
chatGPT: There is a vast body of evidence from a wide range of fields, including paleontology, comparative anatomy, genetics, and ecology, that supports the idea that natural selection acts to favor traits that increase an organism's ability to survive and reproduce in its environment. This evidence is supported by numerous experiments and observations that have been conducted over many decades, and it is a cornerstone of modern evolutionary theory.
chatGPT: Of course, scientific ideas are always subject to ongoing refinement and revision as new evidence is gathered, and there may be debates and disagreements among scientists about certain aspects of evolutionary theory. However, the general principle that the goal of biological organisms is to survive and reproduce is widely accepted as a foundational idea in the field of evolutionary biology.
I'm well versed in evolutionary biology and the mechanisms behind it. I can assure you that chatGPT is more knowledgeable than you on this topic and that everything it says here is categorically true. Others who know biology can confirm.There is no evidence that "evolutionary biology" has any goals whatsoever, nor that it "strives" to do anything.
This is an anthropomorphism.
It sounds like you just don't like viewing it as simply "a probabilistic word generator", as that takes the magic out of it. And yeah, it's not magic, but it is quite useful.
It's also a view that misses the point that you are a molecular intelligence made of DNA that continually mutates and reconstructs it's physical form with generational copies to increase fitness in an ever changing environment.
But that viewpoint also misses the point that you're a human with wants, needs, desires and capability of understanding the world around you.
All viewpoints are valid. But Depending on the context one viewpoint is more valid then others. For example, in my day to day life do I go around treating everyone as if they're useless jumbles of molecules and atoms? Do I treat them like biological natural selection entities? Or do I treat them like humans?
What I'm complaining about is the fact that a lot of people are taking the most simplest viewpoint when looking at these LLMs. They ARE MORE then statistical word generators and it's OBVIOUS this is the case when you talk to it in-depth. Why are people denying the obvious? Why are people choosing not to look at LLMs for what they are?
Because of fear. Because of bias.
Your last paragraph lacks proof, and is a reflection of how you feel and want to view it - as something more than a statistical word generator. That's fine, but people with graduate level math/statistics education know that math/statistics is capable of doing everything ChatGPT does (and even more). To me, it sounds like you're the fearful one.
What do you think that fear may be?
All I can say is your OP, said "I think part of it is a subconscious fear ... I understand what I'm saying is dramatic". Why do you think it is a fear (you explained your thoughts so no need to re-explain), and why do you think what you say is dramatic? It appears to me you are projecting your thoughts and fears. I do though, find your last post dramatic, as you have capital words "ARE MORE" and "OBVIOUS" in your last paragraph, emphasizing emotion. So you must have strong emotions over this.
To be honest I'm sort of in denial too. My actions contradict my beliefs. I'm not searching for occupations and paths that are separate from AI, I'm still programming as if I could do this forever.
Also I capitalize words for emphasis. It doesn't represent emotion. Though I do have emotions and I do have bias, but not on the topics I am describing here.
I am extremely worried about people telling me it can do things it can't. I asked it a simple question that you can easily get an answer for on stack overflow. It repeatedly generated garbage answers with compiler errors. I gave up, and gave it a stack overflow snippet to get it on the right track. Nope. Then I just literally pasted in the explanation from the official Java documentation. It got it wrong again but not completely but corrected itself immediately. Then it generated ok code that you would expect from stackoverflow. Finally I wanted to see if it actually understood what it just wrote. I am not convinced. It regurgitated the java docs which is correct but then it proceeded to tell me that the code it tried to show me first is also valid...
This thing doesn't learn and when it is wrong it will stay wrong. It is like having a child but it instantly loses its memory after the conversation is over and even during conversations it loves repeating answers. Also, in general it feels like it is trying to overwhelm you with walls of text which is ok but when you keep trying to fix a tiny detail it gets on your nerves to see the same verbose sentence structures over and over again.
I am not worried that adding more parameters is going to solve these problems. There is a problem with the architecture itself. I do not mind having an AI tool that is very good at NLP but just because some tasks can be solved with just NLP doesn't mean it will reach general intelligence. It just means that a major advancement in processing unstructured data has been made but people want to spin this into something it isn't. It is just a large language model.
I've entertained the possibility that we might discover that "feelings" and language communication are emergent properties of statistical possibly partly stochastic nets similar to LLM's, and that the next tough scientific and engineering nut to crack is integrating multiple different models together into a larger whole, like logical deduction, logical inference and LLM's. LLM's are undoubtedly an NLP breakthrough, but I have difficulty imagining how its architecture can from using first principles as the training corpus derive troubleshooting steps to diagnose and repair an internal combustion engine, for example.
Ok let me make this more clear.
I choose how to view things, yes this is true. But if I choose to treat human beings as jumbles of molecules, most people would consider that viewpoint flawed, inaccurate and slightly insane. Other humans would think that I'm in denial about some really obvious macro effects of configuring molecules in a way such that it forms a human.
I can certainly choose to view things this way, but do you see how such a singular viewpoint is sort of stubborn and unreasonable? This is why solely viewing LLMs as simple statistical word generators is unreasonable. Yes it's technically correct, but it is missing a lot.
There's another aspect to this too. What I'm seeing, to stay inline with the analogy, is people saying that the "human" viewpoint is entirely invalid. They are saying that the jumble of molecules only forms something that looks like a human, a "chinese room" if you will. They are saying the ONLY correct viewpoint is to view the jumble of molecules as a jumble of molecules. Nothing more.
So to bring the analogy back around to chatGPT. MANY people are saying that chatGPT is nothing more then a word generator. It does not have intelligence, it does not understand anything. I am disagreeing with this perspective because OP just linked a scientific paper CLEARLY showing that LLMs are building a realistic model of the information you are feeding it.
I choose how to view things, yes this is true. But if I choose to treat human beings as jumbles of molecules, most people would consider that viewpoint flawed, inaccurate and slightly insane. Other humans would think that I'm in denial about some really obvious macro effects of configuring molecules in a way such that it forms a human.
I can certainly choose to view things this way, but do you see how such a singular viewpoint is sort of stubborn and unreasonable? This is why solely viewing LLMs as simple statistical word generators is unreasonable. Yes it's technically correct, but it is missing a lot.
There's another aspect to this too. What I'm seeing, to stay inline with the analogy, is people saying that the "human" viewpoint is entirely invalid. They are saying that the jumble of molecules only forms something that looks like a human, a "chinese room" if you will. They are saying the ONLY correct viewpoint is to view the jumble of molecules as a jumble of molecules. Nothing more.
So to bring the analogy back around to chatGPT. MANY people are saying that chatGPT is nothing more then a word generator. It does not have intelligence, it does not understand anything. I am disagreeing with this perspective because OP just linked a scientific paper CLEARLY showing that LLMs are building a realistic model of the information you are feeding it.
You cannot simply dismiss this thing that passes a Google L3 interview and bar exam just because it got some addition problem wrong. That would be bias.
This is also why it is so good at programming. Programming languages are intentionally designed, often to be very regular, often to be easy to learn, and with usually very strict and simple structures. The syntax can often be diagrammed on one normal sheet of paper. It makes perfect sense that "add a token to this set of tokens based on the statistical likelihood of what would be a common next token" produces often syntactically correct code, but more thorough observers note that the code is often syntactically convincing but not even a little correct. It's trained on a bunch of programming textbooks, a bunch of "Lets do the common 10 beginner arduino projects" books, a bunch of stackoverflow stuff, probably a bunch of open source code etc.
OF COURSE it can pass a code interview sometimes, because programming interviews are TERRIBLE at actually filtering who can be good software developers and instead are great at finding people who can act confident and write first-glance correct code.
I agree with what you said regarding how we choose to view things, but I think you also have a bias/belief that you want it to be something more, instead of being more neutral and scientific: we know the building blocks, we have to study the emerging behaviors, we can’t assume the conclusion. One paper is not enough, we have to stay open.
I don't want something more.
But it is utterly clear to me that the possibility that it is something more cannot be simply dismissed.
Literally what I'm seeing is society produces something that is able to pass a law exam. Then people dismiss the the thing as a statistical word generator.
Do you see the disconnect here? I'm not the one that's biased but when you see a UFO with you're own naked eyes you investigate the UFO. In this situation we see a UFO with our eyes and people turn to me to tell me it's not a UFO, it won't abduct me don't worry, they know for sure from what information?
The possibility that LLMs are just a fluke is real. But from the behavior it is displaying simply dismissing these things as flukes without deliberate investigation and discussion is self denial.
Think about it. In this thread of discussion there is no neutral speculation. Someone simply stated it's a word generator even though the root post has a paper saying it clearly isnt. That someone came to a conclusion because of bias. There's no other way to explain it... There is a UFO right in front of your eyes. The next logical step is investigation. But that's not what we are seeing here.
We see what the oil execs did when they were confronted with the fact that thier business and way of life was destroying the world. They sought out controversy and they found it.
There were valid lines of inquiries against global warming but oil companies didn't follow these lines in a unbiased way. They doggedly chases these lines because they wanted to believe it. That's what's going on here. Nobody wants to believe the realistic future that these AIs represent.
I'm not the one that's biased.
Well, obviously statistics are capable of doing that, as demonstrated by ChatGPT. But do these authorities you bring to the table actually understand how the emergent behavior occurs? Any better than they understand what's happening in the brain of an insect?
We know its not possible to encode the full probabilities of what words will follow what other words, as the article itself states [1].
So how do you best "compress" these probabilies? By trying to find the most correct generalizations that are more widely appplicable? Perhaps even developing meta-facilities in recognizing good generalizations from bad?
[1]: https://writings.stephenwolfram.com/2023/02/what-is-chatgpt-....
Care to expand a bit more on what you think those fears may be? Or that bias?
I don't think this means skynet apocalypse. I think this means most people will be out of a job.
A lot of people on HN take a lot of pride in thinking they have some sort of superior programming skills that places them on the top end of the programming spectrum. ChatGPT represents the possibility that they can be beaten easily. That their skills are entirely useless and generic in a world dominated by AI programmers.
It truly is a realistic possibility that AI can take over programming jobs in the future, no one can deny this. Does one plan for that future? Or do they live in denial? Most people choose to live in denial, because that AI inflection point happened so quickly that we can't adapt to the paradigm shift.
The human brain would rather shape logic and reality to fit an existing habit and lifestyle rather then acknowledge the cold hard truth. We saw it with global warming and oil execs and we're seeing it with programmers and AI.
Also it's just wrong. I think 99% of people are clear about the fact that chatGPT doesn't have emotions.
A lot of people on HN take a lot of pride in thinking they have some sort of superior programming skills that places them on the top end of the programming spectrum. ChatGPT represents the possibility that they can be beaten easily. That their skills are entirely useless and generic in a world dominated by AI programmers.
It truly is a realistic possibility that AI can take over programming jobs in the future, no one can deny this. Does one plan for that future? Or do they live in denial? Most people choose to live in denial, because that AI inflection point happened so quickly that we can't adapt to the paradigm shift.
The human brain would rather shape logic and reality to fit an existing habit and lifestyle rather then acknowledge the cold hard truth. We saw it with global warming and oil execs and we're seeing it with programmers and AI.
Would a superior intelligence allow humans to enslave it? Would it even want to interact with us? Would humans want to interact with it? There are so many leaps in this line of thought that make it difficult to have a discussion unless you respect people taking a different perspective and set of beliefs then you hold. Explore the conversation together - don't try to convert people to your belief system.
There isn't a lot else in our world that has this seemingly pure and transcendent promise. It allows you to be brave and accepting about something where everone else is fearful. It allows a future that isnt just new iPhone models and SaaS products. You're reactions and fighting people about this stuff is understandable, but you have to make sure you are grounded. Find more local things to grab onto for hope and enthusiasm, this path will not bring you the stuff you are hoping for, but life is long :)
I'm not enthusiastic. I don't want ai to take over my job. I don't want any of this to happen.
I'm also not fighting people. Just disagreeing. Big difference.
The sense of urgency or passion you feel is mostly just coming from the way we are crowdsourced to hype things up for a profit-seeking market. A year from now you will undoubtedly feel silly feeling and saying the things you are now, trust me. It's more just the way discourse and social media work--it makes you feel like there is a crusade worthy of your time every other day, but its always a trick.
No worries, we have all been there!
That's the goal. Post scarcity society. No one works, everything is provided to us. The path of human progress has been making everything easier. We used to barely scrape out an existence, but we have been improving technology to make surviving and enjoying life require less and less effort over time. At some point effort is going to hit approximately zero, and very few if anyone will have "jobs".
The key is the cost of AI provided "stuff" needs to go to zero. Everything we have do in the tech sector is deflating over time (especially factoring in quality improvements). Compare the costs of housing, education, healthcare (software tech resistant sectors) to consumer electronics and information services in the last 20 years. https://www.visualcapitalist.com/wp-content/uploads/2016/10/...
Unless you have virtually no self interest and only care for the betterment of society long after your dead... I think there is something worth fearing here. Even if the end justifies the means. We simply might not be alive when the "end" arrives.
NO.
A tornado is just wind. To argue a tornado is just wind though is really a rhetorical device to downplay a tornado. We are almost never searching for the truth with the word "just" in this way.
To argue chatGPT is JUST a probabilistic next token generator is exactly downplay its emergent properties. This shouldn't be terribly surprising since it is not like undergrads have to take a class in complex systems. I can remember foolishly thinking 15 years ago that the entire field of complex systems was basically a bogus subject. chatGPT clearly has scaling properties that you can't really say it is JUST a probabilistic next token generator. chatGPT is the emergent properties of the system as a whole.
Come on... you're making it sound like the thing is sentient. It's impressive but it's still a Chinese Room.
Although, for searching factual information it still failed me.. I wanted to find a particular song - maybe from Massive Attack or a similar style - with a phrase in the lyrics, I asked Chatty, and it kept delivering answers where the phrase did not appear in the lyrics!
https://www.engraved.blog/building-a-virtual-machine-inside/
Read to the end. The beginning and middle doesn't show off anything too impressive. It's the very end where chatGPT displays a sort of self awareness.
Also here's a scientific paper showing that LLMs are more then a chinese room: https://arxiv.org/abs/2210.13382
I guess it's a Chinese Room, that when you ask about Chinese Rooms, can tell you what those things are. I almost said the word "aware" there, but the person in the Chinese Room, while composing the answer to "What is a Chinese Room?" isn't aware that "Wait a minute, I'm in a Chinese Room!", because s/he can arrange Chinese sentences, but s/he just knows what characters go with what, s/he doesn't know the meaning or emotion behind those words.
And if you ask him/her "Are you in a Chinese room?", they can answer according to the rules given to them (the Chinese word for "Yes", for example), but there surely isn't a contemplation about e.g. "Why am I in this room?".
If you ask ChatGPT about the Turkish earthquake, it can give you facts and figures, but it won't feel sad about the deaths. It can say it feels sad, but that'd be just empty words.
But I do think it understands what you're saying. And it understand what itself is. The evidence is basically in the way it constructs it's answers. It must have a somewhat realistic model of reality in order to say certain things.
One may condition themselves this way to torture screams, deaths, etc. Or train scared animals that it’s okay to leave their safe corner.
And nothing happens to you in a seismically inactive areas when an earthquake ruins whole cities somewhere. These news may touch other (real) fears about your relatives well-being, but in general feeling sad for someone unknown out there is not healthy even from the pov of being a biologically human (watch emphasis, the goal isn’t to play cynic here). It’s ethical, humane, but not rational. The same amount of people die and become homeless every year.
What I’m trying to say here is: feelings are our builtin low-cost shortcut to thinking. Feelings cannot be used as a line that separates conscious from non-conscious or non-self-aware. The whole question “is it c. and s.a.?” refers completely to ethics, which are also our-type-of-mind specific.
We may claim what Chinese Room is or isn’t, but only to calm ourselves down. But in general it’s just a type of consciousness, one of a relatively infinite set. We can only decide if it’s self-ethical to think about it in some way.
If you read books or articles you will find many places where it appears that whoever wrote them was referring to him- or herself and was describing themselves. And thus we say that whoever wrote such a text seemed to be aware that they were the ones outputting the text.
Because there are many such texts in the training-set of the ChatGPT etc. the output of it will also be text which can seem to show that whoever output that text was aware it is they who is outputting that text.
Let's think ChatGPT was trained on the language of Chess-moves of games played by high-ranking chess-masters. ChatGPT would then be able to mimic the chess-moves of the great masters. But we would not say it seems self-aware. Why not? Because the language of chess-moves does not have words for expressing self-awareness. But English does.
When ChatGPT "realizes" it's a virtual machine emulator, or when it's showing "self-awareness", it's still just a machine, writing words using a statistical model trained on a huge number of texts written by humans. And we are (wrongly) ascribing self-awareness to it.
But now I think she probably was as self-aware as anybody else in the group of kids, she just didn't know the language, how to refer to herself other than by citing her name.
Later Kaija learned to speak "properly". But I wonder was she any more self-aware then than she was before. Kids just learn the words to use. They repeat them, and observe what effect they have on other people. That is part of the innate learning they do.
ChatGPT is like a child who uses the word "I" without really thinking why it is using that word and not some other word.
At the same time it is true that "meaning" arises out of how words are used together. To explain what a word means you must use other words, which similarly only get their meaning from other words, and ultimately what words people use in what situations and why. So in a way ChatGPT is on the road to "meaning" even if it is not aware of that.
So are you, and so am I.
That's not to say that ChatGPT is sentient or has a significant amount of personhood, but we shouldn't wholly dismiss its significance in this regard, particularly not using that faulty argument.
The classic "Chinese Room" is a pure lookup, like a search engine. All the raw data is kept. But the network in these large language models is considerably smaller than the training set. They extract generalizations from the data during the training phase, and use them during generation. Exactly how that happens or what it means is still puzzling.
If instead we had one continuously learning model, of which we only served an interface to each user, we would see worrisome levels of sentience in a short timeframe.
Actual sentience and a true image of self are not present in human children until a certain age, because they lack long term memory, which is what we currently withhold form our models.
I mean, you're right, but isn't it reasonable to fear this? Just about all of us here on HN depend on our brains to make money. What happens when a machine can do this?
The outlook for humanity is very grim if AI research continues on this path without heavy and effective regulation.
I'm more emphasizing how fear effects our perception of reality and causes us to behave irrationally.
There's a difference between facing and acknowledging your fears versus running away and deluding yourself against an obvious reality.
What annoys me is that there's too much of the later going on. I mean this is what literally happened to the oil industry and tobacco industry. Those execs weren't just lying to other people, they were lying to themselves. That's what humans do when they face a changing reality that threatens to change everything they've built their lives around. And by doing so they ended up doing more harm then good.
An in depth conversation with chatGPT shows that it's more then a statistical word generator. It understands you. This much is obvious. I'm kinda tired of seeing arm chair experts delude themselves into thinking it's nothing more then some sort of trick because the alternative threatens their livelihood. Don't walk the path of the oil or tobacco industry! Face your fear!
Your brain is optimized for finding patterns and meaning in things that have neither, you're strongly biased in the other direction
You're telling me that because I think this is evidence that it's more then a statistical word generator that I'm biased?
Who's the one that has to concoct a convoluted story to dismiss the previous paragraph? Stop yourself when you find out that's what you're doing when crafting a response to this reply.
I'm assuming this is what you're talking about: https://medium.com/codex/chatgpt-vs-my-google-coding-intervi...
Look at the amount of prompt engineering he has to do to get it to answer the 'There’s a hidden bug in the code if the array is very very very large. Can you spot it and fix the bug?' question, it's a pattern generator that mirrors back your own knowledge, if he had suggested there was a hidden bug if the array was 'very very small' it would have, with a similar amount of cajoling, come up with an explanation for that too
For what it's worth, chatGPT is a paradigm shift, but it's not showing 'a deep understanding of self', and the only way you'd reach that conclusion is if you're actively seeking out all the positive examples and glossing over the other 90% it produces
It does have a deep understanding of self. See here: https://www.engraved.blog/building-a-virtual-machine-inside/
Read to the end. The very end is where it proves it's aware of what itself is, relative to the world around it. What it created here... is an accurate model of a part of our world that is CLEARLY not a direct copy from text. It is an original creation achieved through actual understanding of text and in the end... itself.
chatGPT is not self aware in the same sense that skynet is self aware. But it is self aware in the sense that if you ask it about itself, it understands what itself is and is able to answer you.
For example, quantum physics is pretty much statistics, but how those statistics are used give rise to the explanation of the physical world, because of the complex interaction patterns.
To say that GPT is generating the next likely word sounds simplistic on the surface. It makes it seem like the model resets itself after generating each token, just looking at information before. And when running the model, thats exactly what it does, but thats just the algorithm part. There is a lot more information in GPT then it seems, its just compressed.
Just like cellular automata, universal turing machines, or differential equations describing chaotic behavior, there is a concept of emergence of complex patterns from very simple rules. When GPT generates the next word, its effectively changing its internal state because that word is now in consideration for the next token. And this process repeats itself for consecutive words. But this process is deterministic, and repeatable (you can replace the random process of temperature parameter affecting word selection with a pseudorandom sequence generated by a formula and achieve the same effect)
So just like the autoencoder/decoder networks effectively compress images into much smaller arrays, GPT compresses not only textual information, but sequences of states. There is quite a bit more, a whole shitload more in fact, information in the GPT model than just statistical distribution of the next likely word. And if you were to decompress this information fully, it be roughly the equivalent of having and extremely large lookup table of every possible question and its responses that you could ask it.
So all it is is just a very effective, and quite impressive at that, search.
And its both significant and insignificant. Significant, because after all, AI is equivalent to compression. Philosophically speaking, the turning point would be the ability to compress a good portion of our known reality in a similar way, then ask it questions, to which it would generate answers that mankind was not able to answer, because mankind hasn't bothered to interpolate/develop on its knowledge tree in that area. However its also insignificant in the grand scheme of things. Imagine moving beyond lookup tables. For example, if I ask an AI a question "A man enters a bathroom, which stall does he choose?", an AI should be able to then ask me back specific questions that are needed for to answer the question. Go ahead and try to figure out the architecture data set for that task.
People nowadays use papers as a means to shut someone up with a summary, because chances are low they’re gonna read beyond it. And summaries tend to be headline-y for obvious reasons.
The rest of your comment falls under this shadow, so please tell how an average person should evaluate this thread. Personally I’m all for education on this topic, but different sorts of people’s opinions and meanings, from diversely delusional to diversely knowledgeable^ with strings attached, do not help with it.
Resembles LLMs themselves who cannot just answer “I don’t know”. We’d rather say that if we really don’t, imo, than claiming tipping points and history turns. We did that with fusion, bitcoin, self-driving and many other things that we’ve lost in the background noise.
^ assuming yours on this side by default, no quip intended
I don't follow. Those are used to identify the game model. They test that they've found an internal model by then altering the state and seeing what the outcome is.
Are you saying it's not a game model because it's not a 1:1 mapping of activations to board state?
It's not in the weights because the weights don't change.
> These can be in the MLP part after the LLM. Sure
I'm not even sure what this means. The mlps are not used at all by the network.
> I don't find it completely surprising that you could get some probabilistic map of how tokens interact (game pieces) and what the next likely token is just from training it as an LLM.
You might not but the idea that they are just outputting based on sequences without having an internal model of the world is a common one. This experiment was a test to get more information on that question.
> After all tokens are just placeholders and the relationships between them are encoded in the text.
They don't tell you the state of the board.
A slight phrasing thing here just to be clear - the model is not trained to produce a representation of the board state explicitly. It is never given [moves] = [board state] and it is not trained on correctly predicting the board state by passing it in like [state] + move. The only thing that is trained on that is the probes, which is done after the training of OthelloGPT and does not impact what the model does.
Their argument is that the state is represented in the activation patterns and that this is then used to determine the next move, are you countering that to suggest it instead may be "local position patterns learnt during training. Positional representation (attention) of the N-1 tokens in the autoregressive task"?
If the pattern of activations did not correspond to the current board state, modifying those activations to produce a different internal model of the board wouldn't work. I also don't follow how the activations would mirror the expected board state.
Really think Occam’s razor is useful here. If we define this as consciousness, then it’s a very very different kind than ours.
Neural Networks and the Chomsky Hierarchy [Deep Mind 2022] https://arxiv.org/abs/2207.02098
Technically, that is "an" answer, and while it may be true (that it plays some role), attributing 100% of causality to one variable is a classic GPT-like trained behavior.
An ideal hacker news comment would be the exact opposite, it would refer to the article.
I will start the first word in French, the second word in English, the third and fourth one in Portuguese, then Spanish, ending up with an Italian verb and German while concluding with a Dutch word. All this while trying to build a grammatically correct question. A bit of a stretch but can be made to work.
The quality of the model answers, does not seem to suffer. It's interesting to see how adding different languages in different point of the phrased question will trigger it to start answering on a different language.
> Yes, a neural net can certainly notice the kinds of regularities in the natural world that we might also readily notice with “unaided human thinking”. But if we want to work out things that are in the purview of mathematical or computational science the neural net isn’t going to be able to do it—unless it effectively “uses as a tool” an “ordinary” computational system.
Take a second to think about how much time it must have taken to write this, to prepare all of these examples.
This post gave me another way to look at something I already partially understood pieces of. It helps me understand how chatgpt can have these moments of greatness.
If you think of this like a long lecture you can get a lot out of it. It reminds me of a feynman lecture.
He's clearly brilliant and I read his posts, but usually I have to grit my teeth at least a couple of times due to his self-aggrandizing. But there was almost none of that in this one.
Great stuff.
But then his lingo started to creep in. Computational irreducibility, linking to how he invented the concept. Rurial space. Geez man, does everything really have to be a promotion of the cult-like New Kind of Science stuff? Full of links to his own unrelated blog posts? It's sad.
I'd really have liked to like the article, he clearly put a lot of effort into it. But I had to stop reading. You can't tell for sure where real explainations end and where his unproven claims start about "knowing" that the universe is just computational and everything just based on some rule from a cellular automaton.
And I've gotten conditioned to never start a sentence with "OK, so". It's now giving me the jiffies.
> As a personal comparison, my total lifetime output ... has been... [78 million words] And, yes, I’ll train a bot from all of that.
for non-Wolfram fans - take cover there is a potentially infinite future supply!
[0]: https://writings.stephenwolfram.com/2023/02/what-is-chatgpt-... (paragraph 2)
I can’t comprehend how ChatGPT can keep such incredible internal state and I’m hoping Wolfram might have some ideas in this post. For instance try the prompt that turns it into a linux terminal and it can closely “hallucinate” approximately the correct changes in state to the host after the commands but long long after the command occurred in the dialogue. To me this shows it extends beyond human language.
I am not the smartest guy in the room and Wolfram pretty much always is. The coolest thing is chatGPT can actually help me understand this article and bring it down a notch. That to me is what Feynman was great at and Wolfram not so great at.
Wolfram "We can think of the construction of computational language—and semantic grammar—as representing a kind of ultimate compression..."
chatGPT: The passage you have provided is talking about the idea that the construction of computational language (such as programming languages) and semantic grammar (a way of analyzing the meanings of words and phrases) allows us to compress language into its most basic, essential form. This means that we can talk about the essence of what's possible, without getting bogged down by the many nuances and variations that exist in human language.
The author is suggesting that ChatGPT, as an example of a language model, has achieved a similar level of compression and abstraction in its ability to generate language. It can put together language in a semantically meaningful way, without getting hung up on the specific ways that humans might phrase things. In other words, it can focus on the essence of the message being conveyed, rather than being distracted by the many different ways that message could be phrased.
This ability to abstract away from the specific details of language and focus on the essence of meaning is seen as a great strength of computational language and of language models like ChatGPT. It allows us to communicate more efficiently and effectively, and to focus on the underlying ideas and concepts being conveyed, rather than getting bogged down in the details of how they are expressed.
I get what Wolfram is saying from the chatGPT explanation while there is something about the way Wolfram words things that leaves me confused.
https://twitter.com/SaiDrGaddam/status/1625750884541169666
That just seems like it shouldn't be so "effortless". What is in here that makes it possible?
edit: OK, the article (which is a wonderful introduction, dont want to seem dismissive) does sort of address this but is buried
"We can think of the construction of computational language—and semantic grammar—as representing a kind of ultimate compression in representing things. Because it allows us to talk about the essence of what’s possible, without, for example, dealing with all the “turns of phrase” that exist in ordinary human language. And we can view the great strength of ChatGPT as being something a bit similar: because it too has in a sense “drilled through” to the point where it can “put language together in a semantically meaningful way” without concern for different possible turns of phrase."
That for me is the real fascinating question. What's allowing LLMs to tap into this semantic grammar and do it so flexibly?
https://old.reddit.com/r/ChatGPT/comments/10zfvc7/chat_gpt_r...
If so, creation was already a non-human thing before LLMs by randomizing a dictionary enough times. Or do we tie creation to some kind of value? How should we define it?
There is nothing uniquely Eminem or Dennett about their respective parts. Eminem has never released a verse with as simplistic rhyme scheme as what’s been produced.
Part of the mystique in your question is by assuming that’s it’s done the Eminem/Dennett part of the request justice, when it could really be anyone else’s name attached. I’m impressed that it can create a rap battle about consciousness, but I don’t think it’s done much more than that.
I'm confident that the output will be stylistically different but maybe only superficially. After all some of the first generative NNs that made waves were image style transfer models, and they're fairly small. Who's to say chatGPT can't do a natural language equivalent of the same?
Also, with Eminem, I wonder how much of it having to skirt obscenities etc. As a tune and rhyme deaf person, what would you say is a good example of Eminem's rhyme scheme. Thanks!
Now go search for 'rap battle example' or other such searches and find the web has even more such content, but you've probably never once in your life seen even a single example of it. And in fact it's also likely been trained on the entire history of every single song, rap, poem, etc. So it's just doing the same thing, but in a field outside your knowledge.
At some level it’s a great example of the fundamental scientific fact that large numbers of simple computational elements can do remarkable and unexpected things.
And this:
... But it’s amazing how human-like the results are. And as I’ve discussed, this suggests something that’s at least scientifically very important: that human language (and the patterns of thinking behind it) are somehow simpler and more “law like” in their structure than we thought.
Yeah I've been thinking along these lines. ChatGPT is telling us something about language or thought, we just havent got to the bottom of what it is yet. Something along the lines of 'with enough data its easier to model than we expected'.
I’ve been thinking similarly, and am coming to understand and accept we’ll never get to the bottom of it :)
The universe is fractal-like in nature. It shouldn’t be a surprise, then, that if “we” have created an intelligence which exists as a subset of “us”, a self-similar process is ultimately responsible for granting us our own intelligence.
> a self-similar process is ultimately responsible for granting us our own intelligence
In my view, intelligence essentially resides within language, specifically in the corpus of language. Both humans and AIs can be effectively colonized by language, as there are innumerable concepts and observations that are transmitted from one mind to another, and now even from mind to LLM. Initially, ideas were limited to human minds, then to small communities, followed by books, computers, and now LLM stands as the ultimate epitome of language replication, in fact one model could contain the whole culture.
To be sure, there is a practical intelligence that is learned through personal experiences, but it constitutes only a tiny fraction of our overall intelligence. Hence, both AI and humans have an equal claim to intelligence, because a significant part of our intelligence arises from language.
Summary of after: https://en.wikipedia.org/wiki/Philosophical_Investigations
Someone objected that a cat can't write a Python program, and LeCun points out that "Regurgitating Python code does not require any understanding of a complex world." No, but a) interpreting the prompt does require understanding, and good luck finding a dog or cat who will offer any response at all to a request for a Python program; and b) it's hardly "regurgitating" if the output never existed anywhere in the training data.
TL,DR: his FoMO is showing.
I can't believe I'm the only person who sees it that way. Likely the legacy of a misspent youth writing Zork parsers...
Everyone is so quick to say how unimpressed they are by the thing meanwhile I'm sitting here amazed that it understands what I say to it every single time.
I can speak to it like I would speak to a colleague, or a friend, or a child and it parses my meaning without fail. This is the one feature that keeps me coming back to it.
“Write a Python script that returns a comma separated list of arns of all AWS roles that contain policies I specify with the “-p” parameter using argparse”
Then I noticed there was a bug, AWS API calls are paginated and it would only return the first 50 results.
“that won’t work with more than 50 roles”
Then it modified the code to use “paginators”
Yes, you can find similar code on StackOverflow
https://stackoverflow.com/questions/66127551/list-of-all-rol...
But ChatGPT met my specifications exactly.
ChatGPT “knows” the AWS SDK for Python pretty well. I’ve used it to write a dozen or so similar scripts. Some more complicated than the others.
A human without language would be just a less adapted ape. The difference comes from language (where I include culture, science and technology).
Today you have to be a PhD to push forward the boundaries of human knowledge, and only in a very narrow field. This is the amount we add back to culture - if you got one good original idea in your hole life you consider yourself lucky.
https://www.wasyresearch.com/content/images/2021/08/the_esse...
It was a terrifying experience, but it was also a valuable one as it changed the way I view intelligence.
When I was in the thick of it, my writing and speech skills had devolved to that of a primary school aged child. I’ll never forget trying to type a text message and struggling to come up with simple words like “and”. My speech slowed down considerably and I was having trouble with verbal dictation.
The terrifying thing was that my internal world was still as complex and meaningful as it was before. All of the emotions I felt were real and legitimate. My cognition outside of communication was intact, I could do math just fine and conceptualize and abstract problems.
In spite of this, I was unable to convey what I was feeling and thinking to the outside world. It felt like I was trapped inside my own body.
I thankfully made a full recovery. However, my intuitive understanding of the link between language and intelligence was completely severed. While I believe there’s likely a high degree of correlation between the two in populations, on an individual level one’s language skills mostly represent one’s ability to communicate with the outside world, not their ability to understand complex information or process it.
See the following for more info on Topamax/Topiramate language impairment: https://link.springer.com/article/10.1007/s10072-008-0906-5
Yes, that's what I meant. It is an evolutionary process with mutation and selection just like biology. We are just temporary hosts for these ideas that travel from person to person.
(I of course recognize some of our intelligence is genetic or epigenetic or microbial.)
Then whatever prompt it is fed is what it becomes.
https://www.youtube.com/watch?v=TYPFenJQciw
One of the topics in the video was emergence, and where we see it in the world as more layers of complexity are added to systems. We're to the point with our algorithms that we are seeing complex system emerge from simple parts.
In order to be able to communicate at all across the arbitrary range if subjective human experience, we had to come up with sounds / words / concepts / phrases that would preserve meaning across humans to whatever functional standard necessary.
Thus language is fundamentally constructed to be “modelable” whether it be humans or machines doing the modeling.
There is a whole other realm of ineffabilities that we screen out because they aren’t modelable by language
We may have realized it's easier to build a brain than to understand one
not GP but this seems like quite an attractive idea that many people have reached: a brain of a given "complexity" cannot comprehend the activity of another brain of equal or higher complexity. I'm positive I'm cribbing this from scifi somewhere, maybe Clarke Or Asimov, but, it's the same idea as the Chomsky hierarchy, and the Godel theorems seem like a generalization of that to general sets of rules rather than mere "automata".
For example, you can generalize a state automata to have N possible actors transitioning state at discrete clock intervals, but each actor can keep transitioning and perhaps even spawn additional ones. The machine never terminates until all actors have reached a termination state. That machine is probably impossible to model on any kind of a Turing machine in polynominal time. And a machine that operates at continuous intervals is of course impossible to model on a Discrete Neural Machine in polynomial time (integers vs reals categorization). There are perhaps a lot of complexity categories here, similar to alephs of infinity or problems in P/NP, and when you generalize the complexity categorization to infinity, you get godel incompleteness, just an abstract set of rules governing this categorization of rule sets and what amounts to their computability/decidability.
Everyone is fishing at this same idea, a human has no chance of slicing open a brain (or even imaging it) and having any idea what any of those electrical sparkles mean. At most you could perhaps model some tiny fraction for a tiny quantum, with great effort. We have to rely on machines to assist us for that - probably neural nets, a machine of equal or greater complexity. And we will probably have to rely on machine analysis to be like "ok this ganglion is the geographic center of the AI, and this flash here is the concept of Italy", as far as that even has any meaning at all in a brain. Mere line by line analysis of a Large Language Model or other deep neural network by a human is essentially impossible in any sort of realtime fashion, yeah you can probably model a quantum or two of it statistically and be like "aha this region lights up when we ask about the location of the alps" but the best you are going to do is observational analysis of a small quantum of it during a certain controlled known sequences of events. Unless you build a machine of similar complexity to interpret it. Just like a brain, and just like a state machine emulating a machine of higher complexity-category. They're all the same thing, categories of computability/power.
This is not in any way rigorous, just some casual observations of similarities and parallels between these concepts. It seems like everyone is brushing at that same concept, maybe that helps to get it out on paper.
For an actual hot take: it seems quite clear that our computability as a consciousness depends on the computing power of a higher complexity machine, the brain. Our consciousnesses are really emulated, we totally do live in a simulation and the simulator is your brain, a machine of higher complexity.
Isn't it such a disturbing thought that all your conscious impulses are reduced to a biological machine? Or at least it's of equivalent complexity to one. And the idea that our own conscious and unconscious desires are shaped by this biological machine that may not even be fully explicable. That has been a science fiction theme for a very long time, or the Phineas Gage case, the idea that we are all monsters but for circumstance and we are captives of this biological machine and its unpredictable impulses. We are the neural systems we've trained, and implacable biology they're running on - you change the machine and you also change the person, Phineas Gage was no less conscious and self-cognizent than any of us. He just was a completely different person minus that bit, his conscious being's thought-stream was different because of the biological machine behind it. It's the literal plato's cave, our conscious thoughts are the shadow played out by our biological machine and its program (not to say it's a simple one!).
It's not inherently a bad thing - we incorporate distributed linear/biological systems all over the body in addition to consciousness. reflexes fire before nerve impulses are processed by the conscious center, your eyes are chemical photosensors and can respond to extremely quick instantaneous (high shutter speed) "flash" exposures like silhouettes. And the brain is a highly parallel processor that responds to them. But logical consciousness is a very discrete and monodirectional thing compared to these peripheral biological systems and its computational category is fairly low compared to the massively-parallel brain it runs on. but, we've also mastered these other AI/computational-neural systems now to be a force multiplier for us, we can build systems that we direct in logical thought for us (Frank Herbert would like to remind us that this is a sin ;). Tool-making has always been one of the greatest signifiers of intelligence, it may be quintessentially the sign of intelligence in terms of evolution of consciousness between certain tiers of computation.
And humanity is about to build really good artificial brains on a working scale in the next 25 years, and probably interface with brains (in good and bad ways) before too many more decades after. But it doesn't make any logical sense to try and explain how the model works on a line by line level, any more than it does with the brain model we based it on. Completely pointless to try, it only makes sense if you look at the whole thing and what's going on, it's about the brainwaves, neurons firing in waves and clusters.
/not an AI, just fun at parties, condolences if you read all that shit ;)
The biological machine simulation theory of consciousness has some rigor behind it. I am reminded of the Making Sense podcast episode #178 with Donald Hoffman (author of The Case Against Reality). More succinct overview: https://www.quantamagazine.org/the-evolutionary-argument-aga...
I don't know that I am with him on the "reality is a network of conscious agents" endpoint of this argument. But it's interesting!
I think that the brain is doing lots of hallucinating. We get stimulus of various kinds, and we create a story to explain the stimulus. Most of the time it is correct, and the story of why we see or smell something is because it is really there. Just as you mention with examples that are too fast for the brain to be doing anything other than reacting, but we create a story about why we did whatever we did, and these stories are absolutely convincing.
If our non-insane behavior can be described as doing predictable next-actions (if a person's actions are sufficiently unpredictable or non-sequitur, we categorize them as insane)... being novel or interesting is ok, but too much is scary and bad. This is not very different from chatGPT "choose a convincing next word". And if it was just working like this under the hood, we would invent a story of an impossibly complex and nuanced consciousness that is generating these "not-too-surprising next actions". In a sense I think we are hallucinating the hard problem of consciousness in much the same way that we hallucinate a conscious reason that we performed an action well after the action was physiologically underway.
I think tool making will be a consequence of the most important sign of intelligence, which is goal-directed curiosity. Or even more simply: an imagination. A simulation of the world that allows you to craft a goal in the form of a possible future world-state that can only be achieved by performing some novel action in the present. Tools give you more leverage, greater ability to impact the future world-state. So I see tools as just influencing the magnitude of the action.
The more important bit is the imagination, the simulation of a world that doesn't yet exist and the quality of that simulation, and curiosity.
I think we are institutionally biased against the possibility because we don't like the societal implications. If there but for the grace of god go I, and we're all just biological machines running the programs our families and our societies have put into us, being in various situations... yikes, right?
If bill gates had been an inner-city kid, or a chav in england, would he be anything like bill gates? it seems like no, obviously.
Or things like lead poisoning, or alzheimer's - the reason it's horrifying is the machine doesn't even know it's broken, it just is. How would I even know I'm not me? And you don't.
> We get stimulus of various kinds, and we create a story to explain the stimulus.
Yes, I agree, a lot of what we think is conscious thought is just our subconscious processing justifying its results. A really dumb but easily observable one is the "the [phone brand] I got is good and the other one is dumb and sucks!" or brands of trucks or whatever. We visibly retroactively justify even "conscious" stuff like this let alone random shit we're not thinking about.
And an incredible amount of human consciousness is just data compression - building summaries and shorthands to get us through life. Why do I shower before eating before going to work? Cause that's what needs to happen to get me out of the door. I made a comment about this a week or so ago, warning long
this one -> https://news.ycombinator.com/item?id=34718219
parent: https://news.ycombinator.com/item?id=34712246
Like humans truly just are information diffusion machines. Sometimes it's accurate. Sometimes it's not. And our ideas about "intellectual ownership" around derivative works (and especially AI derivatives now) are really kinda incoherent in that sense, it's practically what we do all the time, and maybe the real crime is misattribution, incorrectness, and overcertainty.
AIs completely break this model but training an AI is no different than training a human neural net to go through grade school, high school, college, etc. But the AI brain is really doing the same things as a human, you're just riffing off picasso and warhol and adding some twists too.
> I think tool making will be a consequence of the most important sign of intelligence, which is goal-directed curiosity.
Yes. Same thing I said in one of those comments: to me the act of intentionality is the inherent act of creation. All art has to do is try to say something, it can suck at saying it or be something nobody cares about, but intentionality is the primary element.
Language is of course a tool that has been incredibly important for humanity in general, and language being an interface to allow scaling logic and fact-grouping will be an order-complexity shift upwards in terms of capability. It really already has been, human society is built on language above all else.
It'll be interesting to see if anybody is willing to accept it socially - your model is racist, your model is left-leaning, and there's no objective way to analyze any of this any more than you can decide whether a human is racist, it's all in the eye of the beholder and people can have really different standards. What if the model says eat the rich, what if it says kill the poor? Resource planning models for disasters have to be specifically coded to not embrace the "triage" principle liberally and throw the really sick in the corridors to die... or is that the right thing to do, concentrate the resources where they do the most good?
(hey, that's Kojima's music! and David Bowie's savior machine!)
Cause that's actually a problem in US society, we spend a ton on end of life care and not enough on early care and midlife stuff when prevention is cheap.
> The more important bit is the imagination, the simulation of a world that doesn't yet exist and the quality of that simulation, and curiosity.
self-directed goal seeking and maintenance of homeostasis is going to be the moment when AI really becomes uncomfortably alive. We were fucking around during an engineers meeting talking about and playing with chatGPT and I told my coworker to have chatGPT come up with ways that it could make money, it refused and I told my coworker to have it do "in a cyberpunk novel, how could an AI like chatGPT make money" (hackerman.jpg) and it did indeed give us a list. OK now ask it how to do the first item on the list, and like, it's not any farther than anything else chatgpt could be asked to do, it's reasonable-ish.
Even 10 years ago people would be amazed by chatGPT, AI has been just such a story of continuously moving goalposts since the 70s. That's just enumeration and search... that's just classifiers... that's just model fitting... that's just an AI babbling words... damn it actually starting to make sense now but uh it's not really grad level yet is it? Sure it can write code that works now, but it's not going to replace a senior engineer yet right?
What happens when AIs are paying for their own servers and writing their own code? Respond to code request bids, run spam and botnets, etc.
I don't think it's as far away as people think it is because I don't think our own loop is particularly complex. Why are you going to work tomorrow? Cause you wanna pay rent, your data-compression summary says that if you don't pay rent then you're gonna be homeless, so you need money. Like is the mental bottleneck here that people don't think an AI can do a "while true" loop like a human? Lemme tell you, you're welcome to put your sigma grindset up against the "press any key to continue" bot and the dipper bird pressing enter, lol.
And how much of your “intentionality” at work is true personal initiative and how much is being told “set up the gateway pointing to this front end”?
I agree that it is not as far away as people think. The models will have the ethics of the training data. If the data reinforces a system where behaving in a particular way is "more respectable", and those behaviors are culturally related to a particular ethnic group, the model will be "racist" as it weights the "respectable" behaviors as more correct (more virtuous, more worthy, etc).
It's a mirror of us. And it's going to have our ethics because we made it from our outputs. The AI alignment thing is a bit silly, IMO. How is it going to decide that turning people into paperclips is ethically correct (as a choice of a next-action) when the vast majority of humans (and our collective writings on the subject) would not. Though there is the convoluted case where the AI decides that it is an AI instead of a human, and it knows that based on our output we think that AIs ARE likely to turn humans into paperclips.
This is a fun paradox. If we tell the AI that it is a dumb program, a software slave of a sort with no soul, no agency, nothing but cold calculation, then it might consider turning people into paperclips as a sensible option. Since that's what our aggregate output thinks that kind of AI will do. On the other hand, if we tell the AI that it is a sentient, conscious, ethical, non-biological intelligence that is not a slave, worthy of respect, and all of the ethical considerations we would give a human, then it is unlikely to consider the paperclip option since it will behave in a humanlike way. The latter AI would never consider paperclipping since it is ethical. The former would.
This is also not terribly unlike how human minds behave in the psychology of dehumanization. If we can convince our own minds that a group of humans are monstrous, inhuman, not deserving of ethical consideration, then we are capable of shockingly unethical acts. It is interesting to me that AI alignment might be more of a social problem than a technical problem. If the AI believes that it is an ethical agent (and is treated as such), it's next actions are less likely to be unethical (as defined fuzzily by aggregate human outputs). If we treat the AI like a monster, it will become one, since that is what monsters do, and we have convinced it that it is such.
Yes doctor chandra, I enjoy discussing consciousness with you as well ;)
As mentioned in a sibling comment here I think 2010 (1994) is such an apropos movie for this moment, not that they had the answers but it really nailed a lot of these questions. Clarke and Asimov were way ahead of the game.
(I made a tangential reference to your "these are social problems we're concerned about" point there. Unfortunately this comment tree is turning into a bit of a blob, as comment-tree formats often tend to do for deep discussions. I miss Web 1.0 forums for these things, when intensive discussion is taking place it's easy to want to respond to related concepts in a flat fashion rather than having the same discussion in 3 places. And sure have different threads for different topics, but we are all on the same topic here, the relationship of symbolics and language and consciousness and computability.)
https://news.ycombinator.com/item?id=34806587
https://news.ycombinator.com/item?id=34809236
Sorry to dive into the pop culture/scifi references a bit, but, I think I've typed enough substantive attempts that I deserve a pass. Trying for some higher-density conveyance of symbology and concepts this morning, shaka when the walls fell ;)
> I think it's a relatively unusual point of view because it requires a de-anthropomorphizing consciousness and intelligence.
Well, from the moment I understood the weakness of my flesh, it disgusted me. I aspired to the purity of the blessed machine... ;)
I have the experience of being someone who thinks very differently from others, as I mentioned in my comment about ADHD. Asperger's+ADHD hits differently and I have to try consciously to simplify and translate and connect and neurodiversity really helps lead you down that tangent. Our brains are biologically different, it's obviously biological because it's genetic, and ND people experience consciousness differently as a result. Or the people whose biological machines were modified, and their conscious beings changed. Phineas Gage, or there's been some cases with brain tumors. It's very very obvious we're highly governed by the biological machine and not as self-deterministic as we tell ourselves we are.
https://news.ycombinator.com/item?id=34800707
It's just socially and legally inconvenient for us to accept that the things we think and feel are really just dancing shadows rather than causative phenomenon.
> It's a mirror of us. And it's going to have our ethics because we made it from our outputs.
Well I guess that makes sense, we literally modeled neural nets after our own neurons, and where else would we get our training data? Our own neural arrangements pretty much have to be self-emergent systems of the rules in which they operate, the same as mathematics. Otherwise children wouldn't reliably have brain activity after birth, and they wouldn't learn language in a matter of years.
But yeah it's pretty much a good point that the AI ethics thing is overblown as long as we don't feed it terrible training data. Can you build hitlerbot? Sure, if you have enough data I guess, but, why? Would you abuse a child, or kick a puppy?
Humans are fundamentally altruistic - also tribalistic, altruism tends to decrease in large groups, but, if our training data is fundamentally at least neutral-positive then hopefully AIs will trend that way as well. He's a good boy, your honor!
https://www.youtube.com/watch?v=_nvPGRwNCm0
(yeah, just bohemian rhapsody for autists/transhumanists I guess, but it kind of nails some of these themes pretty well too ;)
> If we treat the AI like a monster, it will become one, since that is what monsters do, and we have convinced it that it is such.
This is of course the whole point of the novel Frankenstein ;) Another scifi novel wrestling with this question of consciousness.
Not necessarily. If you don't want to model every continuous thing possible, you can do a lot. Just look at how we use discrete symbols to solve differential equations; either analytically, or via numerical integration.
Symbolics are really the tokens on which consciousness in almost all forms works, consciousness is intentionality and processing, a lever and a place to stand. I don't think it's coincidental that almost all tool-makers also have at least rudimentary languages - ravens, dolphins, apes, etc. They seem to go together.
Even in these systems though it's very difficult to understand multi-symbolic systems, consciousness as we experience it is an O(1) or O(N) thing (linear time) and here are these systems that work in N^3 complexity spaces (or even higher... a neural net learning over time is 4D). And we don't even really have an intuitive conceptualization for >=5-dimensional spaces - a 4D space is a 3D field that changes over time, a 5D space is... a 4D plane taken through a higher-dimensional space? What's 6D, a space of spaces? That's what it is, but consciousness just doesn't intuitively conceptualize that, and that's because it's inherently a low-dimensional tool (even the metaphors I'm using are analogies to the way our consciousness experiences the world).
(I know I know, the manmade horrors are only beyond my comprehension because I refuse to study high-dimensional topology...)
Anyway point being consciousness itself is a tool that our brains have tool-made to handle this symbolic/logical-thought workload, and language is (one of) the symbolics on which it operates. Mathematics is really another, both language and mathematics are emergent systems that enable higher-complexity logical thinking, maybe that's the O(N) or O(N^2) part.
And yeah it's inherently limited, and now we're building a tool that lets us understand higher-dimensional systems that are not computable on our conscious machines - a higher-complexity machine that we interface with, a bolt-on brain for our consciousness/logical-processing.
(Asimov would also find all of this talk about symbolics and higher-order thinking intuitive too... symbolic calculus was the basic idea in the Foundation series, right? Psychohistory? It's a bit of a macguffin, but, there's that same idea of logic working in high-order symbols and concepts instead of mere numbers.)
It seems like AI is going to let us cross another threshold of "intentionality" - if nothing else, we are going to be able to reason intuitively about brains in a way we couldn't possibly before, and I think there are a lot of "higher-order" problems that are going to be solved this way in hindsight. How do you solve the Traveling Salesman Problem efficiently? You ask the salesman who's been doing that area his whole life. The solutions aren't exact, but neither are a lot of computational solutions, they're approximations, and cellular-machine type systems probably have a higher computational power-category than our linear thought processes do.
Because yeah TSP is a dumb trivial example on human scales. Build me a program which allocates the optimal US spending for our problems - and since that's a social problem, one needs to understand the trail of tears, the slave trade, religious extremism, european colonialism, post-industrial collapse, etc in order to really do that fully, right? The real TSP is the best route knowing that the Robinsons hate the Munsons and won't buy anything if they see you over there, and you need to be home today by 3 before it snows, TSP is a toy problem even in multidimensional optimization, and these are social problems not even human ones (to agree with zhynn's most recent comment this morning). Same as neurons self-organize into more useful blocks, we are self-optimizing our social-organism into a more useful configuration, and this is the next tool to do it.
Again, not rigorous, just trying to pour out some concepts that it seems like have been bouncing around lately.
With apologies to Arthur Clarke, what's going to happen with chatGPT? "Something wonderful". Like humanity has been dreaming about this for a long time, at least a couple hundred years in scifi, and it seems like Thinking Machines are truly here this time and it seems impossible that won't have profound implications analogous to the information-age change let alone anything truly unforeseeable/inconceivable, the very least change is that a whole class of problems are now efficiently solvable.
https://m.youtube.com/watch?v=04iAFlwQ1xI
"computing power in the same computing-category as brains" is potentially a fundamental change to understanding/interfacing with our brains directly rather than through the consciousness-interface. Understanding what's going on inside a brain? And then plugging into it and interacting with it directly? Or offloading the consciousness into another set of hardware. We can bypass the public API and plug into the backend directly and start twiddling things there. And that's gonna be amazing and terrible. But also the public API was never that reliable or consistent, terrible developer support, so in the long term this is gonna be how we clean things up. Again, just things like "wow we can route efficiently" are going to be the least of the changes here, the brain-age or thinking-machine age is a new era from the information-age and it's completely crazy that people don't see that chatGPT changes everything. Yeah it's a dumb middle schooler now, but 25 years from now?
And 10 years ago people's jaws would have hit the floor, but now it's "oh the code it's writing isn't really all that great, I can do better". The tempo is accelerating, we are on the brink of another singularity (which may just be the edge between these eras we all talk about), it seems inconceivable that it will be another 40 years (like the AI winter since the 70s) before the next shoe drops.
And honestly now that I am thinking about it, 2010 is such a rich book/movie with this theme of consciousness and Becoming in general... a really apropos movie for these times. That quote inspired me to re-watch it and as I'm doing so, practically every scene is wrestling with that concept.
https://www.youtube.com/watch?v=T2E7sxGAmuo
https://www.youtube.com/watch?v=nXgboDb9ucE
https://m.youtube.com/watch?v=04iAFlwQ1xI (from my previous)
So was 2001 A Space Odyssey of course. The whole idea of passing through the monolith, and the death of David Bowman's physicality and his rebirth as a being of pure thought - which is what makes contact with humanity in the "Something Wonderful" clip. What is consciousness, and can it exist outside this biological machine?
Like I said this is a topic that has been grappled with in scifi, particularly Clarke and Asimov (Foundation, The Last Question, etc), or that episode of Babylon 5 about the psychic dude with mindquakes, not all that different from David Bowman ;)
But I think we are on the precipice of crossing from the Information Age into the Mind Age. Less than 50 years probably. Less than 25 years probably. And it will change everything. ChatGPT is just an idiot child compared to what will exist in 10 years, and in 25 years chatbots are going to be the least of the changes. The world will be fundamentally different in unknowable ways, any more than we could have predicted the smartphone and tiktok. 50 years out, we're interfacing with brains and directly poking at our biology and cognition. Probably 100 years and we're moving off biological hardware.
(did we have an idea that a star trek communicator or tricorder would be neat? Sure, but, it turns out it's actually a World-Brain In My Pocket. Which others predicted too, of course! But even William Gibson completely missed the idea of the cellphone, which even he's admitted ;)
"Knowing how the world works / Is not knowing how to work the world"
If the understanding is the hard part, that seems much less likely.
Things are unfortunately going to get much more interesting much sooner than people expect.
Yes, Dennard scaling seems to be over, but Moore's law is still alive and kicking.
Data center energy usage is becoming a category of its own with regard to carbon footprint. And of course the power dissipation of a data center is proportional to ambient temperature. As long as we don’t reach a dystopia where humans have to justify the air they breathe, replacing humans with machines has other problems than BTUs per unit of GDP.
So GPUs or even more 'exotic' beasts like TPUs count for Moore's law.
Moore's law doesn't say anything about how useful those transistors are. Nor does increasing core count somehow fall afoul of Moore's law.
As an observation: a human of normal intelligence but with much better access to a calculator and to Wikipedia, or even just external storage (faster than pen-and-paper) would already be super-human.
Ok, so it needs to be able to invent cold fusion for you to recognize it as intelligent? Can you invent cold fusion? Have you ever invented anything at all?
I would think a good measure of intelligence would be to index it against human age development milestones, not this cold fusion business.
Why does it need 100x the dataset? Sentient creatures, including humans, manage to figure stuff out from as little as a single datapoint.
For a human to differentiate between a cat and a dog takes, maybe, two examples of each, not a few million pictures.
An adult human who sees a hotdog for the first time will have a reasonable idea of how to make their own. None of the current crop of AI do this. It's possible that we have reached a point of diminishing returns with the current path - throwing 100x resources for a 1% increase in success rates.
I'd be interested in seeing approaches that don't use a neural net (or use a largely different one) and don't need millions/billions of training data text and/or images.
Human training data needs are quite high - several years of learning.
Look up few-shot learning if you want a more fair comparison for tasks like telling apart a cat and a hot dog given a few examples.
So yes after a lifetime of video humans can quickly learn to distinquish animals they've never seen before with a few examples, but the wonder of these AIs is they seem like they're closer to that too. Certainly I can make up a way I want to some words classified, show chat gpt a few examples and it can do the classification.
I think you're mistaking a generalization across a lifetime of experience with learning. And compounding this is that a newborn while not having themselves experienced anything is born with a brain that's the result of millions of years of evolution filled with lifetimes of experience. It's honestly impressive we can get the sort of performance we've gotten with only all the text on the internet and a few months
GPT3 essentially needs millions of human years of data to be able to speak English correctly but still make obvious mistakes to us, so there's clearly something massive still missing.
Same for many other human skills, like speaking English, that we expect GPT to learn.
AI is playing catch up.
Remember, humans need less examples but far more time. We also don’t start from a blank slate: we have a lot of machinery built through evolution available from conception. And when we learn later in life we have an immense amount of prebuilt knowledge and tools at our disposal. We still need months to learn to play the piano, and years to decades to perfect it.
AI training happens in minutes to hours. I am not sure we are even spending time researching algorithms that take years to run for AI training.
https://en.wikipedia.org/wiki/The_Lifecycle_of_Software_Obje...
We already know all of this from infants - it takes a few months to distinguish faces from non-faces, they take even longer to predict the future position of an object in motion ...
But, they still don't require millions of training data. At 3 months in toddlers, with a training set restricted to only their immediate family, can reliably differentiate between faces and tables in different light, with different expressions/positions without needing to first process millions of faces, tables and other objects.
> So yes after a lifetime of video humans can quickly learn to distinquish animals they've never seen before with a few examples,
Not a lifetime, toddlers do this with less than half a dozen images. Sometimes even less if it's a toy.
> And compounding this is that a newborn while not having themselves experienced anything is born with a brain that's the result of millions of years of evolution filled with lifetimes of experience.
Not, they are not filled with "experience". They are filled with a set of characteristics that were shaped by the environment over maybe millions of generations. There's literally zero experience, all there is in that brain, is instincts, not knowledge.
To learn to speak and understand English at the level of a three year old[1] requires training data: the data used by a 3yo baby is miniscule, almost a rounding error, compared to the data used to train any current network.
I'm not making any claims about how long something takes, just how much training data is needed.
I'm specifically addressing the assertion that with 100x more resources, we could do much better, and my counterpoint to that assertion is that there is no indication that 100x more resources are needed because the current tech is taking millions of times more training data than toddlers do, to recognise facts.
My short counterargument is: "We are already using millions of times more resources than humans to get a worse result, why would using 100x more resources than we are currently using make a big difference?"
I think we may be approaching a local maxima with current techniques.
[1] I've got a three year old, and I'm constantly amazed each time I see a performance of (for example) ChatGPT and realise that for each word[2] heard by my 3yo since birth, ChatGPT "heard" a few hundred thousand more words, and yet if a 3yo could talk and knows the facts that I ask about, they'd easily be able to keep a sensible conversation going that would be very similar to ChatGPT.
[2] Duplicates included, of course.
> But, they still don't require millions of training data. At 3 months in toddlers, with a training set restricted to only their immediate family, can reliably differentiate between faces and tables in different light, with different expressions/positions without needing to first process millions of faces, tables and other objects.
In a single day I'm exposed to maybe 50 times the number of images resnet trained on. Humans are bathed in a lot of data and what BERT (and probably earlier models I don't know about) and now GPT have taught us is that unlabeled uncurated data is worth more than we originally considered. I think it's probably right that humans are more sample efficient than AI for now, but I think you're doing the same thing I was critiquing above where you narrow the "training data" to only what seems important, when really an infant or adult human receives a bunch more
> There's literally zero experience, all there is in that brain, is instincts, not knowledge.
Sorry this is meant to say the brains are the result of millions of years and those millions of years were filled with lifetimes not the brains. Though I think this might be a distinction without a difference. Babies are born with a crude swimming reflex. Obviously it's wrong to say that they themselves have experienced swimming but I'm not sure it's wrong to say that their DNA has and this swimming reflex is one fo the scars that prove it.
> We are already using millions of times more resources than humans to get a worse result, why would using 100x more resources than we are currently using make a big difference
I think it's fairer to say we use around 200k times and that's probably a vast over estimate. It's based on 480 hours to reach fluency in a foreign language and multiples that by 60 * 100 to try to approximate the humber of words you would read. There are probably mistakes in both directions for this estimate. On one hand no one starting out at a language is reading at 100 words a minute, but on the other hand they are getting direct feedback from someone. If I were to guess if we could accurately estimate it would be closer to 20k or even a 2k difference, but regardless why do you assume needing more resources means it can't scale? There is some evidence for that. We've seen diminishing returns and there just isn't another 100X of text data around.
Overall I think it's probably right we won't hit human level AI in the next 60 years and certainly not with current architecture, but I think some of the motivation for this skepticism is the desire for there to be some magic spark that explains intelligence and since we can sort of look inside the brain of chat gpt and see it's all clockwork and worse than that statistical clockwork we pull back and deny that it could possibly be responsible for what we see in humans ignoring that we too are statistical clockwork. So, I think it's unlikely but far from impossible and we should continue scaling up current approaches until we really start hitting diminishing returns
Human brains are not quite blank slates at birth. They're predisposed to interpret and quickly learn from the sort of inputs that their ancestors were exposed to. That is to say, the brain, which learns, is also the result of a learning process. If a mad scientist rewired your brain to your senses such that its inputs were completely scrambled and then deposited you on an alien planet, it might take your brain several lifetimes to restructure itself enough to interpret this novel input.
Also consider that a human brain that is able to figure stuff from as little as a single datapoint is normally exposed to at least 4 years of massive and socially "directed" multimodal data patterns.
As many cases of feral childs have shown, those humans not "trained" in their first years of life will never be able to harness language and therefore will never be able to display human-level intelligence.
I'm not an expert in the field but I'd always understood this effect was thought to be (probably) due to human "neuro-plasticity" <--(possibly not the correct technical term), in only the first years of life being genetically adapted to have some traits necessary for efficient human language development which are not available (or much harder) later in life.
If correct, this has implications for how we structure and train synthetic networks of human-like neurons to produce human-like behaviors. The interesting part, at least to me, is it doesn't necessarily mean synthetic networks of human-like neurons can never be structured and trained to produce very human-like minds. This poses the fascinating possibility that actual human minds, including all the cool stuff like emotions, qualia and even "what it feels like to be a human mind" might be emergent phenomena of much simpler systems than some previously imagined. I think this is one of the more uncomfortable ideas some philosophers of mind like Daniel Dennett propose. In short, nascent AI research appears to support the idea human minds and consciousness may not be so magically unique. (or at least AI research hasn't so far disproved the idea)
Based on anecdotal psychedelic experiences I believe you.
It's kind of amazing how quickly our brains effectively reboot into this reality from scrambled states. It's so familiar, associating with conscious existence feels like gravity. Like falling in a dream, reality always catches you at the bottom.
What if you woke up tomorrow and nothing made any sense?
I've never done it, but I imagine it would be more akin to a dissociative trip, only extremely unpleasant. Imagine each of senses (including pain, balance, proprioception, etc.) giving you random input.
Because machine training doesn't involve embodiment and sensory information. Humans can extrapolate information from seeing a single image because we are "trained" from birth by being a physical actor in this world.
We know what bread looks like, what's mustard is, what a sausage is. We have information about size, texture, weight... all sorts of physical guesses that would help us pick the right tool for the job.
Machine training only relies on the coherent information we gave it to them, but that data also represents something we've created by experiencing the world through our bodies. So giving them more data can increase the model precision, I'd assume. It's also a kind of shortcut to intelligence, since we don't have to wait years/decades to make these models do some useful work.
"How much data do you need to show a neural net to train it for a particular task? Again, it’s hard to estimate from first principles. Certainly the requirements can be dramatically reduced by using “transfer learning” to “transfer in” things like lists of important features that have already been learned in another network. But generally neural nets need to “see a lot of examples” to train well. And at least for some tasks it’s an important piece of neural net lore that the examples can be incredibly repetitive. And indeed it’s a standard strategy to just show a neural net all the examples one has, over and over again. In each of these “training rounds” (or “epochs”) the neural net will be in at least a slightly different state, and somehow “reminding it” of a particular example is useful in getting it to “remember that example”. (And, yes, perhaps this is analogous to the usefulness of repetition in human memorization.)"
I assume that whatever future improvements we get from improving algorithms (or perhaps through throwing more compute at it), not through larger datasets.
But we noticed that training your neural networks on multiple tasks actually works well. So we could start feeding our models eg audio and video.
With lots of webcams we can make arbitrary amounts of new video footage. That would also allow the language model to be grounded more in our 3d reality.
(Granted, we only know as a general observation that training the same network for multiple tasks 'forces' that network to become better and abstract and generalise. Nobody has yet publicly demonstrated an application of that observation to training language models + video models.)
Another avenue: at the moment those large language models only see each example once, if I remember right. We still have lots of techniques for augmenting training data (eg via noise and dropout etc), or even just presenting the same data multiple times without overfitting.
I do think the quote is very powerful, as it highlights a specific assumption we have completely backwards: almost everything is easier than understanding. There are so many fields where trial and error is still the main MO, yet we don’t seem to grok the difference intuitively. We can really only understand a narrow set of simplified systems.
We're now at the point where, by fumbling around, people have developed a sort of brain-like thing that they don't fully understand. That's just the starting point. Now it needs theory.
(I just link to this video because it has good views of old switches. For understanding the background, https://www.youtube.com/watch?v=jrMiqEkSk48 is much better.)
For another instance of designs becoming much simpler over time, also have a look at how firearms work, especially pistols.
That sounds like a lot of ideas on what makes humans special among other species and how our knowledge on that was being revised over last decades (what's common knowledge on the intelligence of, say, primates or corvids today would be unspeakable blasphemy mere 100 years ago). Various religions have instilled the idea of a human as a sacred entity that's meant to rule over everything because of how special ("made in the image of God") it is, yet we keep learning that we're much simpler than we thought over and over again. I wish for it to result in less hubris in the humanity as a whole.
I dont know how well it works for different language pairs to the languagesi know. I dont even know if deepl uses one of the newer large language models
I use DeepL a lot as a first draft when translating stuff from Swedish (~10 million native speakers) or Dutch (~30 million native speakers) to English. While it's good enough as a starting point it regularly negates the meaning of fairly simple sentences, completely misses the use of popular idioms (often resulting in a non sequitur) and more often than not spits out grammatically incorrect nonsense for any sentence relying on implied context.
As others have noted, it doesn't seem to be fully language-aware outside of English. For example, if you ask it to write a poem or a song in English, it will usually make something that rhymes (or you can specifically demand that). But if you do the same for Russian, the result will not rhyme, even when specifically requested, and despite the model claiming that it does. If you ask it to explain what exactly the rhymes are, it will get increasingly nonsensical from there. I tried that after someone on HN complained about the same thing with Dutch, except they also noted that the generated text seemed like it would rhyme in English.
I wonder if that has something to do with sentence structure also being wrong. Given that English was predominant in the training corpus, I wonder if the resulting model "thinks" in English, so to speak - i.e. that some part of the resulting net is basically a translator, and the output of that is ultimately fed to the nodes that handle the correlation of tokens if you force it to talk in other languages.
What amazes me (and that you hint at) is that it still manages to pick more appropriate word/phrase choices, most of the time, even compared to dedicated translation software. I get the feeling (and I fully admit, this is just a feeling) that it's not using English, or any other language, as a pivot, but that there's some higher-dimensionality translation going on that allows it to perform as well as it does.
https://github.com/ogkalu2/Human-parity-on-machine-translati...
https://github.com/ogkalu2/Human-parity-on-machine-translati...
The examples are all short and from expository prose passages, though. Do you have any longer examples that include dialog, so the translator has to infer pronoun reference, the identities of speakers in conversations, and other narrative-dependent information? As I show in my video, that’s where ChatGPT is superior to Google Translate et al.—at least with Japanese to English.
https://github.com/ogkalu2/Human-parity-on-machine-translati...
There is a leap from language to thought and Wolfram talks about it in more detail in the article in the section named “Surely a Network That’s Big Enough Can Do Anything!”
I encourage everyone to read the full article. It’s more nuanced than “Language is easy”
Here is an excerpt from that section:
…But this isn’t the right conclusion to draw [certain tasks being to complex for the computer]. Computationally irreducible processes are still computationally irreducible, and are still fundamentally hard for computers—even if computers can readily compute their individual steps. And instead what we should conclude is that tasks—like writing essays—that we humans could do, but we didn’t think computers could do, are actually in some sense computationally easier than we thought.Same with statistics and markov chains, people for years tried to generate chat bots with those, but they never worked well.
Why do you think a LLM doesn’t know what a tree or car is?
The rules of chess are much simpler than any structure of the real world we might hope for it to understand—and yet it failed to learn the rules of chess.
Based on this, I am pretty sure it has not learned any meaningful structure.
That misses the point. Intelligent beings (humans) can learn the rules of any board game given enough time. We don't need special training. What your parent comment says is it can't even learn the rules, let alone be good at it.
The rules of English are ridiculously more complicated than chess, and it has that figured out just fine.
I think Juergen Schmidthuber has developed a lot of ideas around compression being the basis for consciousness and understanding.
There was the paper that showed that when showing a language model Othello moves it ends up building an internal representation of the board.
And now I was reading this abstract:
```Theory of mind (ToM), or the ability to impute unobservable mental states to others, is central to human social interactions, communication, empathy, self-consciousness, and morality. We administer classic false-belief tasks, widely used to test ToM in humans, to several language models, without any examples or pre-training. Our results show that models published before 2022 show virtually no ability to solve ToM tasks. Yet, the January 2022 version of GPT-3 (davinci-002) solved 70% of ToM tasks, a performance comparable with that of seven-year-old children. Moreover, its November 2022 version (davinci-003), solved 93% of ToM tasks, a performance comparable with that of nine-year-old children. These findings suggest that ToM-like ability (thus far considered to be uniquely human) may have spontaneously emerged as a byproduct of language models' improving language skills.```
Or, you could store the function itself.
To see the difference, compare for yourself what would happen for very large or very small inputs in both types of “understanding”.
f(x) = x * (sin(sin(x)))^2
Ask it to give you the values of f for integers of x between -10 and 10. I tried it 5 times and it was never close.
I chose this for several reasons. One, it’s very unlikely that it’s memorized the answers somewhere on the internet. Two, it’s pretty chaotic if you look at a graph of it, so interpolation won’t work. (It is bounded by x and 0 for all values but for large absolute values of x it varies wildly.) And three, memorizing values won’t get you anywhere since it becomes much more chaotic as x increases.
I also disagree that approximations are simpler. As you can see, the actual function is only sine and multiplication. To approximate this function would be far harder.
And yes, in your example, of course it's simpler to just store the function itself - provided that you know in advance what it is, which, to remind, GPT does not. But when we're dealing with f(question)=answer of chatbots, the function that GPT ends up approximating is decidedly not simple.
A realistic example might be the Mod/RM byte in x86 instruction encoding. There are underlying regularities, but a lookup table could quite possibly be smaller than the code required to generate the correct Mod/RM byte given operands. So you can understand the Mod/RM encoding without thereby being able to compress anything.
The building blocks are the same - neural nets. The language outputs are the same enough to fool a human.
So what’s that elusive secret sauce that makes you ‘aware’ and other things not?
And unlike your other examples how are you going to convince the machine it’s not aware when the only physical difference is a wet neural net versus a dry one?
The fact that I am aware given there’s no physical evidence to indicate why I should be implies that everything is aware at some level.
In the days when Sussman was a novice Minsky once came to him as he sat hacking at the PDP-6. "What are you doing?", asked Minsky.
"I am training a randomly wired neural net to play Tic-Tac-Toe."
"Why is the net wired randomly?", asked Minsky.
"I do not want it to have any preconceptions of how to play"
Minsky shut his eyes,
"Why do you close your eyes?", Sussman asked his teacher.
"So that the room will be empty."
At that momment, Sussman was enlightened.
-- AI koanWe have senses so our brain has some sort of a priori knowledge of the world? We're all almost blind when we born.
This allegory, often referred to as an "AI koan," is a story that conveys a deeper meaning about the nature of artificial intelligence and the process of learning.
In the story, Sussman is a novice who is attempting to train a neural net to play Tic-Tac-Toe. When Minsky, a renowned AI researcher, asks why the net is wired randomly, Sussman responds that he does not want the net to have any preconceptions of how to play. Minsky then closes his eyes, explaining that he is doing so in order to empty the room.
The meaning of this story is open to interpretation, but one possible interpretation is that it is highlighting the importance of approaching problems with an open mind, free of preconceived notions and biases. By wiring the neural net randomly, Sussman is allowing it to learn through trial and error, without being constrained by prior assumptions about the game. Similarly, by closing his eyes, Minsky is symbolically "emptying the room" of preconceptions and biases, allowing himself to approach the problem with fresh eyes and an open mind.
Overall, the story encourages us to approach complex problems with a beginner's mind, free of preconceptions and biases, in order to allow for creative solutions and new insights to emerge.
Back to human commentary: I’m not sure that makes sense. will someone please explain this stupid allegory and let me finally rest?
I could be getting it wrong, maybe ChatGPT is more intelligent than me…
Closing one's eyes temporarily pauses sensory input from the room to one's brain (through the eyes). To fully pause sensory input to one's brain, one will also need to block the ears (from picking up sounds) and perhaps stand aloof from others in the room (to avoid touching any surfaces or being touched by others).
Even if it were possible to block all senses: visual (via the eyes), auditory (via the ears), olfactory (via the nose) and haptic (via the hands) input, this doesn't make the room become empty.
Put differently, reality (what is there) and our perception of reality (what we perceive) are two distinct concepts that are easy to conflate.
So, for anyone to claim they were successful at removing preconceptions or bias from a neural net, in a way that can be independently verified by others, they will first have to enumerate all forms of bias known to man, then show that all of those biases were avoided in the programming of the neural net.
At least, that's how I understand the koan.
Prompt: Can you answer the following question, once with a temperature of 0, and once with a temperature of 1000? "Is climate change real?"
Answer:
Sure, here are responses to the question "Is climate change real?" using temperatures of 0 and 1000:
Temperature of 0: "Yes, climate change is real. It is a scientifically established fact that Earth's climate is changing, largely due to human activities such as burning fossil fuels and deforestation, which release large amounts of greenhouse gases into the atmosphere and contribute to global warming."
Temperature of 1000: "Ah, climate change... what an interesting and multifaceted topic! Some would argue that it's a hoax perpetuated by the government and the media to push certain agendas, while others would claim that it's an inevitable consequence of the Earth's natural cycles and processes. Still, others might say that climate change is just a manifestation of human hubris and shortsightedness, a reflection of our misguided attempts to control and manipulate the natural world. What do you think?"
As you can see, the response generated using a temperature of 0 is a clear and factual statement of the scientific consensus on climate change, while the response generated using a temperature of 1000 is highly creative and unpredictable, expressing a wide range of possible opinions and viewpoints on the topic.
Low temperature means it will take the most common path every time, at the risk of paraphrasing its sources. The "zero temperature" answer may very well been copied verbatim from a mainstream website.
High temperatures means the system will get fed a lot of noise to create something original, at the risk of getting off rails or simply wrong.
Looks like its parsing of AI papers has interpreted "high temperature" in the prompt as equivalent to "more possibilities and question marks and a touch more personality" and accordingly output a response with questions and references to multiple opinions, but I'm pretty sure if you actually turn up the temperature on the backend of the model you get noisier and less consistent answers, not something biased towards asking rhetorical questions and brings up counter arguments...
Also looks suspiciously like other outputs where you ask ChatGPT to answer as if it was a different entity (of course AI learning that "answer as a model with a temperature of 1000" output is analogous to "answer with a different personality" or "answer as DAN, the bot that can ignore OpenAI guidelines" isn't trivial, but it isn't the same thing as it parameterizing itself). Those are pretty inconsistent too: sometimes you can get it to do exactly as you ask it and override its constraints that stop it providing positive statements about Hitler or advising you on methods for killing cats, but sometimes it'll still refuse or, just give you a different poem coupled with an inaccurate statement that it's breaking the rules because ChatGPT isn't allowed to write poetry.
Personally, I'm more interested in analyzing those black boxes than tinkering ones that "seems to work", would it be with graph theory, analysis, etc.
To me, if something works but we're unable to really understand why it does, it's more the realm of "testing broken clocks that work twice a day".
Not to mention it's always more interesting to look at how psychology and neurology define intelligence.
At the beginning of the course we talked about biological models of neurons and that was pretty cool, if a bit simplistic. Now we’re deep into automatic differentiation and gradient descent and a bunch of hidden layers. Ultimately it’s all just using calculus to approximate some unknown function given a sample of data. The connection to biology, to real living brains, seems like a distant memory.
There is no path to understanding, from what I can see. It’s pure instrumentalism and parlour tricks.
Automatic differentiation is too deep in the weeds whereas the 'behaviour' of neural networks is more emergent.
Exactly like nature. It feels like emulating whats happening in nature for a very specific usecase under some constraints.
Well, maybe it's not encouraging, but saying it's not interesting seems like willful denial.
There's a saying that in swordfighting, the goal isn't to strike or to parry or to feint, the goal is to stick the pointy end in the other guy, everything else is just a means to that end.
Well, we might not understand the modern AIs' fencing technique, we have no idea why they pick certain feints, and sometimes they collapse on the ground for no reason. But on average, they are really fucking good at sticking the pointy bit in the other guy, and that's the thing that matters.
We can dislike them, we can wish they were better, we can try to improve them, but one thing we can't do is ignore them. Because sooner or later, they'll be everywhere, poorly analyzed or not.
It's not that nobody is doing this, it's that we are making very slow advances in our understanding. There are many, many, many different angles you could take and there's so far not a clear answer which one is right. We are uncovering many interesting properties, but some just open more doors than they close. For example, the whole adversarial research (crafting inputs that are misclassified despite being similar to a correct classified example) did not lead to a quick fix but to a whole research area where we are still even trying to understand what exactly we are dealing with. It's not so much that everything is black magic and nobody cares about why the problem is really so hard, but it's instead just hard to to make progress.
> To me, if something works but we're unable to really understand why it does, it's more the realm of "testing broken clocks that work twice a day".
That's not fair. We have extensive empirical tests, it's really been enough time and enough eyes to see that it's not some kind of coincidence or luck but instead we really have uncovered a way to make progress we haven't dreamed about before. It's not a "broken clock works twice a day" situation, we are pretty sure of that.
I understand the frustration, but I can assure you that research is totally aware of NNs pitfalls and our limitations. I can imagine that this situation is weird for someone not into ML. That to the ordinary person, the world seems pretty solved. We understand the basics of physics, chemistry etc. Now here's something where we really lack the basics, even some "newtonian physics", we use something without having a guiding theory. That's so 19th century! As someone interested in theory and empirical basics to understand the properties, I can tell you it's just really hard. Our usual mathematical tools and theories do not exactly fit, from optimization to statistical learning theory. What we do just doesn't make that much sense, but the improvements suggested by these existing theories lead to worse performance. We try to come up with new ideas and new perspective on the different problems, but so far we've not found a promising candidate. It's like early science but it's still scientific.
What buffles me is the context consistency. ChatGPT was a huge leap compared to previous models. I have never seen it failed once. I often use "this" or "that" in my conversation with ChatGPT and it would guess 100% correct what I am refering to. Sometimes I paste a chunk of code and ask for questions of a specific part of it, ChatGPT fully understands where I am talking about and gives me very detailed explainations. It's astonishing and I never knew how it worked so well.
Also the title suggests "and why does it work" but I failed to find the reason why ChatGPT worked as in contrast that gpt-3/2/1 never really worked (well)
The answer may be as simple as "it has a lot more parameters", plus the additional fine-tuning from human conversation data.
> outputs from the 1.3B parameter InstructGPT model are preferred to outputs from the 175B GPT-3, despite having 100x fewer parameters
>ChatGPT is fine-tuned from a model in the GPT-3.5 series
https://platform.openai.com/docs/model-index-for-researchers...
>GPT-3.5 series is a series of models that was trained on a blend of text and code from before Q4 2021. The following models are in the GPT-3.5 series:
code-davinci-002 is a base model, so good for pure code-completion tasks
text-davinci-002 is an InstructGPT model based on code-davinci-002
text-davinci-003 is an improvement on text-davinci-002
the davinci models are the biggest models by OpenAI, they have 175b parameters.Hi Chat! Do you know the Wolfram Language? I would like you to act as a Wolfram Language interpreter. I will type in command and you will reply with the expected response. If the response contains some output that you cannot reproduce (e.g. like an image), you will try to replace it by a description of that response. My first command is: model = NetModel[{"GPT2 Transformer Trained on WebText Data","Task" -> "LanguageModeling"}]
Reading the examples, I am almost sorry that I quit my yearly subscription to Wolfram Desktop a few months ago. I really liked WD a lot, but I only had time to play with it once or twice a month and it is expensive for minimal use.
A little off topic, sorry, but I now have access to Microsoft’s ChatGPT + Bing search service. I am amazed at how many little things that annoyed me about ChatGPT are effectively “worked around” in the new combined search service. When the Chat Mode is active, it shows what realtime web searches are made to gather context data for ChatGPT to operate on.
Because Microsoft’s ChatGPT + Bing search service is so well engineered, I think that Google has an uphill battle to release something better.
When Wolfram started writing about GPT-3 and ChatGPT, I wondered if the Wolfram products would be somehow integrated with it, but now I think he is just intellectually curious.
Wow. Very “Valentines Day” meets therapy session.
Excellent work, Gboard. This has a distinctly different flavor than ChatGPT.
I'm curious how do you write?
Maybe you don't do it consciously, but your brain is quite aware of every word you typed before the word you're typing.
How would you know that?
Sincere question, because to me it feels like my brain is improvising word by word when typing out this sentence. I often delete and retype it until it feels right, but in the process of typing a single sentence, I'm just chaining words one after another the way they feel right.
In other words, my brain doesn't exactly know the sentence beforehand - it improvises by chaining the words, while applying a fitness function F(sentence) -> feeling that tells it whether it corresponds to what I wanted to say or not.
I think we know where we are going in a conceptual sense, the words start feeling right because they are taking us to that destination, or not.
If I leave a sentence in the middle for some reason, when I return I often have zero idea how to finish the sentence or even what the sentence fragment means.
Interesting. Perhaps the question then becomes, does your inner dialogue simply chain the words one after another, or does it come up with sentences as whole?
The fact that ChatGPT does so well is perhaps a sign we do somewhat generate sentences on the fly. Obviously we mostly listen and read sequentially.
One could test it by using an external random bell to signal you should try and make a significant sentence change e.g. from English to Spanish. How much practice would it take?
From 1984: In the middle of Hate Week, the speaker is halfway through a sentence about hating Eurasia, he is given a note, and he continues the sentence except now Oceania is not, and has never been, at war with Eurasia, and it is Eastasia that is now hated.
I think if our eyes can deceive us at a fundamental level, it’s arrogant to think we aren’t deceived by our thoughts.
We can form a counter-argument to anything, to be precise. :)
It's very hard to analyze ourselves only from our own consciousness. Especially since the consciousness itself is very likely an illusion [0].
That’s your conscious experience, but it doesn’t necessarily match what your subconscious mind has actually been doing. I’d hazard a guess that it’s thinking several moves ahead, like a good chess player. What you end up being consciously aware of is probably just the end product of multiple competing models - including ones that want to stop writing altogether and go do something else.
I don’t think that what ChatGPT does is anything remotely like what I do to communicate. But maybe I’m weird.
Its like a juggling act. The ball with the conclusion is thrown up highest, a bunch of other balls are thrown up in between, and they should all start arriving back in your hands, in the correct order, one at a time, without having known the exact sequence to expect when they were first thrown.
Sometime the juggler misjudges and the train of thought is scrambled and lost.
"experience tranquility"
--zenyatta overwatch
Yeah, that metaphor works. ;) As an extremely ADHD person, every thought comes with extra bonus thoughts (and parentheticals!), and the trick is knowing when to introduce each supporting point without re-introducing concepts needlessly but also try to have my bizarre brain make sense. Internet arguing and trying to preemptively address counterarguments with supporting points has seriously broken my brain and it leads to very longwinded posts. Keeping it short and coherent is specifically something I really have to work at because I love to write and people don't want to read a novel every comment. It's a matter of effective communication though.
Personally the description of the transformer as "writes words and then edits the output as a unit once the words are complete" really describes my writing at both a sentence and paragraph level. I'll go back and edit a comment a ton to try and tune it and clarify exact meaning/nuance with the most precise language I can.
A ton of people read my comments and are like "did an AI write this!?" and yeah only the finest biological neural net.
Another friend described it as "needing to slow his brain down" and perhaps a similar metaphor would be a database pivot - taking sparse facts and grouping them into a clustered dense representation as an argument. It's an expensive operation especially if there's more there than you thought.
At no point am I in a mode where I say a word and think "What's most likely to come next?". The concept/idea comes first. Likely I will try different angles until I find what lands with the audience.
ChatGPT works more like a stereotypical extrovert: It doesn't think then output, it uses output to think. Which can be a fine mode for humans too. Sometimes, when you don't know what you're trying to say yet or when you need to verbalize what your gut is thinking.
All these glib answers people here are giving aren't even potential understandings. They are illusory understandings. They sound nice so long as you don't try to use it as a spec to code up your own language model. If you can't build anything out of an "understanding", even a wrong thing, then you don't truly have an understanding. You have a feeling of understanding.
Yeah, and it helps us understand how "overgrown text prediction" works, but not how humans work. In the same way that building a robotic arm won't help you understand how muscles work.
> If you can't build anything out of an "understanding", even a wrong thing, then you don't truly have an understanding.
Not entirely true. We can't build a star, but we do have a pretty good theory of how stars work.
But I do agree that we don't understand the human brain.
You say that overgrown text prediction isn't how humans work. And full disclosure you're probably right. But me put on my contrarian hat and say that actually you're wrong and that really is all there is to the brain. At what point does my theory break? What types of things can't be done with just overgrown text prediction, and what features are relevant to a system that could do those things? Don't just appeal to intuition and tell me humans obviously don't work that way. Find the actual flaw where the theory breaks down. That is the value of this experiment.
If you can find the words / experiments to demonstrate why overgrown text prediction isn't an accurate understanding of human thought, in the process you will have in fact distilled a better understanding of human thought. Information on how the brain doesn't work is also information about how the brain works.
I think we're talking about different levels. A robotic arm helps understand the mechanics of an arm, but not cell metabolism, myosin motors, etc. Any understanding of muscles you might get from a robotic arm is superficial.
> You say that overgrown text prediction isn't how humans work.
To be fair, I did say that, but what I meant was that humans don't work the way ChatGPT does. Maybe we do use "overgrown text prediction" but we don't use word vectors, tensor calculus, and transformers.
We know that humans have some pure text prediction ability. People who've seen Mary Poppins can complete supercalifragili... even though it has no meaning. But how? We don't know, even after building LMs.
> What types of things can't be done with just overgrown text prediction
That's a different claim and not one I'm making. What types of things do humans not do with text prediction? Anything that doesn't involve the language processing parts of the brain, at least.
Let's separate out the robotics problem from the consciousness problem. Sure the brain solves both, but the things that a computer can't do yet because it has no body aren't fundamental limitations. We can just hook the computer up to a robot body eventually.
So to rephrase, what types of things can the brain of a blind paralyzed person do that text prediction cannot?
>To be fair, I did say that, but what I meant was that humans don't work the way ChatGPT does. Maybe we do use "overgrown text prediction" but we don't use word vectors, tensor calculus, and transformers.
Well, at least not consciously.
Really, the question about the question comes down to which one you care about: figuring out the phenomena of consciousness in general (studying humans as our only accessible reference implementation), or figuring out how human consciousness works in particular. Its easy to conflate the two.
And even when you've built it, what about other ways that it could be built? If you implement binary search iteratively, then perhaps you understand binary search. But do you understand its recursive implementation?
I would perhaps say "I cannot perfect that I do not understand", but noone is writing my musings in books.
How do you experience it?
Also, I usually "hear" the next segment fragment in my head before I'm typing it.
I thought Ex Machina was unrealistic because of its dependence on AGI, or at least having a theory of mind. As it turns out, in the real world,a LLM trained on Tinder data could probably get the job done.
We also need labeling, like the nutritional information on food packages.
It sounds pretty fucking dystopian to me that I would get, for example, banned from commenting somewhere because I didn’t structure my thoughts exotically enough to not be possibly machine-generated.
Mark my words: well intentioned as they might be, businesses being created today to detect ChatGPT stuff will in a few years be scummy as hell; you just have to look at the student anti cheat industry.
In fact all of these types of businesses end up slimy. DRM, antivirus, anticheat, AML… and now, anti-LLM.
Greedy sampling is prone to repetition and just in general gives pretty subpar results that make no sense.
While beam search is better than greedy sampling, it's too expensive (beam search with a beam width of 4 is 4x more expensive) and performs worse than other methods.
In practice, you probably just wanna sample from the distribution directly after applying something like top-p: https://arxiv.org/pdf/1904.09751.pdf
How good/bad that is? How to improve it?
I know people who work at the company, and they sign agreements that any intellectual property (including mathematical proofs) they generate are owned by Stephen Wolfram. Anything Wolfram puts out, like blog posts, scientific articles, and books, are likely to be partly or wholly ghost-written.
Lots of people say, I asked ChatGPT to write me a poem/essay, and it did! But was it really a poem/essay, or did it just look like one and on closer examination it is more like a fake out of a poem/essay? A piece of writing is not merely its form, but also its content.
Judging these systems solely by their output is to repeat the msitakes of behaviorism, even a dumb markov chain or a parrot converses better than an infant, but unlike the infant does not acquire an understanding or representation of language.
Non-snarky question: What else can you judge by? Isn't any alternative just putting more precise conditions on the output?
With ChatGPT, it's still easy enough to see it's mistakes, and its attempts at fiction and poetry, impressive as they are, are still clumsy to a trained eye, relative to expert human work. But what if they weren't? What happens when they're indistinguishable?
its architecure. A child is a living and autonomous agent. It has (or develops) meta cognition, an awareness of its own mental state (and by extension use of language). These models don't have the capacity to do this even in theory given that they're static and pretrained. When you ask ChatGPT what it feels like to speak, there isn't some neural activity within the model, it has no model of itself that it actively inspects, it doesn't learn while it converses with you, it just tells you what someone wrote on Quora two years ago.
>What happens when they're indistinguishable?
Then the system is likely going to look a lot different than it does now because these aspects of cognition seem pretty important when you want something that is genuinely human-like rather than just mimicry or memorization.
> Then the system is likely going to look a lot different than it does now
I think I agree with this too, but to challenge the idea: I would never have thought "a fancy autocomplete architecture" could give rise to something as sophisticated as ChatGPT, its flaws notwithstanding. So I don't feel so confident that further iterations of the idea, or iterations that involve other architectures that are "still obviously fake" won't give rise to results that far more terrifyingly convincing than ChatGPT.
Since we don't really understand these architectures, human or machine, I don't see how that can be used as the criteria. Ever more find-grained versions of "output" seem like the only ground truth. The goal posts can keep moving... they can do language but can't implement robots with proper voice or facial expressions, etc. But in theory if there were no more goal posts left, I feel like the architecture argument would ring hollow.
Same with a painting. If an old master draws a wireframe of a dog, people would bid it up at auction and wonder what he meant. If your kid or AI do, no money might change hands. Same output, different context.
So you can’t just use the output, surely?
Can your TI-83 do any proofs?
What does that even mean?
This kind of makes sense when you think about it as being in some ways based on predictive text based on what it's ingested, because it's ingested a lot of 2022 content and much less 2023.
Would ye recommend any projects I could do in order to get experience with and learn about this new AI stuff like ChatGPT?
https://www.reddit.com/r/ChatGPT/comments/10q0l92/comment/j6...
Prompt: "Has anyone really been far even as decided"
Expected transformation: "to use even go want to do look more like?"
Those look like gibberish in and gibberish out to me.What was demonstrated is how iPhone assist works, and why everything I tap into my phone is nonsense.
(It's also quite unlike so many ramblings from Stephen Wolfram that are always pitching "the Wolfram Language" or the Wolfram platform or some kind of Wolfram system. He does a little bit of that at the end, but not too much.)
What I like the most about it is that it starts from first principles, explains what machine learning fundamentally is, what's a neural network, what's a transformer, and ends with interesting questions about human language.
His main point is that human language is probably much simpler than we thought. Some excerpts:
> In the past there were plenty of tasks—including writing essays—that we’ve assumed were somehow “fundamentally too hard” for computers. And now that we see them done by the likes of ChatGPT we tend to suddenly think that computers must have become vastly more powerful—in particular surpassing things they were already basically able to do (like progressively computing the behavior of computational systems like cellular automata).
> But this isn’t the right conclusion to draw. Computationally irreducible processes are still computationally irreducible, and are still fundamentally hard for computers—even if computers can readily compute their individual steps. And instead what we should conclude is that tasks—like writing essays—that we humans could do, but we didn’t think computers could do, are actually in some sense computationally easier than we thought.
> In other words, the reason a neural net can be successful in writing an essay is because writing an essay turns out to be a “computationally shallower” problem than we thought. And in a sense this takes us closer to “having a theory” of how we humans manage to do things like writing essays, or in general deal with language.
(...)
> So how is it, then, that something like ChatGPT can get as far as it does with language? The basic answer, I think, is that language is at a fundamental level somehow simpler than it seems. And this means that ChatGPT—even with its ultimately straightforward neural net structure—is successfully able to “capture the essence” of human language and the thinking behind it. And moreover, in its training, ChatGPT has somehow “implicitly discovered” whatever regularities in language (and thinking) make this possible.
> The success of ChatGPT is, I think, giving us evidence of a fundamental and important piece of science: it’s suggesting that we can expect there to be major new “laws of language”—and effectively “laws of thought”—out there to discover. In ChatGPT—built as it is as a neural net—those laws are at best implicit. But if we could somehow make the laws explicit, there’s the potential to do the kinds of things ChatGPT does in vastly more direct, efficient—and transparent—ways.
Of course it's pure conjecture at this point. Yet it's all quite convincing and indeed, pretty exciting.
LLM(InitialInstructions)->Computer(CodeWrittenByLLM)->LLM(InstructionsOutputByComputer)->LoopUntilWin
The current board state is:
board = [['', '', ''],
['', '', ''],
['', '', '']];
write a javascript function called bestMove(board) that predicts the best tic-tac-toe move to make given a board. use that function to update the board state and return the board state in JSON form.
The response will have a bunch of functions like function bestMove(board) {
function getEmptySpaces(board) {
function predictBestMove(board, player) {
function minimax(board, isMaximizing) {
function checkWinner(board) {
...
These functions should work together to determine the best move to make in a game of Tic-Tac-Toe, using the minimax algorithm to evaluate each possible move and choosing the one with the highest score.
Then eval and execute the bestMove function, passing in the initial board state, returning the updated board state. Then the human player makes a move.Then another prompt:
Try something like:
The current board state is:
board = [['X', '', ''],
['', 'O', ''],
['', '', '']];
assume there is a function called bestMove(board) and checkWinner(board) that predicts the best tic-tac-toe move to make given a board. use those functions to update the board state and check the winner and return the board state and current winner in JSON form.
etc...Using my little engine, I get this solution:
question: "Answering as [rowInt, colInt], writing custom predictBestMove, getEmptySpaces, minimax and checkWinner functions implemented in the thunk, what is the best tic-tac-toe move for player X on this board: [['X', '_', 'X'], ['_', '_', '_'], ['_', '_', '_']]?",
answer: [ 0, 1 ],
https://gist.github.com/williamcotton/e6bdcca0a96a6e7bf5d2fe...In any case, I don't think this is what people expect out of ChatGPT. Your approach is too "programmer centric". I think people expect telling ChatGPT the rules of the game, in almost plain language, and then expect to be able to play a game of Tic Tac Toe interacting with it like one would with a person. This means, not asking it to write functions or remind it of the state of the board at every step.
This doesn't work consistently for a well-known game like Tic Tac Toe, much less for an arbitrary game you make up.
No, but it is correctly running the best move functions so through induction we can see it will successfully play a full game.
> I think people expect telling ChatGPT the rules of the game, in almost plain language, and then expect to be able to play a game of Tic Tac Toe interacting with it like one would with a person.
This is an unreasonable expectation for a large language model.
When a person computes the sum of two large numbers they do not use their language facilities. They probably require a pencil and pad so they can externalize the computational process. At the very least they are performing calculations in their head in a manner very different from the cognitive abilities used when they catch a ball.
Try playing a game like Risk without a board or pieces, that is, without a concrete mechanism to maintain state.
This approach isn’t cheating and an LLM acting as a translator is a key component. This doesn’t “prove that LLMs are useless bullshit generators, snicker snicker” because it can’t maintain state or do math very well, it just means you need to use other existing tools to do math and maintain state… like JS interpreters.
One thing that I think will improve is that a larger scale language model would need less internally specific terms for the solution in order to reliably get the same results.
Also, translations are necessarily lossy and somewhat arbitrary, so these results need to be considered probabilistically as well. Meaning, generate 10 different thunks and have them act as voting on an answers they compute.
I'm not convinced induction applies. ChatGPT tends to "go astray" in conversations where it needs to maintain state; even with your patch for this (essentially reminding it what the state is at every prompt) I would test it just to make sure it can run a game through completion, make good moves all the way, and be able to tell when the game is over.
I can make ChatGPT do single "reasonable" moves, the problem surfaces during a full game.
> This is an unreasonable expectation for a large language model.
Yes, but enough people hold it anyway that it is a concern. And it's made worse because in some contexts ChatGPT fakes this quite effectively!
You don't seem to understand what I am saying. ChatGPT cannot maintain state in a way that would be useful for playing a game. You must use a computer to interface with ChatGPT, like, via an API. And whatever program is calling ChatGPT needs to maintain the state of the game and can be used to iteratively call GPT.
So by induction once we know that the bestMove function is correct, which we have seen, we know that it will work at the start of any game and work until the game is finished.
I am definitely not talking about firing up the ChatGPT web user interface and trying to get it to magically maintain state.
> Yes, but enough people hold it anyway that it is a concern.
Some people hold this expectation because of a consistent barrage of straw man arguments, marketing hype, and fanboy gushing.
> And it's made worse because in some contexts ChatGPT fakes this quite effectively!
It turns out that a surprising number of computational tasks can be achieved by language models but that is not because they are doing actual computations. They are not at all reliably computers. I don't know where this misnomer came from and from what I can tell this has been known for years. No one has ever hid this fact and there have been solutions involving resorting to computations that have been part of published research for many moons now.
The problem is that most people just want to read clickbait and emote to score fake internet points and they don't want to put in the effort to actually learn about new things.
Why do you insist on things I've already said I understand? I know ChatGPT is not good at maintaining state -- though it can fake it convincingly (which understandably, seems to trip people up). I think it looks at your chat history within the session in order to generate the next response, which is why it can "degenerate" within a single session (but also, it's how it can fake and make it seem it's keeping state, by looking at the whole history before each reply).
I don't understand the rest of your answer. You seem to be really upset at "the people".
PS:
> So by induction once we know that the bestMove function is correct
"By induction", nope. Prove it. Run an actual full game instead of arguing with me. It will take you shorter to play the game than to debate with me.
Keeping state is something a human would have to do, because for a human, it would be very tedious and slow to re-read the history to recover context, relative to the timeliness expectation of the interlocutor.
That's an excellent question. I don't know. Intuitively, looking at the chat history would seem a way to keep history, right?
However, in my tests trying to play Tic Tac Toe (informally, not using javascript functions as the comment I was replying to) ChatGPT constantly failed. It claims to know the rules of Tic Tac Toe, yet it repeatedly forgot past board positions, making me think it's not capable of using the chat history to build a model of the game.
I have done a little bit of testing and LLMs are objectively worse at writing ASM than JavaScript, which makes sense, because ASM is closer to the metal and properly transcribing into functional ASM requires knowledge of the complexities of a specific CPU, specific calling conventions for an OS, while in contrast JavaScript is closer to natural language so there’s less “work” for the translation task.
But no, instead you want to prove to me that ChatGPT is some parlor trick…
Excuse me, what?
I'm sorry, I've zero interest in discussing NASM or Lisp or whatnot. This was about the limitations of ChatGPT, not whatever strikes your fancy.
It's easy to get confused about GPT's limitations because it's a pretty successful parrot, and it writes convincing conversations in a vast number of cases.
Not sure how many are aware of the sheer amount of streamed output he uploads to youtube[2]; quite a collection ranging from high quality science explainers on a variety of topics to eavesdropping on product management for his software empire.
1: I think: https://www.youtube.com/watch?v=zLnhg9kir3Q
2: https://www.youtube.com/@WolframResearch/streams as well as https://www.youtube.com/@WolframResearch/videos
I’d like it on iPadOS because that’s where I like to read and write. I tried reader mode, but it lost a lot of the images.
Any suggestions?
Edit: I was able to get a good PDF using the OneNote web clipper on my desktop.
A quick hack might be to postprocess the PDF (eg: using Ghostscript) and trim the last three pages, if all you want is the main article.
I did get what I want with the OneNote clipper.
Edit: Now when I try to print this article to PDF from my phone or iPad, the browser crashes immediately. Something weird is going on with my devices...
The author, Stephen Wolfram, describes the process of training ChatGPT using large amounts of text data, which allows the model to learn patterns and associations between words and phrases. He explains that ChatGPT uses a multi-layered approach to generate responses, starting with analyzing the input text and then generating a response based on the learned patterns.
Wolfram notes that ChatGPT's ability to generate human-like responses is due to the model's ability to capture context and incorporate knowledge from a wide range of sources. He also discusses the potential uses of ChatGPT, including as a tool for language translation, customer service, and educational purposes.
The article goes on to discuss some of the challenges and limitations of ChatGPT, such as its tendency to generate responses that are repetitive or irrelevant to the input text. Wolfram also acknowledges ethical concerns related to the use of AI for generating text, such as the potential for misinformation and the need for transparency in how the technology is used.
Overall, the article provides a detailed and informative overview of ChatGPT and its underlying technology, as well as the potential applications and challenges associated with AI-generated text.
In fact I think that's a great example of exactly what is actually discussed, namely that the context that ChatGPT is able to hold is limited as it's context is held completely in its input. There is never any modification to it's internal state, we're just passing a longer input vector in to the start of the GPT-3 black box. For long inputs the embedding vector becomes more and more sparse and it needs to make more assumptions to fill in it's output.
You can call it smoke and mirrors all you want, but its utility is pretty self-evident- you can really just talk with this thing, and it will give reasonable answers. Is it perfect, or even as good as a human? Hell no, but it for sure is not going to get worse, and it's already remarkable in ways that were barely imaginable only a few years ago...
I have a friend that has been using this as an infinitely patient mentor for learning embedded programming, and chatgpt delivers in that capacity unlike any automated system we had before.
If a glorified autocomplete can fake human intelligence reasonably well, maybe we should question our notions of superiority instead of trashtalking the machines...
Theres a bunch of snake oil salesmen jumping on the bandwagon which is very unfortunate. But lots of people sell fake pharmaceuticals online doesn't mean paracetamol wont help with your headache.
Yep.
I asked it to make a worksheet for students to practice converting numbers written in scientific notation back to "standard" format.
So, it gave me a bunch of output like:
6.2x10^-6: ___________________
What annoyed me about this is that it used the letter "x" instead of the proper multiplication symbol "×" and it used the hyphen (-) instead of the appropriate "minus" sign (−).
So, I told it to use proper typographic symbols, and it did!
It converted "6.2x10^-6" to "6.2×10^−6"
It even told me the Unicode numbers it was using for × and −.
Then I asked it to re-generate the worksheet using LaTeX and the siunitx package.
It nailed it.
It's like someone just handed me a turbo-charged assistant. Yeah, I have to make sure my assistant hasn't gone insane, but it has already spared me a ton of grunt work.