AI chatbots are having their “tulip mania” moment
salon.com
salon.com
But now that the dust is settling, people are realizing that this technology is not flawless. We haven't fully 100% mastered it, there are unsolved problems, and maybe it's not quite ready for a broad deployment just yet.
That's normal. It's just the way the hype cycle works. People (and the media) quickly get really excited and build unrealistic expectations, then they become aware of the flaws and limitations. That doesn't mean these kinds of systems can't be better and even more useful in 1-2 years. There's a lot of really smart people working on them with many possible paths to address the current limitations.
Siri doesn't need AI to be useful - it really just needs this $2t company to invest in some basic usability around a set of prompts that people could actually use.
Yesterday's ridiculous Siri failure: "hey siri, set alert volume to 25%" Siri: "media volume set to 25%" That is not bad AI, that is a total failure of UX and it makes it clear that Apple has simply failed to invest resources into making Siri useful.
We have tons of patterns for CLI and GUI usability. But in both cases, we have a easy enough way for users to discover what is possible—in the CLI, you can present a list of commands or categories for commands. In the GUI, we have menus, dialogs, and windows that let you explore. Voice is a much more difficult problem in the first place. Voice commands are often modeless, discovery is difficult, and commands are given with troublesome natural language patterns which must somehow be adapted to multiple languages.
Take a look at all the different voice assistants—Siri, Alexa, and Google. They all kinda suck, despite enormous amounts of money being poured in. From recent news, we know that Alexa managed a loss close to $10 billion in a single year. By all signs, it contains some brilliant engineering, but brilliant engineering doesn’t matter if it doesn’t solve the underlying (difficult) problems well enough.
(I don’t think tools like ChatGPT are going to push the needle much here, at least for a couple years, because integration with all the different software systems takes such a massive amount of coordinated work.)
1. The novelty of the thing!
2. Start using for some simple stuff like weather, timer, alarm, etc.
3. Try more complicated things like skills, really managing controls/media in the car, reliably playing some album.. and largely give up because it's frustrating.
4. Mostly go back to 2.
My example of a Siri fail is not at all indicative of a "major UX problem. The fact that Siri doesn't understand iOS nomenclature and that iOS doesn't have the proper hooks to do something that a user would obviously want to do, is indicative of a total failure to invest resources in making Siri useful.
>hey siri, play Chet Baker Sings on the Apple TV
>playing the best of chet baker sings on Apple TV
>Siri, play the album “chet baker sings” on Apple TV
>playing the best of chet baker sings on Apple TV
Those are not the same album…and only one of them is saved in my library, the one I asked for. It failed repeatedly, too. The only response Siri would produce was to insert two new words to my request and then serve that.
> hey siri, play the art of fugue by bach on downstairs
> now playing Love The Way You Lie featuring Rihanna, by Eminem, on downstairs
And like have apple’s ML people just never even heard of Bayes or what? if i have, in the last few days, asked siri to play debussy, ravel, liszt, and chopin, and today I say “play debussy on downstairs” - and i have never, not once, played any kind of rap or hiphop - why would it... “now playing Ain’t No fun - by Snoop Dogg - on downstairs”?
It seems to be more about willingness than capability.
Now that ChatGPT is there there is no choice other than to expose what they have of become irrelevant.
Prompt: "Write a three-paragraph story about someone who cannot see."
Response:
Maria was born without the ability to see. Growing up, she always felt like she was missing out on something that everyone else took for granted. She would hear her friends describe colors and the beauty of nature, and while she could imagine what they were talking about, she knew she could never truly experience it for herself. Despite this, Maria never let her blindness hold her back. She learned to read Braille and used her heightened sense of hearing and touch to navigate the world around her. She even took up piano and became an accomplished musician, playing by ear.
One day, Maria met a man named Alex who was fascinated by her ability to navigate the world without sight. He would often accompany her on walks and describe the beauty of the world around them, pointing out flowers and trees and the changing colors of the sky. Maria found herself falling in love with Alex, and although she knew that their lives would be different because of her blindness, she was willing to take the risk. Alex was patient and understanding, and he took the time to learn how to communicate with Maria in a way that she could understand. He even started taking piano lessons so that he could play duets with her.
Years went by, and Maria and Alex built a life together filled with love and happiness. While Maria never regained her sight, she realized that she didn't need it to experience the beauty of the world. Through her other senses and the love of those around her, she had everything she needed to live a full and meaningful life.
To illustrate what I mean, do you think it's likely/possible ChatGPT ever comes up with a story about a person who goes to the grocery store, examining produce for ripeness, chatting with another parent about parenting, then going home and feeding their child? None of those things require sight, so it would fit the prompt as long as it portray the person as being able to see.
Throw more details in there and put it in situations where it has to actually make the blind person navigate the world and see if it starts to violate the constraint or not.
Alexa isn't getting any better at recognizing my voice? At 8am everyday, it hasn't learned I'm likely asking for NPR, not something else? It's not consistent at all.
Controlling Hue lights with it is atrocious too.
Alexa is purely a dumb voice assistant and there is either a complete lack of investment in it, or that money is invested elsewhere like making Alexa more likely to order too much product off Amazon so someone can point to a chart that shows a bug increased sales.
Then Siri then just forgets about it and doesn't do anything.
I have five echo devices through my house and have only found about ten uses for them.
They’re useful but 1) finding uses for them and 2) twisting your words around to make it understand you is an exercise in frustration.
It seems like it would be an easy data mining exercise for these companies:
If someone asks for something and stops, don’t flag it. If someone asks for something 3+ times in a row, flag it to figure out why they didn’t get what they wanted the first time. If someone says, “Alexa you’re an idiot” don’t just play a snarky comeback. Flag that for serious review. Then iterate.
There seems to be almost no useful iteration on these virtual assistants.
It's tremendously coherent.
The chance that chatgpt just hears your example and than does the right thing is nearly a no brainer.
It already understood much more complicated prompts from my tests with ease.
Siri and other agents struggle tremendously and the implication after such a long time has to be a reflection of either a tremendous Missmanagement or that those classical approaches are just too hard.
Or that the chatgpt people are genius.
But that would be even worse: it would show how much impact chatgpt and co will have sooner than later.
ChatGPT can't do anything. Someone has to, somehow, wire it up to actual physical or digital controls. Intent recognition and slot filling still have to happen somewhere, and that means defining discrete intents and slots.
GPT may be better at mapping from speech to intents, but it can't magically interface with APIs that haven't been defined.
EDIT: To elaborate a bit, the problem in OP's story isn't caused by a bad language model, it's caused by no one at Apple thinking to define different "alert volume" and "media volume" intents. Current language models are plenty good enough to recognize the distinction, so simply adding ChatGPT won't be enough to make any feature work unless someone at Apple predicts the need and writes the interface.
If I tell it to "shtudown", it said "did you mean 'shutdown'?
There is an AI failure in distinguish "mispronounced command I know" from "command I don't know".
Basic intent recognition models are trained to produce a single neuron per intent as the output, which makes it pretty easy to use the activation levels of the output to decide whether to perform an action, confirm an action, or ask for clarification. You just need to check if the certainty is below a certain threshold.
With ChatGPT you'd have to encode the intents as text of some kind (JSON?) and hope that it doesn't just hallucinate an intent that your APIs don't have when it's faced with ambiguous input. You could probably have hallucinated intents map to a decent-sounding error message, but that feels more brittle to me than the existing approaches.
Ask it to write an SQL query and it will.
Give it upfront context and it will be able to produce API commands.
The voice recognition can trigger chatgpt API 'create API command for the following text's.
The practical problem is cost and hardware.
But this is closer than we ever were.
A much safer approach would be to have it produce a JSON object that is interpreted by the wrapping code. This allows you to inspect the interpreting code and provide guarantees about what can and cannot be done by users.
But neither execution model solves my main point: someone still has to write the interpreter (or the library functions if you go the dangerous route). Someone has to think ahead and guess what the user is going to try to do and provide APIs that allow that. And that work is most of the work: our existing intent recognition models are actually not bad at all at their job, the failings these voice assistants have are almost all to do with predicting the ways in which users will try to interact with the system.
The code generation and execution is trivially problematic. I don't need to try it out to know that if someone can coerce one of these bots to spit out its prompt, someone could coerce it to execute dangerous code if it's wired up to an interpreter.
As for the other part of my analysis, in order to empirically test it I would have to build a full-featured voice AI, release it to millions of users, and see if I was able to predict everything they tried to do. Given that my prediction is that even Apple can't do that well, I'm not sure why you think I would bother to try it out.
If an engineer told you that a bridge would collapse if built a certain way, you wouldn't insist on trying to build it anyway just to be sure. Most of the time in engineering you can't run a full-scale test, you have to make do with analyses. If you want to critique my analysis, I'm all ears, but "you haven't actually tested it" is not a critique.
All of my analysis is based on my usage of ChatGPT and my understanding of the underlying model. The hype is largely driven by people who don't understand how it actually works and think they're interacting with an artificial general intelligence.
It's a very impressive language model with a lot of applications, but with many fewer applications than the hype would suggest.
I'm hyped because I threw normal text (sentences) against it and it always did what I wanted and it didn't matter if it was bad English or bad German.
I didn't even try promt hacking or promt tuning because natural language already worked so good.
It's a tremendous good 'normal human' text interface.
And it doesn't need to be more correct to disrupt industries already it just needs to be more correct than humans.
The most crazy thing is that it's already so good and coherent that it makes totally sense to train a ml model like chatgpt instead of training humans! Because it scales.
We never had this.
And.i already mentor a few junior people I don't mind training an ai instead.
If it is "correct enough" to lull people into a false sense of security, the inevitable failures will be worse than if it people remained on guard.
The particular way in which these LLMs fabricate information makes them incorrect enough to be dangerous, not "correct enough to be useful."
People keep mistaking text completion for intelligence and understanding.
At some point it is a question of rate of error. The ML model will keep improving as the net get bigger and the dataset larger.
But I agree they are good at having a very good surface understanding while having little of dept at the moment.
Ex: a human making a mistake of visual interpretation while driving will examine in his mind the error he made and thing of the consequences in a more dangerous situation. The NN will not bother (at the moment)
In much the same way that Teslas are still literally killing people when they encounter vehicles parked across roadways in 2023, just like they did back in 2016, the difference between the appearance of understanding and actual understanding is a much wider chasm than it first appears.
The LLM will not "consider" anything. Even transcripts we've all seen in which LLMs seem to acknowledge errors (rather than doubling down and inventing false sources) are still just text-prediction, modeling what it would be like if someone acknowledged an error.
When a particular LLM seems to take the corrections offered in a widely-publicized transcript on board and not repeat them in future transcripts, that's because someone acted to modify the model, adding filters or weights to ensure that the text prediction path goes differently in the future.
It can reproduce the right output of 3 insert statements with a select *.
It can interpretate an SQL procedure too.
Whatever the architecture is behind this, is slightly more than just text prediction.
I don't think 'text prediction ' does this type of state management justice.
But hey let's see. Perhaps this is just an emerging feature from a tremendously huge text/language model.
To be able to code a program form a short description, combine different elements to make a solution, understanding is required.
The understanding is in the weights and pattern detection capabilities.
Sure, some part are missing, like the capacity we have to hold a few thought in our mind, correct ourselves, remember the mistakes we made, etc.
The same is true of NN that can differentiate cats and dogs, they detect patterns is a similar ways that humans do.
In the openAI api you can see the level of confidence the NN has on each words it output, it’s just not exposed on the chat interface.
I think the larger missings parts are by design, to give the system memories, agency, access to the internet, to programming would be careless at this point, but it’s happening anyway.
I feel it missing our capacity to examine our thought, to replay events in our head to reprogram itself when the results are not optimal.
And just an hour ago I asked chatgpt how to do something specific with docusaurus (Facebook static page generator) and it just told me the answer.
You know how often I explain things in my current position?
How slow some developers are? How often they forget things I told and explained them? How often they still get things wrong?
It's a slow process and doesn't scale very well at all.
If the table turns and we all teach one system instead ooohh boy.
It will make experts faster and better and potentially removes a certain amount of people in every industry faster than we can imagine.
You know the people who are adding some value but not that really but it's still better to have them than not having them?
Cloud probably got rid of plenty of basic sysadmins.
These new systems break through tasks were no one had an idea how we will break through.
Of course you want to do that with people too. Shouldn’t be too hard to bake in a self check module.
The idea is “we will not give you a way to make our conversional UI incomprehensibly quiet”.
It is an especially dumb policy decision on anything that’s not a HomePod, and it’s also dumb to willfully ignore the user’s intent and focus on the one thing the device is willing to do for you regarding volume, rather than explaining “sorry, I’m always going to be loud, my makers are worried you’ll complain I’m inaudible if you try to get me to stop yelling my responses” but it is deliberate.
It's also faulty to assume that improvements will happen linearly. We had a self-driving car mania in 2013-2017 with billions poured into it and people just forgot about it.
Also consider that if easy linear improvements were possible, they would've delayed launching it. AI (software in general?) seems to be the kind of thing where once R&D hits the exponential difficulty wall, they make it generally available.
You equate lack of ubiquitous self driving cars with lack of progress in self-driving technology, but even if we've had super-linear progress in terms of self-driving "driving skill," if that skill is below the threshold where the sales from pushing the software are less than the potential legal liability, companies aren't gonna push it.
On improvements— one of the things that has impressed is how much Chat GPT has improved since it was released. I don’t think it’s going to be a perfect AI in two years. But it already is working better for some use cases for me.
And more and more puzzle peaces are falling in place.
We can see were it is going.
And when you follow Nvidia with their digital twin topic, we are converging.
Very, very, very, very hard.
https://github.com/williamcotton/empirical-philosophy/blob/m...
Cars haven't made it yet because the assumption about allowable errors was wrong. Turns out it isn't "good enough" for self-driving cars to be safer than humans.
Of course it's not flawless, Google isn't and people aren't... and those are the two targets for this tech. People lie, make mistakes and mislead and Google directs to plenty of problematic content.
The thing to watch isn't the position of GPT or any disruptive technology today, but its trajectory. The trajectory of GPT is explosive growth, and that growth will be disruptive to imperfect human jobs and imperfect search engines. GPT does not have to be perfect to be disruptive.
This is especially threatening to Google because they have an organization that has never had to compete, like pandas they just consume. They don't hunt. Look at their vast graveyard of platforms they couldn't sustain even with their monopolistic power and endless money.
Consumers and ad purchasers want competition in search because like Uber and Lyft demonstrate, when you have two technology products competing head to head the margins get compressed toward zero.
Edit: spelling
The bigger problem for Google is that we've got AI which can produce millions of times more shitty blog articles to flood the internet with, that their SEO algorithms can't keep up with in order to surface anything actually relevant to humans.
Also the problem for all social media and reddit as it becomes easier to write bots that you can't casually distinguish from humans (given how many humans on reddit read like bots that are desperate for attention).
It could be great, but until an assistant of this sort runs on your hardware, and works only for you (really hard to show), it won't be your assistant. It will be some corporation's Carnival Crier.
That parenthetical "almost" is doing an awful lot of heavy lifting.
What is "ChatGPT-3" and which experts exactly regard LLMs as unimpressive?
MY feeling so far is more that laymen are less impressed of them than experts - because for people not from the field, the fact that they produce so much bullshit seems to trump all other aspects. Whereas I think for people who have kept track of AI developments a bit more, the fact just how coherent the bullshit is is still extremely impressive.
We've had random text generators for decades, starting with markov chains way back. Up until a few years ago, you were amazed if those things produced gramatically correct sentences - not even dreaming of producing any kind of coherent meaning.
In the potentially more difficult field of understanding text, there are a million small subfields, but nothing that could parse arbitrary english text and extract actionable, semantic meaning, in the face of idioms, figures of speech, complicated coreferences, etc.
Compare this to today, where you can give an LLM a written instruction in freeform and it will just execute that instruction - the answer may be glaringly incorrect, but it's exactly the kind of answer that fits your question.
In the field of understanding text, this is a progress which had seemed largely impossible before.
The NYT literally just published an embarrassingly credulous account of how the Bing AI wanted to seduce the author away from his wife and commit various acts of violence, that fully took all of the interactions at face value.
The non-experts are imagining full-on sentience where the experts correctly recognize mere word association shenanigans. Your feelings, in short, are ass backwards.
The current model (which is an early version; computers are new on earth and exist only for a statistical error length of time on humanity scale, on earth scale, let's not talk universe scale) is already vastly superior than many humans I know in almost every way (outside manual dexterity and some other fringe stuff you can easily fix with external systems, like we humans do). We have no good definition for what sentience is either; maybe our brain is word association shenanigans; connect up 2 chatgpts and call one 'inner voice' and the other 'external voice'. It will start claiming sentience in no time flat; the same as you. Why are you right? It's a feeling yeah?
This is just stupid.
We've pumped more English through GPT-3 than any existing English-speaking human has absorbed in their entire lifetime, and what we've ended up with is something that very very clearly has so utterly failed to generalize even the most basic level of understanding of basic concepts that if you ask it to count the number of letters in a word it will cheerfully pump out the wrong answer ("there are thirteen letters in the word 'twelve'") because its dataset correlates the two words with one another and it has learned precisely sweet fuck all about what it means to count, something a child's brain picks up with exposure to many orders of magnitude fewer language examples.
To imagine your brain is just an LLM is to mistake your reflection for another person in the room. Utterly daft. Get off the LLM hype train and start looking at these things objectively and critically. They're nowhere near what you're imagining.
Not to be too blunt, but you seem to be talking out of your ass. I'm tired of the overly dismissive comments here on hackernews by people who didn't bother to do the bare minimum of research. LLMs do not work with individual characters, they use tokens (i.e multiple characters, or sometimes entire words).
Seriously, slow down and read it again, all the way to the end of that sentence.
Never have I said such a thing. Also, BPEs aren't a natural limitation of any transformer-based model, it's a trick to save compute for LLMs.
> We have no good definition for what sentience is either; maybe our brain is word association shenanigans;
Sure sounds like you're suggesting the brain is an LLM to me, and I can't blame bonsaibilly for thinking that.
snark aside, I think he does have a point - afaik, we don't know what intelligence is, so it's kinda hard to make any argument about fundamental differences between "true" intelligence and LLM intelligence. I do feel like there should be something else. I saw somebody describe their mind as consisting of a "babbler" and a "critic" (a GAN, basically). The LLM would be the babbler while the critic is not yet implemented, and this sounds intuitively right to me. Then again, my intuition could be completely wrong and we may be able to get to human level intelligence with further scaling. I haven't seen any solid counterarguments yet. And not even the biggest LLM believers are denying the fact that it's not exactly the same thing as a human brain, but the question is whether it captures the gist of it.
https://github.com/williamcotton/empirical-philosophy/blob/m...
...
LLMs don't understand text. They encode statistical probabilities of relationships between strings of characters. Those are not the same thing. That's the entire damn point of this article!
Honestly, for someone taking pot shots at "laymen", this is a profound thing to misunderstand.
Right now, research is still ongoing what kind of knowledge is actually captured inside the models, but there are some hints that higher-level, "semantic" knowledge might emerge during training: https://thegradient.pub/othello/
The LLM absorbs the manifold (high-dimensional shape) of the data. The manifold contains the underlying concepts of grammatical structure, abstract thought, and inductive reasoning, which the LLM (within reasonable capacity and structural capability) captures because it is a more efficient representation of the underlying data.
Transformers are simple and contain the correct building blocks to efficiently capture these representations during training.
This allows for out-of domain generalization.
This is most, to my knowledge, of what happens in most deep learning processes and is likely 90% of the main content that one needs to know to understand neural networks at their core. I think it's pretty basic but it often gets buried in unconscionable mathematical symbols and fancy-speak to be of use to anyone reading.
We know it's more complex than just "empirical probability of word x given the n words before", because that would result in a markov chain and we know those don't generate the kind of output we are seeing.
We also see that it's able to map descriptions of tasks to its execution, even for unseen tasks. E.g., I can tell it "Write a limerick that contains the names of all living US presidents and format is as a JSON array inside a python script.". The result will probably contain some dead US presidents or some canadian prime ministers or whatever, and the text may be a haiku and not a limerick - but the output will usually be a python function with a JSON array with a poem with some names in it.
I don't see how that could work without a more abstract representation of the concepts "president's names", "poem", "json" and "python", so it can combine them meaningfully into a single response.
So it is estimating the markov chain, just in a way that is compressible and according to the inductive biases as we define them (i.e. what we nearly force the network towards with our architectural and otherwise decisions).
(Also I think the "last n tokens" term is a bit misleading: ChatGPT seems to have an n of "approximately 4000 tokens or 3000 words" [1, 2] which would amount to ~6 pages of text [3].
I've seen very few conversations even approaching that length - and in the ones that did, there were reports of it breaking down, e.g. having continuity errors in long RP sessions, etc. So I think for practical purposes we can say its "probability of the next token given all the previous tokens".)
Building a naive markov chain with such a large context is infeasible even before compression, you couldn't even gather enough training data: If you have a vocabulary of 200 words, that would give you 200^4000 [4] conditional probabilities to train. You'd have to gather enough data to get a useful empirical probability for each of them. (Even if shortened the length to the ones of realistic prompts, like 50 words or so, 200^50 is still a number too big to have a name)
Which is why, from what I've got, the big innovation in transformer networks was that they don't look at each token in their context window but have a number of "meta models" which select which tokens to look at - the "attention head" mechanism.
And I think there, the intuition of markov chains break down a bit. Those meta models make the selection based on some internal representation of the tokens. But at least I haven't really understood yet what that internal representation contains.
[1] https://www.reddit.com/r/deeplearning/comments/zk5esp/chatgp...
[2] https://help.openai.com/en/articles/6787051-does-chatgpt-rem...
[3] https://capitalizemytitle.com/page-count/1500-words/
[4] this number: https://www.calculator.net/big-number-calculator.html?cx=200...
I'm doing work currently on this and hopefully it will yield some fruit -- at least, the work relating on the internals of what's happening inside of Transformers. Just from slowly absorbing the research over the years I have a few gut hypotheses about what is happening. Hopefully any of the work I do will yield some fruit, I think a good chunk of what is happening is surprisingly standard, just hidden due to the complexity of millions of parameters slinging information hither and yon.
Thanks again for putting all of the thought and effort into your post, I really appreciate it. This is something I love about being here in this particular place! :D
Do you have a blog?
I try to be skeptical about certain possibilities within the field, but I do feel bullish about us being able to at least tease out some of the structure of what's happening due to how some properties of transformers work. At least, I think it'll be easier than figuring out how certain brain informational structures work (which has happened to some tiny degree, and will be I think even cooler in the future)! :) XD :DDDD :)
import json
presidents = ["Biden", "Obama", "Bush", "Clinton", "Carter"]
limerick = "There once were presidents five, " \
"Biden, Obama, Bush, Clinton, and Carter alive. " \
"They served our country with pride, " \
"And kept our democracy alive. " \
"May they continue to thrive!"
print(json.dumps({"limerick": limerick, "presidents": presidents}))great, so it's a politician?
I really hate that this is being heralded as a triumph. Someone being very confident in their wrongness does not make it right. This is something we've been dealing with for some time, but it seems to be orders of magnitude stronger than say 20 years ago. Now we're celebrating how our AI systems are amazing at this. I just want to hang my head in shame at what we've come to accept.
We still don't understand how these models really work, at least not in a way that teaches us anything about language or about humanity.
I find that a strange take. Almost all scientific advancement ends in one thing. Another set of unanswered questions. Quite often we gain more knowledge, but know less because we realize there are thousands more questions we should have been asking.
For example, theory of mind being observed out of just processing language is a big arrow that we may have been thinking about many things wrong in the past and open new avenues for 'scientific' tests that give us answers.
Is there a "hacker news" of theory of mind researchers? Probably some interesting conversations happening over there.
We understand how LLMs work conceptually. Human language, when parameterized as a sequence of tokens, has geometry and structure in a billion dimensional space. A priori, we didn't know that. So in a very real sense, a discovery was made.
FTFY. At least any kind of actually notable scientific advancement like using metals, farming, steam power, electricity, cars, planes, computers, etc.
If a person can use a chatbot to speed up their work then that's a major advancement in productivity right there.
We’ve effectively made like a general teenager that can be up-trained for specific tasks. It’s super useful if you know how to use it.
The unexplained Chain-of-Thought emerges from LLM surely would an huge gain in "knowledge". In fact I'd say it potentially the most important meta-knowledge human race can gain, you know, to unlock the mystery of "conciousness" itself.
Ask for a biography of a random public figure without giving notes and there is 0 chance they will provide an accurate result. This may mean we need systems that are literally memories without the noise driven generalization, and that these need to be exposed to the said LLMs as a resource. It's not insurmountable and it certainly isn't tulips.
I feel like everybody commenting on this thread ought to have to write "it's text completion, not intelligence" on a chalkboard 500 times or so first.
There’s some great chess videos for example where the bots start by making decent moves until you quickly find out they not only don’t know how to play well they don’t even understand the rules or the board. https://youtu.be/rSCNW1OCk_M What’s fascinating is how easy it is to give something the benefit of the doubt because it’s using language, we just aren’t very suspicious of things that can talk.
The models are not operating at a lower or higher level along a scale of human intelligence, so they're not genius-like or child-like. They're not operating on the same scale at all.
They're text prediction models, and many seem to insist on anthropomorphizing them in a way that causes them to forget that every other second or so.
Further if you watch the video it wasn’t moving the pieces incorrectly it was just creating them from thin air. This wasn’t like an 8 year old learning the game this was like your cat randomly hitting the keyboard.
It acts similar to a human that has a button for "recite chess rules" and presses it when you ask them to recite the chess rules. You'd then get a reply with the chess rules without the button-presser getting the same info, so they wouldn't gain any knowledge from it. I'm sure you can find a lot of people who could, if asked, read you the chess rules but even after reading wouldn't have understood them themselves. But are they not generally intelligent?
I get a feeling that people compare ML-models to a smart person that concentrates on what you ask them, and LLMs fail in comparison but do so confidently. If you instead compare them to an average person at their daily average focus level, the ML probably looks pretty smart.
I have a neighbor who is not very bright. He'll see some news on TV and will tell you about it, only he doesn't understand it and will just repeat words and half-sentences as best he can remember. He doesn't know he doesn't understand (or he does but doesn't care), so he'll confidently state utter nonsense. He's not smart and I wouldn't go to him for advice, but I think he qualifies as "intelligent" in the binary sense.
Even a dumb person is vastly smarter than an ant. So when I am saying ChatGPT is dumb I mean compared to having a conversation with a person which seems like the only reasonable scale to use with a chat bot.
ChatGPT’s model requires quite a of of processing power and sophistication to form its responses. It just doesn’t comprehend what it’s saying.
I'm sure you've talked to people before who used certain terms but it was obvious to you they didn't actually know what those terms meant. They're able to use them without understanding the concept. Just like ChatGPT, only that it's more likely to sound like ChatGPT understands the concept.
What I'm trying to say: humans have a very clear limit to what concepts they can grasp. They can punch above their weight by pretending to understand a concept, but the proof is in the pudding, so unless they're able to apply it to solve problems, they haven't understood it. And even if they do understand something, they'll make mistakes. If the test for intelligence is essentially 'can perform certain tasks with a maximum error rate of x, and can learn to perform new tasks', does 'understanding' really matter? And can you even really test for understanding, aren't you only testing for 'gives answers to questions that signal understanding'?
I don’t have specific level beyond that because it’s deficiencies don’t map to human condition.
As to a touring test given enough time and freedom in questions you can easily tell it’s a machine so it just fails. That’s the thing, the touring test isn’t can you tell if this is a machine in 7 seconds while talking about baseball it’s can to tell period.
Which, I suppose, even if the specific facts and arguments in it are wrong does support the headline thesis; every commentary outlet feeling so desperate to say something about it, even if it means rushing hot garbage out the door, is certainly indicative of something of a “tulip mania” moment.
I see a spectrum of responses to GPT, but the really important thing is the rate of change. Even if GPT somehow doesn't impress you currently, why do you think things won't be improved in 6 months? What about 5-10 years from now? I can understand thinking that transformers are maxing out on their sigmoid curve of progress, but every time people complain about perceptrons, RNNs, CNNS, etc - another team invents something new and blows out all the benchmarks. These things won't get worse, so even if you assume a pessimistic rate of change, how different will things look soon?
The observation that these AIs are "mere text generators" is a testament to how insane they are IMO. The fact that they're just predicting the next token and are still this powerful is paradigm-shifting for how I viewed human intelligence. I don't think these things are sentient, but if they can replace me at work in 10 years or less, who cares?
Need I remind you of the AI winter?
Progress isn't linear. You want to believe progress will continue at its current pace. I'd suggest it's every bit as reasonable to assume progress will stall out as current techniques plateau, just as they did in the past.
But there's no reason to believe that this technique will plateau before some noteworthy results are obtained.
Other than being connected by the term 'AI' what makes you believe that the past performance of researchers is indicative of future performance?
Are any of the limitations of the Apollo program inherent limitations of space exploration? How are they connected?
Absolutely there is, and the article specifically touches on it.
Because these models do not encode semantics, there's very little reason to believe the hallucination problem can be solved with current techniques.
In fact, OpenAI themselves have been downplaying expectations about GPT-4, and my suspicion is that's because these techniques are already starting to plateau.
I fully expect that GPT-4 will create even more natural-seeming language, with greater sentence variation, less repetition, etc, and with no change (or possibly even an increase) in the frequency of factual errors in the results.
Let alone other hyped technologies that have failed to deliver, like self-driving cars, voice assistants, AR/VR...
Heck, the AI winter was preceded by a boom in things like expert systems, and I'll bet folks back then thought it'd never end, too.
I'm willing to put my money where my mouth is. I'm willing to enter into a $$$ wager and assert that AI, specifically LLMs or similar systems, will only become more important and more in demand in 5 years.
But I actually played around with chatgpt just today and holy shit.
I never ever talked to a bot like I talked to chatgpt.
Is it perfect? Not at all but I already was able to use it for practical things.
It feels tremendously natural.
This is so close to everything I'm looking for in an AI that it shows all the potential already without much fantasy.
This will not go away.
It will transform every aspect.
I wish I would be in university right now. This thing can explain things.
This thing understands context.
And it's multilingual out of the box.
The coherence chatgpt already shows let me believe that it does create an internal model which is more than just doing basic statistics.
Alexa was great to experience because it showed me already how a voice interface feels like.
But Alexa and chatgpt?
Chatgpt trained on documentation!
Chatgpt / ml like this will be the universal knowledge library the Rosetta stone.
We are experiencing the creation of the Alexandria library.
Let's see how long it will take to combine dall e/Stabel Diffusion, got, chatgpt together with other things like voice synthesis and voice recognition.
Using this is college would be a recipe for disaster, you'd ask chatgpt to explain a thing to you and you, being a student, don't know if it's getting the explanation correct or wrong. You don't know if it's reasoning is correct.
ChatGPT isn't telling you the truth or how the world is, it's just spitting out the statistically more likely next word.
It's far enough on this side of the uncanny valley so you think it's a conscious human unless you know enough about the subject to realize that it's not.
Perhaps "will not go away" will still seem plausible, but I think the rest will age poorly.
It told me the sidebar.js way.
I asked it to tell me another option which I can configure in the doc file directly.
Chatgpt got all of that right. If this is just a text prediction engine, it's crazy impressive how well it just works.
I don't even start designing the prompt to get what I need.
I definitely will follow all of this very closely and already have plenty of use cases for this 'text prediction engine '.
I'm hyped because I'm waiting for ml to break through for a while. Google for example has a huge us hospital under it's cloud platform with tons and tons of data.
We will see more and more expert systems.
Chatgpt will push for more investments.
This is not just a hype it's part of the journey we are for a few years now.
And in comparison to shit like crypto or nft there are already clear benefits.
Stabel Diffusion is fun to use and to explore.
Chatgpt is fun to use and to explore.
Not that this is any indication but the interest in SD was huge but the interest in chatgpt is much bigger. We had multiple internal talks about it and due to the interest in it, more got scheduled.
You will no longer search through documentation. You will just ask the product documentation what you need.
All of those real life use cases pushes more and more money into this topic
Must've been written by one of these "AI"s, then.
But consider that I didn't use Github Copilot only because I wasn't sure if it was free (and ChatGPT functioned well enough). This, even though I kept hitting limits in the system. It's still the case that for a lot of these use cases that people are using ChatGPT for that a chat interface is pretty inefficient for getting things done.
I do think the chat UI is probably overdone right now. But that's just because there's a multi-year usability learning process that still needs to happen. It's why Microsoft is building GPT into each of their products differently and others will do the same. Chat is a short-term design trend for many use cases, just for the moment. That said, the tech is expanding so quickly that we'll likely see more chat rather than less even as other UIs have even more room to grow.
I also don't have a very "pair programming friendly" way of thinking though/style. Maybe that plays a role.
I ask it to tell me about libraries or how a solution might get solved (recommend me known products out there, products as in like, popular open source software you might find off the shelf on GitHub) but that's about it. I can't see it telling me how to architect something at a high level or giving implementation details.
I've asked it about compiler errors before and how it might fix them, which has lead me to incorrect answers, but did help because I could Google/StackOverflow what was returned as it was typically "on the right track" after a few tries. But that was mainly me just "rushing" and not reading documentation/not having a good understanding of what I was working on (like being more familiar with Rust than C# for example)
ChatGPT: 24/7 pair programming partner, universal documentation wizard and instant Stack Overflow answerer. Also good at translating from lang A to lang B.
CoPilot: Autocomplete and suggestions on steroids, on the fly template and boilerplate generator and master refactorer.
But right now, it's so easy (and fun, I've done it myself) to spin up chatbots that there aren't enough folks doing the more difficult UI work for the various use cases, in my view.
"automated calculators", are calculators not insanely useful, especially earlier on in their history?
"but how many companies will be willing to jeopardize their reputation by giving their customers incorrect information?" Yet links to companies doing exactly that.
"carbon emissions" Yet datacenters all together put out 1% of total emissions globally, once again linked.
- - -
The article feels almost like a list of dotpoints which then get elaborated on, regardless of if they are notably relevant, la ChatGPT.
Again. This mania will subside just like the Clubhouse mania did.
LLMs can easily be trained on (token -> .. -> token, factual_accuracy) annotated data, we just need to build a "fact check" data set for common token sequences. That data set will be expensive to build, but we could always build a fact check app that gamifies the process to help build it instead of just paying turks.
However I wonder if having the LLM answer questions based on the knowledge baked in the model is the way forward at all: You quickly run into issues with up-to-date information, as you'd have to do a possibly expensive retraining every few weeks at least - and you also have no way of controlling which information exactly is stored exactly and how the models combines it.
I think an interesting alternative approach could be to split the bot up into two "agents" and one backend database:
- One model which receives user input and translates it into a machine-readable well-defined representation e.g. a series of SQL or SPARQL queries, a series of API calls or whatever.
- Then have a wholly non-AI backend that validates and executes those queries and delivers a result, in machine-readable format as well.
- Finally have a second model (or the first model with different finetuning/prompt) which can translate the query result back into a freetext sentence, which is then delivered to the user.
This way, you can keep tighter control about what exactly the model is answering and you can also update the underlying knowledgebase independently of the model.
https://writings.stephenwolfram.com/2023/01/wolframalpha-as-...
for example (butchered for clarity), (["The", "2022", "election", "was", "stolen"], accuracy=normal(0.0001, 0.1)).
Frankly, I think you're massively underestimating the difficulty in what you're describing (assuming what you're describing would actually work, and I'm not sure it would). To do what you're suggesting, you'd need a human to evaluate and feed into one of these LLMs every known fact. Good luck with that.
And that's ignoring the issue of maintaining that model as those facts change.
Data set annotation is expensive and time consuming, and it's a hard sell to shareholders to spend a billion dollars on data set annotation when there's still more data that can be cheaply obtained, and the liability or cost for producing incorrect answers is intangible or low. After Google's massive share hit from Bard's exoplanet flub though, I don't think the costs are quite as intangible as they were, and I expect accuracy will become a new push in models over the next few years.
I actually made three statements in my comment and you only disputed the first one.
Yours is a good counter-argument to my first statement, and it is a valid rebuttal, I concede the point!
But I maintain the scale of the problem, and the difficulty of maintenance, will make manual curation impossible, while automated "fact generation" is a technology that doesn't exist today.
Additionally, you don't have to manually curate the entire data set, only a large enough chunk to predict accuracy annotations for the stuff you haven't curated yet. Facts that change over time shouldn't be hard to identify if you periodically do cross validation on existing annotations to pick out ones where the predicted accuracy differs significantly from the annotation.
There are a lot of process and engineering questions still to be answered in terms of building high quality AI products, but I don't think there are hidden dragons that are going to lead to another AI winter. There will be hiccups and snafus, but AI is the real deal.
These kind of takes are getting tiring fast. I’d even go as far as to qualify them as midwit takes.
Yes you can reduce LLMs to their output channels, or trivialise what their doing by using words as « statistical ».
Facts: ChatGPT, a rough demo based on a model that is already obsolete, is nearly indistinguishable from the average college student.
https://twitter.com/lemonodor/status/1628270074074398720
https://twitter.com/lemonodor/status/1628420476136730624?cxt...
This headline may be the best indicator that there's actual fire under all the smoke for AI.
It might've seemed like there were more positive voices for web3 because many people stood to gain monetarily from it. It was more "democratized" in that way. With AI, there is a bigger moat and it's unclear whether there's big money to be made by anyone other than those who already own the underlying tech or have tens of billions to catch up.
Yeah, I think people underestimate the prevalence of this effect. At one point I tried figure out the origins of a story (covered by MSNBC) that the web3 game Axie Infinity was lifting people out of poverty in the Philippines. Every story was sourced back to the same report by a related party of the game.
(Axie ended up imploding catastrophically. I haven’t seen any stories about what happened to all the Filipinos dragged along into it)
You clearly weren't paying attention to the media over the past 5-10 years.
Saying that it doesn't understand words is ridiculous.
Prob it is a copy and paste from when gun powder was discovered.
> Write an essay about AI chatbots having their "tulip mania" moment.
> The Tulip Mania of the 17th century was a period in which tulip bulbs in the Netherlands were traded at exorbitant prices, resulting in a speculative bubble that eventually burst. The phenomenon is often used as an example of irrational economic behavior, with investors caught up in a frenzy of buying and selling without considering the actual value of the commodity they were trading.
> In a similar way, it is possible that AI chatbots could experience their own "tulip mania" moment in the future. Chatbots are becoming increasingly popular in many areas of business and communication, from customer service to healthcare. They offer the promise of cost savings, improved efficiency, and 24/7 availability. However, there are concerns that the hype surrounding chatbots could lead to unrealistic expectations and irrational investment.
> One potential trigger for an AI chatbot tulip mania could be a breakthrough in the technology that makes them significantly more capable. For example, if a chatbot was developed that could convincingly mimic human conversation and intelligence, it could generate a wave of excitement and investment. This could result in a rush to develop and implement chatbots in every industry, regardless of whether they are truly effective or necessary.
> Another potential trigger could be a high-profile success story of a chatbot that generated a significant return on investment. If a company or organization was able to demonstrate that a chatbot had dramatically improved their customer service or reduced costs, it could create a rush to invest in similar solutions. This could lead to a flood of new chatbots being developed and implemented without sufficient testing or consideration of their actual effectiveness.
> However, just as with tulip bulbs in the 17th century, the hype surrounding AI chatbots is unlikely to last forever. There are already concerns about the limitations of chatbot technology, including the difficulty of programming them to understand complex or ambiguous language, the potential for biases to be introduced, and the risk of chatbots making errors or providing inaccurate information. These limitations could become more apparent as chatbots become more widespread, leading to a burst of the hype bubble.
> In conclusion, it is possible that AI chatbots could experience their own tulip mania moment in the future, driven by hype and unrealistic expectations. While chatbots offer many potential benefits, it is important for investors, businesses, and developers to approach them with a critical eye and evaluate their actual effectiveness and value. Only by doing so can we ensure that chatbots fulfill their potential as a useful tool for communication and customer service, rather than a passing fad.
Comparing them to tulips just seems naive.
>Recorded before an audience at the Bristol Festival of Economics (11/17/2022)
> The Dutch went so potty over tulip bulbs in the 1600s that many were ruined when the inflated prices they were paying for the plants collapsed - that's the oft-repeated story later promoted by best-selling Scottish writer Charles Mackay. It's actually a gross exaggeration.
>Mackay's writings about economic bubbles bursting entertained and informed his Victorian readers - and continue to influence us today - but how did Mackey fare when faced with a stock market mania right before his eyes? The railway-building boom of the 1840s showed he wasn't so insightful after all.
> For a full list of sources used in this episode visit Tim Harford.com See omnystudio.com/listener for privacy information.
On the one hand, there is no way that I could just hand ChatGPT output straight to an editor.
At the same time--at least for certain types of stories--ChatGPT output is going to save some time relative to a blank page and it may even prompt you with some points you haven't thought of. You have to fact check and make it less formulaic and otherwise spruce things up. But I bet it could save me a couple hours on some articles.