Skimming the abstract as a senior academic in the area. This looks like preliminary work and a limited investigation for a single (non-standard) task. Thus far from a strong result published at say a top-tier conference or journal. Still, interesting direction and if expanded upon could absolutely be impactful. I should also mention that I am not familiar with the related literature, so it could very much be that there is similar (better?) work out there exploring the same question.
Honestly, this whole "they are not intelligent" argument is becoming ridiculous.
might as well argue that a plane isn’t a real bird or a car isn’t a real horse.
There's nothing perfect. Even computers and computer networks need to have error-correcting code because information gets randomly corrupted.
Our whole reality is probabilistic.
And us humans are way worse than AI at consistency. We even overwrite our own memories all the time, so we can't even be sure what we remember is actually what happened! (btw, this is currently being used in therapy to re-write traumatic memories and help people overcome PTSD).
https://www.npr.org/sections/health-shots/2014/02/04/2715279...
Even when I give it the benefit of the doubt, this sentence makes no sense to me. Do we have a list of tasks a language model can perform? To the best of my knowledge, they can arguably perform any language task.
> Large enough LLMs crush anything else for any NLP task. and evidently they beat top humans too.
Yes, they are certainly (rightfully) the go-to model for most tasks at this point if your concern is outright performance. Have I indicated otherwise? As for beating “top” humans, I am sure that can be investigated, but it is a fairly nuanced research question. It is inarguable that they are amazingly good though, especially relative to what we had just a few years ago.
> Honestly, this whole "they are not intelligent" argument is becoming ridiculously obtuse. > > might as well argue that a plane isn’t a real bird or a car isn’t a real horse.
Which is a claim and argument that I never made – hallucinating? How about you calm down a little and get back on the ground? You are talking to someone that has argued in favour of these kinds of models for about a decade. But that does not mean that I am willing to spout nonsense or lose track of what we know and what we do not yet know.
no-one wants to call a spade a spade yet but the sentiment is obvious in recent research. directly being called General purpose technologies from the jobs paper, general artificial intelligence from the creativity paper. That last one is particularly funny, they just switched the two words.
They aren’t though… They are far superior at specific things birds and horses are known for, but they can’t do everything that birds and horses can, so they aren’t even artificial birds and horses.
For the people who are now salivating at their potential, or dreading the possibility of being made redundant by them, these large language models are already intelligent enough to matter.
Handwringing bout some non-existent difference between "true understanding" and "fake understanding" which by the way nobody seems to be able to actually distinguish (I mean wow such a supposed huge difference and you can't even show me what that is. a distinction you can't test for is not a distinction ) is so far beyond the point, it's increasingly maddening to read.
It’s clear that at the least, they can decipher very numerous patterns across a wide range of conceptual depths — it’s an architectural advance easily on the the level of the convolutional neural network, if not even more profound. The idea that NLP is “solved” isn’t a crazy notion, though I won’t take a side on that.
That said, it’s equally obvious that they are not AGI unless you have a really uninspired and self-limiting definition of AGI. They are purely feedforward aside from the single generated token that becomes part of the input to the next iteration. Multimodality has not been incorporated (aside from possibly a limited form in GPT-4). Real-world decision-making and agency is entirely outside the bounds of what these models can conceive or act towards.
Effectively and by design these models are computational behemoths trained to do one singular task only — wring a large textual input though an enormous interconnected web of calculations purely in service of distilling everything down to a single word as output, a hopefully plausible guess at what’s next given what’s been seen.
And you want to know the crazier thing? Evidently a lot of researchers feel similarly too.
General Purpose Technologies ( from the Jobs Paper), General Artificial Intelligence (from the creativity paper). Want to know the original title of the recent Microsoft paper ? "First contact with an AGI system".
The skirting around the word that is now happening is insanely funny. Look at the last one. Fuck, they just switched the word order. Nobody wants to call a spade a spade yet but it's obvious people are figuring it out.
I can you show you output that clearly demonstrates understanding and reasoning. That's not the problem. The problem is that when I do, the argument Quickly shifts to "it's not true understanding!" What a bizzare argument.
This is the fallacy of the philosophical zombie. Somehow there is this extra special distinction between two things and yet you can't actually show it. You can't test for so called huge distinction. A distinction that can't be tested for is not a distinction.
The intelligence arguments are also stupid because they miss the point entirely.
What matters is that the plane still flies, the car still drives and the boat still sails. For the people who are now salivating at their potential, or dreading the possibility of being made redundant by them, these large language models are already intelligent enough to matter.
I'm definitely not contesting that.
I've always considered the idea of "AGI" to mean something of the holy grail of machine learning -- the point at which there is no real point in pursuing further advances in artificial intelligence because the AI itself will discover and apply such augmentations using its own capabilities.
I have seen no evidence that these transformer models would be able to do this, but if the current models can do so do then perhaps I will eat my words. (Doing this would likely mean that GPT-4 would need to propose, implement, and empirically test some fundamental architectural advancements in both multimodal and reinforcement learning.)
By the way, many researchers are equally convinced that these models are in fact not AGI -- that includes the head of OpenAI.
AGI went from meaning Generally Intelligent to as smart as Human experts and then now smarter than all experts combined. You'll forgive me if I no longer want to play this game.
I know some researchers disagree. That's fine. The point I was really getting at is that no researcher worth his salt can call these models narrow anymore. There's absolutely nothing narrow about GPT and the like. So if you think it's not AGI, you've come to accept it no longer means general intelligence.
Are you talking about large language models (LLMs)? Because those are narrow, and brittle, and dumb as bricks, and I don't care a jot about your "No True Scotsman". LLMs can only operate on text, they can only output text that demonstrates "reasoning" when their training text has instances of text detailing the solutions of reasoning problems similar to the ones they're asked to solve, and their output depends entirely on their input: you change the prompt and the "AGI" becomes a drooling idiot, and v.v.
That's no sign of intelligence and you should re-evaluate your unbridled enthusiasm. You believe in magick, and you are loudly proclaiming your belief in magick. Examples abound in history that magick doesn't work, and only science does.
I'm an old hat hobby programmer that played around with ai demos back in the mid to late 90s and 2000s and chatgpt is nothing like any ai I've ever seen before.
It absolutely can appear to reason especially if you manipulate it out of its safety controls.
I don't know what it's doing to cause such compelling output, but it's certainly not just recursively spitting out good words to use next.
That said, there are fundamental problems with chatgpt's understanding of reality, which is to say it's about as knowledgeable as a box of rocks. Or perhaps a better analogy is about as smart as a room sized pile of loose papers.
But knowing about reality and reasoning are two very different things.
I'm excited to see where things go from here.
I'm not sure where this idea of LLMs being intelligent even comes from. It took me a whopping 9 prompts (genuine questions, no clever prompt engineering) of interacting with ChatGPT to conclude it does not understand anything. It doesn't understand addition, what length is, doesn't remember what it said a second ago, etc.
The output of ChatGPT is clearly just a reflection of its inner workings - predicting the next word based on training data. It's clever and undoubtely useful for a certain set of repetitive problems like generating boilerplate but it's not intelligence, not by any reasonable definition.
There once was a dog from the pound
Whose bark had a curious sound
With a wag and a woof,
He'd jump on the roof,
Delighting the folks all around.It's like saying 4 months after the first useful car was manufactured. "If these are so good, how come there are still horses? Clearly the market disagrees with you".
[1]: https://arxiv.org/abs/2302.04761
Other research questions, but without concrete papers to reference to that current keeps me up at night: 1.) to which degree can we train substantially smaller LLMs for specific tasks that could be run in-house, 2.) it seems like these new breakthroughs may need a different mode of evaluation compared to what we have used since the 80s in the field and I am not sure what that would look like (maybe along the lines of HELM [2]?), and 3.) can AI academia continue to operate in the way it currently does with small research teams or is a change towards what we see in the physical sciences necessary?
I don't see why this follows. We do loads of stuff other than language, it is entirely possible for an AI to be better than us at language but worse at everything else.
The point that would say to me the LLM actually has any "understanding" of what it's saying would be when it's able to reliably say "I don't know the answer to that" instead of making up things from scratch. You see that a lot if you ask Bing/Bard "Who is _____?" Most of them are kind of right but a lot of large details are just completely fabricated. A lot of the facts it gets wrong are things Google is already able to produce when queried like where was Person X born or where did they go to school so the fact these LLMs can't slot in actual available facts says to me they're not really going to be that useful with the kind of tasks we've been working on NLP for.
Both Bing chat and ChatGPT plugins are examples of being able to do just this.
You’re right about how they make up answers though, but humans are often quite prone to that too…
Granted getting the name for something to search is often half the battle in tech.
In fact GPT-4 is quite good at catching hallucinations when the question-answer pair is fed back to itself.
This isn’t automatically applied already because the model is expensive to run, but you can just do it yourself (or automate it with a plug-in or LangChain) and pay the extra cost.
Remember that the model only performs a fixed amount of computation per generated token, so just asking it to think out loud or evaluate its own responses is basically giving it a scratchpad to think harder about your question.
They can: https://arxiv.org/abs/2302.04761
> In this paper, we show that LMs can teach themselves to use external tools
I'm sure it'll be doing that in five years, but not now.
One interesting thing is that's it's nondeterministic, so sometimes 'For chrissakes' turns to ちくしょう (Damn!) but sometimes to クリスのために (for Chris' sake). Sometimes 'the goddamn door' turns into クソドア ('shit door'), sometimes the goddamn changes the phrasing of the whole sentence instead. If you run it five times and take the best sentences out of all five runs it's probably quite good. Maybe prompting would help too, I said "idiomatic Japanese" but it still usually translated it in a very "foreigner Japanese" way typical of US drama/movie translations.
In "Can GPT-4 translate literature?" (Mar 18, 2023) [https://youtu.be/5KKDCp3OaMo?t=377], Tom Gally, a former professional translator and current professor at the University of Tokyo, said:
> …the point is, to my eye at least, the the basic quality of the translation [from Japanese to English] is as good as a human might do, and with some relatively mild editing by a sensitive human editor, it could be turned into a novel that would be readable and enjoyable.
Like he says, there's lots of ambiguity in Japanese that needs to be handled, gender not being specified until later, etc. and an editor would need to spend time going over it - but it saves months of traditional work. There are words and _concepts_ that are hard to translate, there are cultural issues, dialects, slang, registers. So yeah it'll make the media accessible, but it won't be as a good as a skilled translator.
I mean i can if you want (chinese though) but enough people lie on the internet.
I'm a native speaker of Chinese, but not a professional translator.
and you can see talk on English/Japanese here - https://youtu.be/5KKDCp3OaMo?t=377
This.
GPT/ChatGPT is able to even translate between different "accents" or dialects of the same language. For example it can give you the same sentence in Mexican, Argentinean or Peruvian Spanish.
Example:
Me: Give me a sentence in spanish that is different in Mexican, Argentinean and Peruvian Spanish. Write the sentence in each dialect.
ChatGPT: Mexican Spanish: "¿Qué onda, güey? ¿Cómo estás?" Argentinean Spanish: "¿Qué onda, boludo? ¿Cómo estás?" Peruvian Spanish: "¿Qué tal, causa? ¿Cómo estás?"
These sentences mean "What's up, dude? How are you?" in English. The primary difference is the slang term used for "dude" or "friend" in each dialect: "güey" in Mexican Spanish, "boludo" in Argentinean Spanish, and "causa" in Peruvian Spanish.
Similar to how a lot of swear words work in many languages.
It's interesting to see how what matters is not the word, but the intention behind it. At the end we are trying to communicate meaning, and words are just one of our tools to do it.
Of Course i don't know the data for GPT-4
Mandarin isn't the second language I speak but I officially compared with it because I wanted to test also with a model that had more equivalent corpus training than the very lopsided gpt models. And Chinese/English is the only combo that has a model of note in that regard.
For almost everything else, it seems to produce pretty decent and usable translations, even when used against relatively obscure languages.
I used it a some green landic article that was posted on hn yesterday (about Greenland having gotten rid of daylight saving time). I don't speak a word of that language but the resulting English translation looked like it matched the topic and generally read like correct and sensible English. I can't vouch for the correctness obviously. But I could not spot any weird errors or strange formulations that e.g. Google translate suffers from. That matches my earlier experience trying to get chat gpt to answer in some Dutch dialects, Frysian, Latin, and a few other more obscure outputs. It does all of that. Getting it to use pirate speak is actually quite funny.
The reason that I used Chat GPT for this is that Google translate does not understand greenlandic. Understandable because there are only a few tens of thousands of native speakers of that language and presumably there's not a very large amount of training material in that language.
Therein lies the rub. There's a huge gap between what LLMs can currently do (spit back something in a target language that gives you the basic idea, however awkwardly phrased, of what was said in the source language). And what is actually needed for idiomatic, reasonably error-free translation.
By "reasonably error-free" I mean, say, requiring a human correction for less than 5 percent of all sentences. Current LLMs are nowhere near that level, even for resource-rich language pairs.
I ran the abstract of this article through chat gpt. Flawless translation as far as I can see. To be fair, Google translate also did a decent job. Here's the chat GPT translation.
Veel NLP-toepassingen vereisen handmatige gegevensannotaties voor verschillende taken, met name om classificatoren te trainen of de prestaties van ongesuperviseerde modellen te evalueren. Afhankelijk van de omvang en complexiteit van de taken kunnen deze worden uitgevoerd door crowd-werkers op platforms zoals MTurk, evenals getrainde annotatoren, zoals onderzoeksassistenten. Met behulp van een steekproef van 2.382 tweets laten we zien dat ChatGPT beter presteert dan crowd-werkers voor verschillende annotatietaken, waaronder relevantie, standpunt, onderwerpen en frames detectie. Specifiek is de zero-shot nauwkeurigheid van ChatGPT hoger dan die van crowd-werkers voor vier van de vijf taken, terwijl de intercoder overeenkomst van ChatGPT hoger is dan die van zowel crowd-werkers als getrainde annotatoren voor alle taken. Bovendien is de per-annotatiekosten van ChatGPT minder dan $0.003, ongeveer twintig keer goedkoper dan MTurk. Deze resultaten tonen het potentieel van grote taalmodellen om de efficiëntie van tekstclassificatie drastisch te verhogen.
Translating the Dutch back to English using Google translate (to rule out model bias) you get something that is very close to the original that is still correct:
Many NLP applications require manual data annotations for various tasks, especially to train classifiers or evaluate the performance of unsupervised models. Depending on the size and complexity of the tasks, these can be performed by crowd workers on platforms such as MTurk, as well as trained annotators, such as research assistants. Using a sample of 2,382 tweets, we show that ChatGPT outperforms crowd workers for several annotation tasks, including relevance, point of view, topics, and frames detection. Specifically, ChatGPT's zero-shot accuracy is higher than crowd workers for four of the five tasks, while ChatGPT's intercoder agreement is higher than both crowd workers and trained annotators for all tasks. In addition, ChatGPT's per-annotation cost is less than $0.003, about twenty times cheaper than MTurk. These results show the potential of large language models to dramatically increase the efficiency of text classification.
I'm sure there are edge cases where you can argue the merits of some of the translations but it's generally pretty good and usable.
I will be re-assessing my view on general-case translation performance accordingly.
- Retrieval augmented in-context learning
- Better benchmarks
- Last mile for productive application
- Faithful, human-interoperable explanations
It seems like it's a solved issue indeed. AI reasoning has still some way to go, but it seems language understanding is a finished subject.
Everyone shrugs and says, “nope, humans are different”. I’ve commented about 100 times recently asking for detail as to how human language / thought works, yet have not seen an answer.
That's basic linguistics and cognitive psychology. Nothing an LLM has done has invalidated that.