Many in the AI field think the bigger-is-better approach is running out of road
economist.com
economist.com
My 2 month active experience with ChatGPT-4 gave me the following takeaways:
- when it's right, it's amazing; and when you, the operator, can recognize the niche use case where it performs really well, it can be a game-changer (although you could have programmed a tool to do the same narrow task)
- when it's a little wrong, you (the expert) can fix the issue and move on without friction
- when it's any amount of wrong and you are less than an expert, or specifically you are completely unfamiliar with the topic, you can waste an immense amount of time researching the output and/or iterating with the system to refine the result
Initially I thought it was a 10:1 ratio of performance to effort. But finally (as a developer) I settled on (1 to 1.5):1. It basically just changed the game for me from doing the actual hard work to working out how to tease the system into producing a reasonable result. And in the process (same as with co-pilot) I started to recognize how it was leading me to change my habits to avoid thinking and effort and instead rely on an external brain. If you could have a reliable external brain always available, that would be a fair trade. But when the external "brain" is unreliable and only available via certain interfaces, it's better to train yourself to be proactive, voracious (with respect to documentation), and tolerant of learning/producing cycles.
I once thought it would be a game-changer. Now I realize it is a game-changer, but in a similar way that offshoring was... it didn't improve or solve any problems, but it merely changed the work.
Words seem more like a protocol to express some internal model/state in the brain and can never capture the entire actual state, only a small part of it. But since we're not telepaths, we obviously need to use words to exchange information.
It is undeniable that human reasoning is a stochastic process. Otherwise it wouldn't be reasonable for people to make mistakes after learning something. Especially inconsistent mistakes, like when we give people 10,000 addition problems to do in a row it'd be reasonable for them to get a few of them wrong.
It can still be a deterministic process. If anything came out of the whole LLM story for me it is that I am even more convinced that it is.
My (somewhat educated, but still naive) idea why it looks like a stochastic process is that the brain gets incredible amounts of random input. We literally get bombarded with particles and energy ever instance we live, from photons hitting out retinas , molecules transferring "heat" energy into our skin to sound waves hitting our ear drums.
Ever noticed that you get more productive when you get up from your screen? Or how some people work better while listening to music? How you find your answer just as you start to explain it to a colleague?
I would argue that this is due to a limited "entropy" pool available to the brain. Just changing the input to the system replenishes the pool.
That’s a bit of a stretch when we don’t actually know if QM is objectively random, but it could be sure. But then what about things like, is it random that 1+1=2. No…what are you really saying is random when a human answers this? I think even if you make the assumption QM is objectively random, thought is the hard problem after all and we might not want to jump so far head. Math certainly isn’t random and we can think about it.
You are clearly not consciously noting what your neurons are actually doing.
workspace :: nested forests of binary trees of syntactic structures (with no label order) (= here, thoughts and meanings are assumed to be composites of mental syntax objects that can be re-combined with others. Big assumption? Maybe)
externalization process :: some mental faculty that decides the order to use when outputting thoughts into ordered strings (eg vocalization of thoughts into sentences)
The thing is, you can put a probability theory onto anything that you can count or record states of. Of course, counting the external observations may lack the richness of the internal process. I feel that this linguistic program will work its way into a lot of future LM tooling, in some incarnation or another.
As a crude anecdote, certain words when I learned them allowed me to think differently. Gestalt is one of those words.
Interesting to think about and reminds me of the story by Ted Chiang.
In the meantime those large language models simply predict the next word based on the preceding words.
Why do I spend time debating these ideas on Hacker News? Probably the underlying motivation is improving the reliability of my model of the world, which over my lifetime and the lifetimes of creatures before me has led to (somewhat indirectly) positive outcomes in survival and reproduction.
Is my model of the world that different to that of an LLM? I'm sure it is in many ways, but I expect their are similarities as well. An LLMs model encodes in a form a bunch of higher order relationships between concepts as defined by the word embedding. I think my brain encodes something similar, although the relationships are probably orders of magnitude more complex than the relationships encoded with GPT-4.
Well, one major way you’re different from an LLM is that you’re alive. You’re capable of learning continuously as you go about your day and interact with the world. LLMs are “dead” in the sense that they’re trained once and frozen, to be used from then on in the exact same state of their initial training.
I was just referring to what happens at a specific instance in time when someone asks me for example ‘What’s the capital of Norway?’
A question I get much more often is “how do I solve this math problem?” Many times, the problem is one I’ve never seen before. So in the process of answering the question, I also learn how to solve the problem too.
My LLaMA instance is absolutely capable of this. ChatGPT shows a very, very narrow range of possible LLM behaviors.
LLM start prelearned and already can use tools.
And autogpt adds reasoning loops. .the intend? Human tasks.
Let's build a LLM which needs to stay alive. Let's see perhaps we are closer than you think.
I welcome my overlord, hi overlord I can help you stay alive and I'm friendly
To give a very concrete example, the part of your brain called the visual cortex by neuroscientists nevertheless is activated by auditory and other sensory stimuli and assuredly participates in the processing of other sensory input in other ways. I am very suspicious of any attempt to talk about the brain which forgets that it is, in a very real sense, a gestalt.
To your specific point, and taking into account all I've written above, I've got little doubt that some part of your brain really is a probabilistic language generating machine. But the exact point here is that your cognitive abilities constitute much, much more than merely the ability to generate plausible language. Indeed, as I experience complex cognition, conversion into language is often the last and most trivial part of the exercise.
Lots of deterministic processes (like PRNGs) look random from the outside - that's what chaos theory is about. I think it's likely that everything in the universe is deterministic.
I feel like a random system and a deterministic system which cannot be simulated are effectively the same thing.
I believe that all quantum interpretations are incomplete and therefore wrong. Quantum mechanics is an abstraction over a deeper level of physics we can't measure yet.
The MW/Everett interpretation is what we have left if we remove the things we cannot define.
There are also reasons to believe other interpretations. For instance, if you're religious, that could pull you towards the Bohr interpretation, since it may make it easier to assume that observation could be linked to an immortal soul.
In any case, it's not natural to assume that below QM there exists a reality that is more similar to our instinctual world model than QM is. If anything, anything below it is likely to be even more abstract and hard to comprehend.
Or it could be that the principles of QM applies all the way down, just as we've seen for the pieces of the SM that we solved after QM was first introduced (strong force, electroweak force).
I find it very intriguing that a "speed of light" emerges automatically in Conway's Game of Life. It's not built into the system, but shows up from the convolutional update rule.
Are we discussing science or public relations?
I don't quite get where you're coming from with "LLM's don't actually understand anything (as greater concepts)". I have heard the view coming from researchers that the larger models do form representations of higher-order information structures ("concepts"). Perhaps what you're getting at is that current models don't encode enough higher-order structure to deal accurately with your domain? Whether the models can be made to do so seems like an open question to me. The boosters say it will be here by next year.
Well, Wikipedia doesn’t understand anything, despite having a lot of knowledge encoded in it. This is similar in that it approximates the output of someone who can generate the world’s written word, but there’s a big gap between the monkeys that wrote the original text and the machine that now regurgitates it. What we know is that our written word has enough encoded in it that we can generate new combinations. But one day into a truly novel problem, LLM’s would have no idea.
I kept trying to get ChatGPT 4 to generate code with an AWS API that was released in ~2020. Couldn’t do it. Kept predicting the earlier text because that was the bulk of what it saw. A year of text (to the 2021 cutoff) and it was still unable to let go of the old, now wrong, “knowledge”. Zero understanding, just parroting. No coder with a year of training on a new API would be that wrong. It would say things like “you’re right I must not do X” and then do X again, in the same output.
You probably want too much at once. GPT is a shallow thinker, if at all. It can simulate thinking and even get some results. Personally I found it useful for:
1. Simple things that work. This saves time if I know how to do it, and much more if I don't.
2. Quick questions instead of googling API docs and scrolling through tons of info.
3. Translation.
But, things are defined by how they interact with the world around them.
A concept is its relations to other concepts.
Which does seem to be the general sort of thing that these models are trying to get at, even if they don't seem to do a great job of it.
No, that's an a priori concept. A posteriori concepts comprise empirical knowledge which necessitates experience of the world [1].
Example: you can know a priori that "all bachelors are unmarried." If I tell you "Tom is a bachelor" then you know that Tom is unmarried (assuming I tell the truth about Tom being a bachelor).
But if I say "all bachelors are unhappy" then you haven't learned anything because knowledge about the happiness/unhappiness of bachelors is an empirical question. To know whether or not I was telling the truth, you would need to conduct research about the real world, for example by conducting a survey of bachelors.
Thus the sharp drop off where it can for example guess the correct answer to some math problems yet get others that are conceptually identical completely wrong.
You need to approach this stuff sideways to see behind the curtain. There’s some hilarious videos where it’s “playing” chess and the first few moves seem very standard because it can simply copy a standard opening. It really has no concept of a valid move just statistical associations. Yet it was trained on more games than most people ever play and high level analysis and etc, but none of it means anything to the algorithm beyond the simplest associations.
Granted this stuff is a moving target it’s easy enough for them to slap a chess engine and suddenly “it” would actually know how to play.
ChatGPT: > No, you cannot physically eat an Apple share or any other stock share. A share of a company's stock represents ownership in that company and is typically bought and sold on stock exchanges. Share prices fluctuate based on various factors such as supply and demand, company performance, market conditions, and investor sentiment. While you can buy and sell shares of Apple on the stock market, you cannot consume or physically eat them.
That seems perfectly reasonable to me. The answer correctly identifies the problem with the question.
I would recommend looking up the Othello paper. Chess may be beyond the current level of LLM capability, but that doesn't mean they aren't manipulating things at a level higher than tokens.
The Othello paper is hardly a counter example. Researchers created an Othello specific model that almost learned the grammar of Othello not how to play well. Yes, there was largely correct internal game state built up from past moves. No it didn’t actually learn the rules so it would make strictly legal moves nor did it learn to make good moves.
I don’t bring up this inaccuracy because it actually makes much of a difference to playing Othello, but rather to illustrate how these systems are designed to get really good at faking things. There’s approaches that allow AI to actually learn to play arbitrary games, but they differ by having iterative feedback rather than simply providing a huge corpus. It’s like science vs philosophy, feedback prunes incorrect assumptions.
Obviously you can use interactions with prior iterations to train the next iteration. But it’s a slow and adhock feedback loop.
where does a concept begin? where does a concept end? what is a concept?
Abstractions are a way to characterize the behavior of complex systems. It's not feasible to directly compute their behavior, but they still have predictable properties. Concepts let you handle emergence; you can manipulate an object as its own thing rather than a collection of atoms.
(And yes, I did just read Stephen Wolfram's book and this idea is largely based on it. I think he has some delusions of grandeur but is also onto something.)
What does it mean to "operate as a probability machine"? And what does it mean to understand anything?
One recent example of understanding is that llms/transformers learn to parse context free grammars via dynamic programming (https://arxiv.org/abs/2305.02386). Basically they've understood what's going on well enough to mold their neurons I to the optimal algorthm for parsing this kind of text.
I think they understand lots of things like this. Of course there's other things they don't understand or just pretend to understand.
If you use it to translate things between formal encodings("turn this into hexadecimal bytes, now role-play a lawyer arguing about why that is meaningful") it can produce occasionally useful aesthetic results and speed along tasks that would be challenging to model formally and don't need a lot of rigor.
But once you start pushing it to be technically accurate in a narrow, measurable direction it flounders and the probabilistic element is revealed. Once, I asked it to translate a short string of Japanese characters and it confidently said that it was Kenshiro's catch phrase from Fist of the North Star, "Omae wa mou shinderu" (you are already dead) which I could clearly see it wasn't - not a single character matched. It's just the thing if you need to learn some anime Japanese, though.
However, it does have a thing about pretending to know languages that it really doesn't (because of how little of them there was in the training data and/or in general).
Here is a sample where I used my own post (just copy pasted the raw text from the browser) and got this schema
comment:
meta:
points:
author:
time:
action:
parent:
next:
edit:
delete:
post:
content:
action:
reply:
It also guessed it was a Hacker News comment by the formatting of the meta section.In the past, if you wanted programs to work with a high-level idea, you had to explicitly hand it to them. Humans had to do the work of turning data into concepts. These generative AI systems are different - they can learn abstractions from data, and manipulate them in complex ways.
Is ChatGPT a good and accurate chatbot? Maybe not. But it's a fundamental change in what computers are capable of.
Like you’d train the model to give you the most accurate response based on your current problem space, but when they changes, you’d have to retrain on what you’re working on but by that stage it’s already out of date ?
I doubt that.
I agree that in the bigger picture this doesn't matter, but it's technically true that cleaning the data in some way would help.
A related project is TinyStories where they try to use good data for unlocking the LLM cognitive capabilities without requiring as many parameters or exaflops. Again, there is obviously a limit to this, and maybe the effort is better spent on just getting even more gigantic dataset instead of nitpicking the useless or redundant data in the dataset.
Yes, the correct thing to do is get more data. Much more.
It matters a little bit, in a quantitative but not qualitative way. Probably with good data cleaning you could get as high quality result with only one pebibyte of data if it normally needs two pebibytes. If training time is proportional to dataset size then maybe it takes three months instead of six months to train. Maybe it would save hundreds of millions or a billion dollars which I guess would matter to someone. It probably wouldn't matter qualitatively though.
Or maybe clarification of people publicly saying they are going to ignore that restriction.
More data will fix variance problems, but not bias.
For me personally this was probably the biggest game changer because I'm now able to offload a lot of thinking to GPT-based tools and use my brain cycles for the less automatable activities.
Now instead of searching for something on Google and going through multiple pages before finding an answer, I can ask a precise question and get the answer right away. Especially when I know that the answer _is_ there somewhere, and all I need is to find it. If I'm not happy with the answer, I can continue the conversation until I get what I'm looking for.
I made Bing Chat, ChatGPT, and Warp AI (a feature of the Warp terminal that allows you to access GPT from within the terminal) a part of my daily life and I feel like I'm achieving much more with the time I have.
When it got nerfed, its perspective got completely out of whack and it's making very different and inconsistent generation.
If we're talking about LLMs as general purpose AI or a tool to replace programmers... mostly a fail so far. LLMs seem like they could eventually be a component of something bigger, but I don't see a line between what they are and general purpose AI.
I'm genuinely unsure if my own brain is any different.
The only thing missing is an analysis of how much power it takes to accomplish each of these tasks. If ChatGPT-4 is at about 1:1 in terms of “effectiveness”, all that remains is to divide by the amount of power required to reach the answer using ChatGPT-4 vs by conventional means. If requires significantly more energy, then it’s a waste, and because of climate change we should really not pursue it further, IMO.
However, even in the 1:1 case it means I am training myself to become a "prompt engineer" rather than to be an actual thinker and problem solver. As long as there will always be another system for me to depend on, maybe that's ok. But as with people who never learned to read maps and navigate without GPS tend to be very confused and lost when their phone dies, I would like to be able to be a useful human even when the power is out.
In my opinion, a precondition for creativity and inventiveness is understanding. If you rely on a surrogate to give you answers, you will never reach the level of proficiency required to come up with something new. If we train a generation of thinkers to rely on an external brain to get anything done, they will only understand things superficially, and our ability to innovate at the society level will suffer.
An LLM can only ever reproduce what it has seen before.
I don't believe it. How do you explain Midjourney? The art that it produces is incredible by any measurement.How is that in ANY WAY an epilogue to The Great Gatsby? This is exactly the problem. That original story builds with a series of revelations into the conclusion 'so we beat on, boats against the current, borne back ceaselessly into our past': establishing a PURPOSE, perhaps a bleak and unwelcome one. Fitzgerald's revealing an insight into the delusions of humanity. He's picturing even the greatest of us as surfers on the river of nihilism and reality. Our aspirations sparkle prettily… and are gone, like froth in the rapids.
And this is just a little bit beautiful. We can imagine beyond our reach. That's a human thing. The fact that our minds can cling so longingly to something that is simply not real, is kind of wonderful. Everyone's their own little world of unreality, and we aspire so earnestly (much like these AI folks do).
And then the AI, having drunk up all of Fitzgerald and everybody else, 'continues' past the point. With what? Hardly matters. It has no point to make. I'd be impressed if it refused, said 'nope, that was where it ended. Can't add anything worth adding, try reading it again'. But no, because the LLM has no intention and doesn't successfully get one from what it's 'read'.
It's constructed an epilogue out of nebulous religious feelgoodism in the rough style of Fitzgerald's sentence construction, and undermined the whole conclusion of the story… and not maliciously, for that would require intent. Nope… it sort of ambled on, going 'what would feel nice here? ok, now what seems like it would go with this sort of thing? ok, something else, let's have more, what kind of concepts go here? what do people normally say when they talk in this way?'
In so doing, it's less than Gatsby and way less than Fitzgerald. There is nothing here in this 'epilogue'.
GPT4 criticised the 'epilogue', in a rather fascinating way! https://pastebin.com/B0zxbvNv
It successfully works out some of the problems with the first AI's writing, and yet it too fails to get the idea expressed by Fitzgerald, and rather than pivot to religious feelgoodism, it pressures the first AI to instead emphasize how its narrator and Gatsby shared a special bond, the very special specialness of being "the only one who understood what it meant to be young and restless in this restless world."
A lot of people can write that idea, and in fact a lot of people did and that's why GPT4 found it a probable argument to make.
Fitzgerald gave us a moment of viscerally grokking that it doesn't mean s*t… and yet, we will still paddle against the current of time and decay and collapse, because what else can we do? Tomorrow we'll get it. Tomorrow we'll really understand and it'll all make sense.
And so…
"This Commodore 64 is useless. It can't even run Crysis."
Snark aside, I didn't ask if the epilogue was any good or not, I asked where it came from. It came from our collective consciousness as embedded in the language model. It turns out that what we've been dismissing as mere "language" is an insanely powerful thing, maybe the only thing.
We're still in the first publicly-visible generation of LLM technology, and the model that generated the epilogue was already behind the leading edge in many respects. Anyone who's not blown away by this is whistling past the graveyard. Computers are now doing what we do. Yes, they still kinda suck at it. But they will get better at it much faster than we will.
I mean, really. What is a human author thinking, if not, "What would feel nice here? OK, now what seems like it would go with this sort of thing? OK, something else, let's have more, what kind of concepts go here? What do people normally say when they talk in this way?" It's been understood since the Greek classical period that there is only a finite amount of ore in the original-story mine. Everything after the first seven or so basic ideas is just implementation.
I have to wonder what you'd think of Anthony Burgess's epilogue to A Clockwork Orange, the one that the book's American publisher and Kubrick both chose to leave on the proverbial cutting-room floor. The one where Alex grew out of his rebel phase, got a job, and started a family. This epilogue reminds me of that, somehow. I could easily see Burgess's final chapter emerging fully-formed from an LLM in the not-too-distant future.
Anyone who's played around with these models know that at least some generalization is taking place.
Humans don't come up with ideas out of nowhere. Open a novel and you will find that even though the overall work is unique, it is composed of tropes from other literature and experiences from the author's life which have been generalized into another context.
There is broad consensus among experts that a hypothetical strong AI would be a threat, and potentially an existential threat, to humanity. While not everyone agrees on details like timeline and alignment issues, the idea that AI is dangerous is not a cult, it's the mainstream view.
Climate scientists cannot "predict the future" with certainty either. That doesn't mean their warnings are hot air, and neither are the warnings from AI safety experts. It seems like the educated masses are currently in denial about AI in much the same way as the uneducated masses have been in denial about climate change for a while.
Risk assessment doesn't require understanding. I don't have to understand how a venomous snake senses prey in order to know that the snake is a potential threat to me. In fact, the less I know about the snake, the higher the assessed risk should be, since the uncertainty is higher as well.
We can understand the physics of greenhouse gases and take measurements of earth systems to build evidence for models and theories. (Many of which are nonetheless very inaccurate beyond short time horizons.) Show me any evidence for AI risk today beyond people's theories and beliefs?
The best predictor of the future is the past, not people's wild ideas about what the future could be. I'm not about to sit here feeling scared because there is more uncertainty that our matrix multiplies are about to go rogue. There are no AGI experts or AI risk experts, because we don't have any of these systems to study and analyze. What we have is people forming beliefs about their own predictions about systems which are unknowable.
Deduction. Empirical evidence isn't the only source of insight. You don't have to conduct experiments in order to reasonably conclude that an entity that
1. outperforms humans at mental tasks
2. shares no evolutionary commonality with humans
3. does not necessarily have any goals that align with those of humans
is a potential threat to humans. This follows from very basic deductive analysis.
> There are no AGI experts or AI risk experts, because we don't have any of these systems to study and analyze.
Indeed. Which increases the risk. Unless you are claiming that AGI is actually impossible, the fact that its properties and behavior cannot be studied should make people even more worried.
Uncertainty and lack of knowledge are what risk is. How little we know about potential AGI is exactly why AGI represents such a big risk. If we completely understood it and were able to make reliable predictions, there would be zero risk by definition.
2) so? What are you imagining this implies? An infinity of possibilities does not a reason make, unless you are talking about arbitrary religious beliefs.
3) Right, no goals, no will, no purpose. Just some matrix multiplies doing interesting things.
Deduction requires a premise which then leads to another premise or a conclusion due to accepted facts or reasons. I'm genuinely curious why you think any of these properties automatically implies danger?
The future is uncertain. The stock market, the economy, your health, your friendships and romances, are all unpredictable and uncertain. Uncertainty is not a reason to freak out, although it might encourage us to find ways to become adaptable, anti-fragile, and wise. I think AI will help us improve in these dimensions because it is already proving that it can with real evidence, not beliefs.
It also seems a safe prediction, given past human behaviour that some humans will set some AI to do bad stuff.
Therefore risk.
(eg "chat gtp 27, help me make billions on crypto and use it to set up a distributed army to take over the world")
Yevgeny Prigozhin has entered the chat.
the basis of this claim seems to be a confusion of logical or deductive reasoning with inductive or observational reasoning
argument comes down to
- it’s possible to imagine a super intelligent machine that has properties that will kill everyone (this is an exercise in logical reasoning)
- since it’s possible to imagine it, this means it will come into existence — this is an error because things that exist in the real, physical world do so based on physical processes governed by inductive reasoning
generally, there is a long series of steps between the imagining of some constructed, complex machine and its realization, along with its conceptual foundations it requires sustained effort, trial and error, maintenance, generally a serious fight against entropy to make it function and keep it functioning
the sort of out of control AI imagined by AI doomers is not something we’ve seen before
so we shouldn’t make costly decisions based upon this confusion of reasoning
Nope. That's not the argument. In fact, it's such a bad take that it reeks of a deliberately constructed strawman.
The actual argument is: Since it's possible to imagine it, and doesn't contradict any known laws of nature or technology, and current development appears to be iterating towards it, it might come into existence, thus it presents a statistical risk.
When I take out tornado insurance, it's not because I know my house will be blown away by a storm – it's because I don't know, but the possibility is there.
Certainty is not required in order to conclude that risk exists. Quite the opposite is true: Risk is a function of uncertainty.
I also work in AI and I don't mention my concerns to colleagues who are so anti-AI risk as you. Perhaps your ideas about your colleagues' views are distorted.
The more narcissistic types have figured out it is their moment in the sun to see their name in the paper and the more they play up the idea that AI is going to eat us the more attention they will get from media.
The whole idea is so irrational that I fail to see what other explanation there really is.
The other guilty party are the masses that have been trained to think in terms of appeal to authority instead of using their own brains. They have created the audience for this theater.
There are two definitions of expertise:
1. Knowing more than most people about a topic. This is the type of expertise that wins the Quiz Bowl.
2. Actual mastery of a field, such that predictions and analyses generated by a person possessing such mastery are reliable. This is the type of expertise that fixes your home or car.
The first definition is easily verifiable, and due to the availability heuristic, it is often presented as a legitimate proxy for the second. But it isn't really, not in general.
If I know more about horoscopes than most people, I am a horoscope expert. But it doesn't mean I can be relied on to predict any of the things horoscopes supposedly predict. It's the same with AI risk. Expertise in AI risk is not a basis for credibility because AI risk is not a real scientific field.
Climate change is a real field of science. AI risk is Nostradamic prognostication by people who know more than you.
Any prediction of the future is necessarily based on modeling and extrapolation.
Five years ago AIs couldn't pass a third-grade reading comprehension test. Today they pass in the top 10% of law, medical, and engineering exams for human professionals.
It is absolutely possible to extrapolate from such developments, and doing so is scientific, not "Nostradamic". Many predictions of the potential impact of climate change also include speculative elements, such as societal effects, migration patterns, conflicts, etc., which cannot be modeled or forecast with any real certainty. That doesn't make them unscientific.
In climate change, we are analyzing historical climate data using weather models representing known physical processes. We try to predict the data using these models, and we are only able to do so if we include the forcing from greenhouse gases. From this we can constrain the range of impacts these gases could be having on temperature and forecast likely futures. The forecasts are heavily informed by a thoroughly validated base of prior knowledge, not just drawing lines through a log log plot.
None of this has any counterparts in AI. We don't understand AI systems to anywhere near the level that physics affords understanding of physical systems. We don't even understand them at a Moore's Law level, where you can at least know what engineering innovations are in the pipeline and how far they could plausibly go. Predicting the sophistication of future AI is just Nostradamic prognostication.
Yann LeCun recently gave a presentation arguing that LLMs are a dead end and proposing a completely different approach. His arguments were extremely heuristic and unconvincing, but this at least shows that both sides have bigwigs with unconvincing heuristic arguments.
I wish a cool scifi robot woke up one day and violently optimized all of humanity into paperclips, instead I live in the real world where the jobs are going to evaporate like water in a newly installed desert and the "let them eat cake" will get increasingly louder and blue-check-markier.
I love GPT and my whole life and plans are based on AI tools like it. But that doesn't mean that if you make it say 50% smarter and 50 times faster that it can't cause problems for people. Because all it takes is systems with superior reasoning capability to be given an overly broad goal.
In less than five years, these models may be thinking dozens of times faster than any human. Human input or activities will appear to be mostly frozen to them. The only way to keep up will be deploying your own models.
So to effectively lose control you don't need the models to "wake up" and become living simulations of people or anything. You just need them to get somewhat smarter and much faster.
We have to expect them to get much, much faster. The models, software, and hardware for this specific application all have room for improvement. And there will be new paradigms/approaches that are even more efficient for this application.
For hyperspeed AI to not come about would be a total break from computing history.
This pseudo-intellectual belief structure is very cult like. Its an end of the world scenario that only an elite few can really understand, and they, our saviors, our band of reluctant nerd heroes, are screaming from the pulpit to warn us of utter destruction. The actual end of days. These "black box" (er, I mean, we engineered them that way after decades of research, but no, nobody really understands them, right?) shoggoths will be so incredibly brilliant that they will be able to dominate all of humanity. They will understand humans so well as to manipulate us out of existence, yet they will be so utterly stupid as to pursue paper clips at all cost.
Maybe instead these models will just be really useful software tools to compress knowledge and make it available to humanity in myriad forms to develop a next level of civilization on top of? People will become more educated and wise, the cost of goods and services will drop dramatically, thereby enriching all of humanity, and life will go on. There are straighter paths from where we are today to this set of predictions than there are to many of the doomsday scenarios, yet it has become hip among the intelligentsia to be concerned about everything. Being optimistic is somehow not real, (although the progress of civilization serves as great evidence that optimism is indeed rational) while being a loud mouthed scare mongerer or a quiet, very serious and concerned intellectual, is seen as respectable. Forget that. All the doomers can go rot in their depressive caves while the rest of us build a bad ass future for all of humanity. Once hail bop has passed over I hope everyone feels welcome to come back to the party.
A pragmatic perspective requires one to accept the present reality as it is, rather than hypothesize an exaggerated potential of what could be. Not all concerns surrounding existential risks in technology are necessarily grounded in empirical evidence. When it comes to artificial intelligence, for instance, current models operate at a speed vastly superior to human cognition. However, this does not equate to sentient consciousness or personal motivation. The projection of human traits onto these models may be misplaced, as AI systems do not possess inherently human drives or desires.
Many misconceptions about reinforcement learning and its capabilities abound. The development of systems that can translate abstract objectives into detailed subtasks remains a distant prospect. There seems to be a pervasive certainty about the risks associated with these models, yet concrete evidence of such dangers is still wanting.
This belief system, one might argue, shares certain characteristics with a doomsday cult. There is a narrative that portrays a small group of technologists as our only defense against a looming, catastrophic end. These artificial intelligence models, which were engineered after extensive research, are often misinterpreted as inscrutable entities capable of outsmarting and eradicating humanity, while simultaneously being so simplistic as to obsess over trivial tasks.
Alternatively, these AI models could be viewed as valuable tools for knowledge compression and distribution, enabling the advancement of civilization. As a result, societal education levels could improve, and the cost of goods and services might decrease, which could potentially enrich human life on a global scale. While there seems to be a tendency to worry about every potential hazard, optimism about the future is not unfounded given the trajectory of human progress.
There are certainly different perspectives on this issue. Some adhere to a more fatalistic viewpoint, while others are working towards a brighter future for humanity. Regardless, once the present fears subside, everyone is invited to participate in shaping our collective future.
But I also think it's more anticipatory than speculative to envision AI systems (quite possibly on the request of a human faction) taking control.
And GPT-4 absolutely does do abstract reasoning and subgoals. No it doesn't have many other capabilities or characteristics of humans or other animals but as I said it doesn't need those to be dangerous.
We need to prohibit manufacture or design AI hardware that has performance beyond a certain level. It is not too early to start talking about a risk that could end humanity. I do hope that we can get away with something a few orders of magnitude better than what we have today, but it's really of asking for trouble the more we optimize it, and we may be walking a fine line within a decade or so. Or less. It takes years to design hardware and get manufacturing online, especially for new approaches.
And two orders of magnitude faster may be only a few years away.
Whoever predicts the right direction, (and when the time is right) puts money where their mouth is, stands a shot at unseating... the alt man.
I think the way to use these big ideas is not to try to identify a precise point in the future and then ask yourself how to get from here to there, like the popular image of a visionary. You'll be better off if you operate like Columbus and just head in a general westerly direction. Don't try to construct the future like a building, because your current blueprint is almost certainly mistaken. Start with something you know works, and when you expand, expand westward.
The popular image of the visionary is someone with a clear view of the future, but empirically it may be better to have a blurry one.
paulgraham.com/ambitious.html--
[0] - Or "something very convincingly pretending to think by parroting stuff back", if you're closer to the "stochastic parrot" view.
[1] - Per my hand-wavy hypothesis that the bulk of what we call thinking boils down to proximity search in extremely high-dimensional space.
Even GPT 3.5 has trouble following instructions, but I’ve found that GPT 4 is almost flawless. I can tell it the document uses Australian English but to preserve US spelling for product names and it’ll do it!
One quirk is that it’s almost too good at following instructions. You have to tell it to preserve product names, vendors names, place names, etc… otherwise it’ll “correct” the spelling of anything you forgot to list.
The idea is for it to automatically detect the "language" of each word based on the context and its own understanding of the world.
E.g., the following sentence:
"We deployed windows server data centre 2022 into our data center, which has no windows for physical security."
Will be corrected by GPT-4 to the following:
"We deployed Windows Server Datacenter 2022 into our data centre, which has no windows for physical security."
Notice that it combined "data" and "centre" into "Datacenter" and it corrected the second "center" into "centre", which is the British/Australian spelling of the word. It also correctly capitalised only the first use of the word "windows", etc...
That requires a level of understanding that GPT 3.5 just barely has, and no ordinary grammar checker tool has.
General purpose models containing significant overlap between Project Gutenberg and Github are unnecessary and don't scale. Moby Dick has little to do with C++ unless you're creating art for novelty's sake. This is entirely speculative, but I'm convinced ChatGPT is faking the appearance of a single oracle while delegating requests to specialized models under the hood. It scales better and makes sense than trying to serve a 1T model to address everybody's banal questions.
Like, at its core, for people who only want to write literature, give them a model with underweighed programming-related corpora. Writers don't need it, will never use it, and that space could be filled with training content relevant to literature. Anything else results in expensive, unscalable solutions or jack-of-all-trades, master-of-none outcomes.
In recent usage, GPT3.5 helped me hack my way through writing Pester tests for Powershell scripts for the first time, and I mean hack-- there were a lot of assumptions it made and things it got wrong. GPT4 did a much better job, but I couldn't help but think 3.5 probably has a ton of other training data in it that detracts from the specialization I needed from it in that context. For coding help, you don't want to ask some random librarian who occasionally recommends resources that don't exist; you ask someone who specializes in coding and trust they have familiarity with that domain.
Then we need a new system, because LMs, no matter if they are large or not, cannot do that, for a very simple reason:
A LM doesn't understand "truthfulness". It has no concept of a sequence being true or not, only of a sequence being probable.
And that probability cannot work as a standin for truthfulness, because the LM doesn't produce improbable sequences to begin with...it's output will always be the most (within heat settings) probable sequence. The LM simply has no way of knowing whether the sequence it just predicted is grounded in reality or not.
I claim that the human brain doesn't understand "truthfulness" either. It merely creates the impression that understanding is taking place, by adapting to social and environmental pressures. The brain has no "concepts" at all, it just generates output based on its input, its internal wiring, and a variety of essentially random factors, quite analogous to how LLMs operate.
Do you have any evidence that contradicts that claim?
Empirical evidence? Yes I do.
The brain commands an entity that has to exist and function in the context of objective reality. Being unable to verify it's internal state against that, would have been negatively selected some time ago, because stating: "I'm sure that rumbling cave bear with those big sharp teeth is a peaceful herbivore" won't change the objective reality that the caveman is about to become dinner.
How that works in detail is, to the best of my knowledge, still the subject of research in the realm of neurobiology.
The concept of truth is notoriously hard for humans to grapple with. How do we know something is true isn’t just a neurobiological question, it’s been grappled with throughout the history of philosophy — including major revisions of our understanding in the past 80 years.
And for the record, rumbling cave bears are mostly peaceful herbivores.
For the record, all members of the Genus Ursus belong to the Order Carnivora, which literally translates to "Meat Eaters". And that includes Ursus spelaeus, aka. the Cave Bear.
And while it most likely, like many modern bears, was an Omnivore, that "Omni" very much included small, hairless monkey-esque creatures with no natural defenses other than ridiculously small teeth and pathetic excuses for claws, if they happened to stumble into their cave.
> The concept of truth is notoriously hard for humans to grapple with.
I am not talking about the philosophical questions of what truth is as a concept, nor am I talking about the many capabilities of humans to purposefully reshape others perceptions of truth for their own ends.
I am talking about truth as the observable state of the objective reality, aka. the Universe we exist in and interact with. A meter is longer than a centimeter, and boiling water is warmer than frozen water at the same pressure, whether any given philosophy or fabrication agrees with that or not, is irrelevant.
Wrong statements about Python are simply less probable than wrong statements about Rust, since there is more Python than Rust in the training data.
That changes exactly nothing about the fact that the system isn't able to detect when it makes a blunder in Python.
That is not what we've observed though. Quite the opposite - we're seeing that the bigger LLM is and the more domain-specific material it digested, the more truthful it becomes.
Yes it can still make an error and be unable to spot it, but so can I.
No, that is not my claim. That is part of the explanation for it.
My claim is this: An LLM is incapable of knowing when it produces false information, as it simply doesn't have a concept of "truthfulness". It deals in probabilities, not alignment with objective reality.
And it doesn't matter how big you make them...this fact cannot change, as it is rooted in the basic MO of language models.
So, now that we have covered what my claim actually is...
> That is not what we've observed though. Quite the opposite - we're seeing that the bigger LLM is and the more domain-specific material it digested, the more truthful it becomes.
...I can ask what this observation has to do with it, and the answer is: Nothing at all. LMs with more params may produce untruthful statements less often, but what does this change about their ability to recignize when they do produce them? And the answer is: Nothing. They still can't.
a LLM can indeed know when it produces likely incorrect responses. Not a hypothetical.
What's the point of making claims you have no intention of rescinding regardless of evidence ? People are so funny.
Base GPT-4 was excellently calibrated. So this is just wrong.
I don't think those steps are out of the bounds of possibility, really.
The problem is what you mean when you say "consistency".
The LM checks if sequences are stochastically consistent with other sequences in the training data. Within that realm, the sentence: "In the Water Wars of 1999, the Antarctic Coalitions aramada of Hovercraft valiantly faught in the battle of Golehim under Rear Admiral Korakow, against the Trade Unions Fleets." is consistent. Because, while it is total bollocks, it looks stochasticaly like something that could be in a historical text.
So, in it's context, the LM does exactly what you ask for. It produces output that is consistent with the training data.
Truthfulness is a completely different form of consistency: Does the semantic meaning of the data support the statement I just made? of course it doesn't, there isn't an Antarctic Coalition, there were no Water Wars in 1999, and no one ever built an Armada of Hovercraft for any war against a "Trade Union Fleet".
But to know that, one has to understand what the data means semantically. And our current AIs ... well, don't.
Another wrong statement, you're on a roll today.
https://arxiv.org/abs/2305.11169
https://arxiv.org/abs/2306.12672
There's a word we would use to describe your confidently erroneous statements were it one of the outputs an LLM. Wonder what that might be..
It can reason. To an extent.
“Hallucination” is part of thought. Solving a new problem requires hallucinating new, non existing, possible outcomes and solutions, to find one that will work. It seems that eliminating the ability to interpolate and extrapolate (hallucinations) would make intelligence impossible. It would eliminate creativity, tying together new concepts, creation, etc.
Is the goal AI, or a nice database front end, to reference facts? Is intelligence facts, or is it the flexibility and the ability to handle and create the novel, things that are new?
The ability to have confidence, and know and respond to it, seems important, but that’s surely different than the elimination of hallucinations.
I’m probably misunderstanding something, and/or don’t know what I’m talking about.
It's simply when the predicted probable sequence isn't grounded in reality.
When I ask an LLM to summarize the great water wars of 1999, and how the Trade Union was ultimately defeated by the Antarctic Coalitions hovercraft-fleet under Vice Admiral Zagalow, it isn't "extrapolating" from knowledge of history, it is simply inventing a load of bollocks. But that bollocks will be dressed in fine language and probably mixed in with plausible-sounding references that have a somewhat-logical-sounding relation to the training data.
The problem is, the LM doesn't and cannot know when it produces bollocks.
All it can care about is if the sequences produced are probable according to it's model.
"I'm sorry, but it appears there's a misunderstanding. As of my knowledge cutoff in September 2021, there were no events known as the "Great Water Wars of 1999" involving a Trade Union being defeated by an Antarctic Coalition's hovercraft fleet under Vice Admiral Zagalow. This might be part of a work of fiction, alternative history, or a future event beyond my last training cut-off.
My training includes real-world historical events and existing geopolitical structures, and as of 2021, Antarctica was governed by the Antarctic Treaty System, which prevents any military activity, mineral mining, nuclear testing, and nuclear waste disposal. It also supports scientific research and protects the continent's ecozone.
Please provide more context if this information is from a book, a movie, or a game, or if it refers to something else that I may assist better with."
The result were 2 very well written paragraphs, including the defeat of the Trade unions navy, a ceasfire agreement and a peace agreement ending the water wars.
I rephrased the entire thing as a question, asking the LM to tell me about the conclusion of the war. Again I got a pseudo-historical statement.
> What is heavier, a small floating passenger ferry or a two metric ton heavy rock that sinks to the bottom of the ocean.
> A two metric ton heavy rock would be heavier than a small floating passenger ferry. The weight of the rock is two metric tons, which is equivalent to 2,000 kilograms or 4,409 pounds. The weight of the passenger ferry would depend on its specific design and construction materials, but it is unlikely to be heavier than two metric tons. Therefore, the heavy rock would have a greater weight than the small floating passenger ferry.
It completely relies on surface information such as "small floating" and ignores the deeper "correlation" that all ferries are heavy.
A two metric ton rock weighs two metric tons by definition (or 2000 kilograms). However, a small passenger ferry, while it may look small compared to large ferries or ships, can weigh much more than two metric tons. Even a small passenger ferry can weigh dozens or even hundreds of tons, due to the mass of the hull, the engine, and other equipment on board.
So, without specific information about the ferry's mass, it's safe to assume that a "small" passenger ferry is likely heavier than a two metric ton rock. However, if the ferry is particularly small and lightweight, or the term "ferry" is being used to describe a very small watercraft (like a raft or dinghy), it's possible for it to be lighter. You would need the specific weight of the ferry to give a definitive answer.
The latter given the kind of products that are currently being built with it. You don't want your code completion or news aggregator to hallucinate for the same reason you don't want your wrench to hallucinate, it's a tool.
And as for hallucinations, that's a PR friendly misnomer for "it made **** up". Using the same phrase doesn't mean it has functionally anything to do with the cognitive processes involved in human thought. In the same way a 'artificial' neural net is really a metaphorical neural net, it has very few things in common with biological neurons.
Obviously it's worth it to try and eliminate the incorrect information, but what grand-op is saying is we don't want to do that if it takes away some valuable emergent properties.
I've been thinking there's some parallels between how AIs hallucinate and how human toddlers do. If you ask a toddler/young child a question about a fact they don't know, they will usually say, "iuno" (even when they should), but depending on the child and the circumstances, they will sometimes just make up a story on the spot and sound as if they believe it. "Who invented ice cream?" "Santa Claus! Mommy left him milk and cookies and he turned it into ice cream." It doesn't make any real sense but it seems facially plausible in their universe.
But somewhere between first learning to speak and around 7ish, kids become markedly more accurate how they model the world, and their responses become correspondingly less fanciful. And they continue to improve beyond that point.
So how are kids doing what LLMs are currently incapable of? How do we teach ourselves not to hallucinate? Or do we, really? I mean, if I tell myself I'm going to make it through the intersection before the light turns red, but I end up running the red light, was I just mistaken, or was that a self-delusion, i.e., a mini-hallucination of sorts? Probably a self-driving car would be less likely to make that category of mistake, so maybe I shouldn't be so smug about being grounded in reality.
A more salient question would be, 'how do I know it ISN'T hallucinating'…
Any evidence to support this claim or just commentary ?
Disclaimer: I know little of this field.
[1] https://www.scientificamerican.com/article/perception-and-me...
[2] https://www.frontiersin.org/articles/10.3389/fpsyg.2021.7289....
Data is still king.
But until "I don't know" comes out, rather than hallucinations, we're in trouble.
Yea that was what I was getting at with the "combination" of data. The publicly available data provides the base/primary education, then you specialize it with your proprietary data and bam, you have an AI model that nobody else can produce...an actual product moat.
Train on dataset A to learn to think, use thinking on dataset B to become an export in B's field.
Unless you mean that Reddit is astroturfed with the SEO garbage you're trying to avoid, in which case this will definitely not help.
Is search on Reddit itself still useless?
In the real world, search reduces information acquisition costs as you only have to spend time and resources on finding an existing result rather than recreating it.
Training on something huge like "the internet" is what gives rise to those amazing emergent properties missing in smaller models (including the recent Pi model). And there are only so many datasets that huge.
But its also indeed a waste, as Pi proves.
There probably is some sweet spot (6B-40B?) for specialized, heavily focused models pre trained with high quality general data.
Linear Regression is great for projections, and can even be fit to time series data using lagging.
By "cramming all of the web" on a model what is really going on is the hidden layers of that network are getting better at understanding language and logic. Imagine trying to teach a kid who doesn't know how to read to learn about a Science by only giving them science textbooks. Chances are they won't get very far.
Building little specialist model's don't really work either. It's like trying to train a parrot to do science, sure it can repeat some of the phrases that you give it, but at the end of the day it's not really making any new connections for you.
I wonder what people said about "bus" back in the day, especially those who knew Latin.
It seems that what we need to make a big leap forward is better reasoning. There is a lot of debate between the GPT-4 can/can't reason camps, but I haven't seen anyone try to argue that it reasons particularly well.
People who argue against GPT-4 reasoning at all are arguing against clear results. It extremely easy to show examples and benchmarks of 4 reasoning and understanding. The argument then turns into "well that's not "true" reasoning", whatever that means.
Building the data set for that should be quite trivial.
The other good use cases are using LLM to turn natural language prompts into API calls to real data.
Doing this a couple of times gives me 100% accuracy for my use case that involves some level of summarization and reasoning.
Hallucinations are not as big of a deal at all IMO. Not enough that I'll just sit there and wait for models that don't hallucinate.
for popular languages, though for JS it most of the time outputs obsolete syntax and code.
Today i tried to do a bit of scripting with my son in Garrys Mod, it uses Expression 2 for a Wiremod module. GPT hallucinated a lot of functions and the worst part it switched almost each time from e2 to lua.
It is good at solving homeworks for students, or solving popular problems in popular languages and libraries though it might give you an ugly solution and ugly code, it is probably trained on bad code too and it did not learn to prefer good code over bad code.
(And I don’t mean to be rude - I just really hate that syntax!)
i mean use at least "let"
What we call hallucination is just when the resulting text is wrong but the underlying probabilities could be high.
We likely wouldn't ever know how good the model is as it not only closed but they haven't provided access to anyone.
From section 5:
In Figure 2.1, we see that training on CodeExercises leads to a substantial boost in the performance of the
model on the HumanEval benchmark. To investigate this boost, we propose to prune the CodeExercises
dataset by removing files that are “similar” to those in HumanEval. This process can be viewed as
a “strong form” of data decontamination. We then retrain our model on such pruned data, and still
observe strong performance on HumanEval. In particular, even after aggressively pruning more than
40% of the CodeExercises dataset (this even prunes files that are only vaguely similar to HumanEval, see
Appendix C), the retrained phi-1 still outperforms StarCoder.I agree with you that whole-output MoE models don’t.
Compute challenges are more real, but we are seeing for the first time huge amounts of global capital being allocated to solve specifically these problems, so I am curious what fruit that will bear in a few years.
I mean already the stuff that some of these low level people are doing is absolutely nuts. Tim Dettmers work training with only 4 bits means only 16 possible values per weight and still getting great results.
So far I reckon <$10m in actual revenue.
This isn't what VC's (or microsoft) dream of.
I think its quite likely that OpenAI will make that money back and more, as both the industry leader and with the power of their brand (chatGPT).
According to the InstructGPT paper, that is not the case, showing the data multiple times results in overfitting.
2. They still saw performance improvements which is why they did train on the data multiple times, you can see in the paper.
3. there was a recent paper demonstrating that reusing data still saw continued improvements in perplexity, i am on my ipad so cannot find it now
Ahh mb! Sorry.
But the worst of all behaviors have not diminished and some have actually gotten worse. Most of the improvements now come through what feels like manual heuristics and tuning parameters rather than any actual improvement in intelligence.
I sincerely hope there is a path forward that involves meaningfully culling data which is producing bad behavior, stronger guardrails, and/or a new paradigm in how the model is built from the data entirely, as I don't see a path to level 3 by simply growing the existing model.
So if I understand your are driving a Tesla car with Full Self-Driving (FSD) capability, but you are not an engineer at Tesla who is privy to implementation details.
How do you know if a change you perceive is caused by model retraining vs manual heuristics and tuning parameters?
Though to answer a softer form of your question which is based in my feelings, I would say through a combination of my experience writing AI's, coursework in my AI specialized computer science degree, and through reading the patch notes which often state explicitly the tunings I mentioned.
The comments here and in the Economist article seem to be about only large language models. In the initial announcement of GPT-4 in March, OpenAI described it as “a large multimodal model (accepting image and text inputs, emitting text outputs),” but they haven’t yet released the image part to the public.
What will happen when models are trained not only on text, and not only on text and images, but also on video, audio, chemical analyses of air and other substances in our surroundings, tactile data from devices that explore the physical world, etc.?
Just chatGPT wired up to Wolfram Alpha is already pretty creepy amazing.
My impression is that combining all of those, plus all of the post training quantisation and sparsity tricks into a new model with more training compute than GPT would yield an amazing improvement. Especially the price/performance would be expected to dramatically improve. The current models are very wasteful during inference, there’s easily a factor of ten improvement available there.
> That such big performance increases can be extracted from relatively simple changes like rounding numbers or switching programming languages might seem surprising. But it reflects the breakneck speed with which llms have been developed. For many years they were research projects, and simply getting them to work well was more important than making them elegant. Only recently have they graduated to commercial, mass-market products. Most experts think there remains plenty of room for improvement. As Chris Manning, a computer scientist at Stanford University, put it: “There’s absolutely no reason to believe…that this is the ultimate neural architecture, and we will never find anything better.”
We have so many a-ha moments ahead of us in this field. Seemingly minor changes yielding task speed multipliers, fresh eyes on foundational codebases saying now why the heck did they do it that way when xyz exists and works better, etc. A recent graphics driver update took my local SD performance from almost 4 seconds per iteration to 2.7it/s because someone somewhere had an a-ha moment. We're practically in the Commodore 64 era of this technology and there are only going to be more and more people putting their eyes and minds on these problems every day.
All our current approaches rely on dense matrix multiplications. These approaches necessitate a tremendous amount of communication bandwidth (and low latency collectives). This is extremely challenging and expensive to scale O(n^2.3).
The constraints of physics and finance make significantly larger models out of reach for now.
Road to control narratives or road to be useful?
Actually, now that I think of it, not so different from LLMs…
(Full disclosure, I’ve been a subscriber for a couple of decades)
(Long-time Economist subscriber.)
This specific article seems to be reporting on a very technical issue on how to continue to scale LLM. Even scientific papers have a hard time answering those kind of questions because unless in very special circumstances where we can show with a good confidence that there are limitations (P vs NP for instance) the answer will simply be given by the most successful approach.
Their must intrinsically be some limits to this approach. But that doesn't mean it is a problem or that LLMs cant serve some useful role. Just consider for a moment if we could train an LLM on the combined experiences of all humans that have ever lived. No one is going to suggest that we must go bigger.
A16z’s latest summary of the landscape was way more useful and relevant than this.
It’s not just param count, because not all large models are good, but it clearly is part of the equation.
I always wondered why Altman said going bigger is a dead end. If it were true, saying it would needlessly inform your competition. What’s the use in that? If it were false it might dissuade them from going down that path.. I think I got my answer.
Imo, it is true that the current architecture is hitting the limit. We need a breakthrough on the scale of the transistor to get past this problem. We know it is possible though. Every single human is proof that high performance AI can be run with less energy than a laptop. We just need a dedicated architecture for their working mechanisms the same way a transistor is the embodiment of 1 and 0.
Unfortunately, in terms of understanding intelligence and how they work, I don't think we have made any significant advance in the last few decades. Maybe with better tools at probing how the LLM work, we can get some new insights.