LLMs are a game changer that are going to enable a new programming paradigm as models get faster and better at producing structured output. There are entire classes of app that couldn't exist before because there there was a non-trivial "fuzzy" language problem in the loop. Furthermore I don't think people have a conception of how good these models are going to get within 5-10 years.
Pretty sure it's quite the opposite of what you're implying: People see those LLMs who closely resemble actual intelligence on the surface, but have some shortcomings. Now they extrapolate this and think it's just a small step to perfection and/or AGI, which is completely wrong.
One problem is that converging to an ideal is obviously non-linear, so getting the first 90% right is relatively easy, and closer to 100% it gets exponentially harder. Another problem is that LLMs are not really designed in a way to contain actual intelligence in the way humans would expect them to, so any apparent reasoning is very superficial as it's just language-based and statistical.
In a similar spirit, science fiction stories playing in the near future often tend to have spectacular technology, like flying personal cars, in-eye displays, beam travel, or mind reading devices. In the 1960s it was predicted for the 80s, in the 80s it was predicted for the 2000s etc.
https://www.amazon.com/Friends-High-Places-W-Livingston/dp/0...
tells (among other things) a harrowing tale of a common mistake in technology development that blindsides people every time: the project that reaches an asymptote instead of completion that can get you to keep spending resources and spending resources because you think you have only 5% to go except the approach you've chosen means you'll never get the last 4%. It's a seductive situation that tends to turn the team away from Cassandras who have a clear view.
Happens a lot in machine learning projects where you don’t have the right features. (Right now I am chewing on the problem of “what kind of shoes is the person in this picture wearing?” and how many image classification models would not at all get that they are supposed to look at a small part of the image and how easy it would be to conclude that “this person is on a basketball court so they are wearing sneakers” or “this is a dude so they aren’t wearing heels” or “this lady has a fancy updo and fancy makeup so she must be wearing fancy shoes”. Trouble is all those biases make the model perform better up to a point but to get past that point you really need to segment out the person’s feet.)
Of course it doesn't help that people are a bit hand wavy about what that mark exactly is to begin with. We're very good at moving the goal posts. So that 100% mark has the problem that it's poorly defined and in any case just a brief moment in time given exponential improvements in capabilities. In the eyes of most we're not quite there yet for whatever there is. I would agree with that.
At some point we'll be debating whether we are actually there, and then things move on from there. A lot of that debate is going to be a bit emotional and irrational of course. People are very sensitive about these things and they get a bit defensive when you portray them as clearly inferior to something else. Arguably, most people I deal with don't actually know a lot, their reasoning is primitive/irrational, and if you'd benchmark them against an LLM it wouldn't be that great. Or that fair.
The singularity is kind of the point where most of the improvements to AI are going to come from ideas and suggestions generated by AI rather than by humans. Whether that's this decade or the next is a bit hard to predict obviously.
Human brains are quite complicated but there's only a finite number of neurons in there; a bit under 100 billion. We can waffle a bit about the complexity of their connections. But at some point it becomes a simple matter of throwing more hardware at the problem. With LLMs pushing tens-hundreds of parameters already, you could legitimately ask what a few more doublings in numbers here enable.
Insofar as it's simple to throw like six orders of magnitude more hardware at something that has already had a lot of hardware thrown at it.
o Computer/human interfaces may become so intimate that users
may reasonably be considered superhumanly intelligent.
o Biological science may find ways to improve upon the natural
human intellect.
https://edoras.sdsu.edu/~vinge/misc/singularity.htmlI don't believe LLMs will really get us much of anywhere, Singularity-wise. They're just ridiculously inefficient in terms of compute (and thus power) needs to even do the basic pattern-prediction they do today. They're neat tools for human augmentation in some cases, but that's about all they contribute.
I think, even prior to the recent explosion of LLM stuff, that the aggregate of Humans and the depth of their interconnections on the Internet is already starting to form at least the beginnings of a sort of Singularity, without any AI-related topics needing to be introduced. The way memes (real memes, not silly jokes) spread around the Internet and shape thoughts across all the users, the way the users bounce ideas off each other and refine them, the way viral advocacy and information sharing works, etc. Basically the Singularity is just going to be the emergent group consciousness and capabilities of the collective Internet-connected set of Humans.
Blackwell.
> o Develop human/computer symbiosis in art: Combine the graphic generation capability of modern machines and the esthetic sensibility of humans. Of course, there has been an enormous amount of research in designing computer aids for artists, as labor saving tools. I'm suggesting that we explicitly aim for a greater merging of competence, that we explicitly recognize the cooperative approach that is possible. Karl Sims [22] has done wonderful work in this direction.
Stable Diffusion.
> o Develop interfaces that allow computer and network access without requiring the human to be tied to one spot, sitting in front of a computer. (This is an aspect of IA that fits so well with known economic advantages that lots of effort is already being spent on it.)
iPhone and Android.
> o Develop more symmetrical decision support systems. A popular research/product area in recent years has been decision support systems. This is a form of IA, but may be too focussed on systems that are oracular. As much as the program giving the user information, there must be the idea of the user giving the program guidance.
Cicero.
> Another symptom of progress toward the Singularity: ideas themselves should spread ever faster, and even the most radical will quickly become commonplace.
Trump.
> o Use local area nets to make human teams that really work (ie, are more effective than their component members). This is generally the area of "groupware", already a very popular commercial pursuit. The change in viewpoint here would be to regard the group activity as a combination organism. In one sense, this suggestion might be regarded as the goal of inventing a "Rules of Order" for such combination operations. For instance, group focus might be more easily maintained than in classical meetings. Expertise of individual human members could be isolated from ego issues such that the contribution of different members is focussed on the team project. And of course shared data bases could be used much more conveniently than in conventional committee operations. (Note that this suggestion is aimed at team operations rather than political meetings. In a political setting, the automation described above would simply enforce the power of the persons making the rules!)
Ingress.
> o Exploit the worldwide Internet as a combination human/machine tool. Of all the items on the list, progress in this is proceeding the fastest and may run us into the Singularity before anything else. The power and influence of even the present-day Internet is vastly underestimated. For instance, I think our contemporary computer systems would break under the weight of their own complexity if it weren't for the edge that the USENET "group mind" gives the system administration and support people!) The very anarchy of the worldwide net development is evidence of its potential. As connectivity and bandwidth and archive size and computer speed all increase, we are seeing something like Lynn Margulis' [14] vision of the biosphere as data processor recapitulated, but at a million times greater speed and with millions of humanly intelligent agents (ourselves).
Twitter.
> o Limb prosthetics is a topic of direct commercial applicability. Nerve to silicon transducers can be made [13]. This is an exciting, near-term step toward direct communcation.
Atom Limbs.
> o Similar direct links into brains may be feasible, if the bit rate is low: given human learning flexibility, the actual brain neuron targets might not have to be precisely selected. Even 100 bits per second would be of great use to stroke victims who would otherwise be confined to menu-driven interfaces.
Neuralink.
---
> Blackwell.
I'm fucking sorry but there is no LLM or "AI" platform that is even real intelligence, today, easily demonstrated by the fact that an LLM cannot be used to create a better LLM. Go on, ask ChatGPT to output a novel model that performs better than any other model. Oh, it doesn't work? That's because IT'S NOT INTELLIGENT. And it's DEFINITELY not "superhuman intelligence." Not even close.
Sometimes accurately regurgitating facts is NOT intelligence. God it's so depressing to see commenters on this hell-site listing current-day tech as ANYTHING approaching AGI.
I don't know where that "computationally sufficient" line is. It'll always be fuzzy (because you could have a very slow, but smart entity). And before we have a working AGI, thinking about how much computation we need always comes down to back of the envelope estimations with radically different assumptions of how much computational work brains do.
But I can't rule out the idea that current architectures have enough processing to do it.
If a human brain does its magic with 10 petaflops, and you have 1 petaflop, you should be able to make an equivalent to the human brain that runs at 1/10th of the speed but never sleeps. In other words, once you've reached the same order of magnitude it doesn't matter.
On the other hand, Kurzweil's math really comes down to an argument that the brain is using about 10 petaflops for inference, but it also is changing weights and doing a lot more math and optimization for training (which we don't completely understand). It may (or may not) take considerably more than 10 petaflops to train at the rate humans learn. And remember, humans take years to do anything useful.
Further, 10 petaflops may be enough math, but it doesn't mean you can store enough information or flow enough state between the different parts "of the model."
These are the big questions. If we knew the answers, IMO, we would already have really slow AGI.
Ok, let's run this test of "real intelligence" on you. We eagerly await to see your model. Should be a piece of cake.
By that logic most humans are also not intelligent.
And now I'm transforming that through the concept of taking a photograph and applying the clone tool via a light airbrush.
Repeat enough times, and you get uncompilable mud.
LLMs are not going to generate improvements.
I currently expect we'll need another architectural breakthrough; but also, back in 2009 I expected no-steering-wheel-included self driving cars no later than 2018, and that the LLM output we actually saw in 2023 would be the final problem to be solved in the path to AGI.
Prediction is hard, especially about the future.
AFAICT, both are guesses. The low-end estimate I've seen for human brains are ~ 162 GFLOPS[0] to 10^28 FLOPS[1]; even just the model size for GPT-4 isn't confirmed, merely a combination of human inference of public information with a rumour widely described as a "leak", likewise the compute requirements.
[0] https://geohot.github.io//blog/jekyll/update/2022/02/17/brai...
That assumes that you can represent all of the useful parts of the decision about whether to fire or not to fire in the equivalent of one floating point operation, which seems to be an optimistic assumption. It also assumes there's no useful information encoded into e.g. phase of firing.
> Imagine that there's a little computer inside each neuron that decides when it needs to do work
Yah, there's -bajillions- of floating point operation equivalents happening in a neuron deciding what to do. They're probably not all functional.
BUT, that's why I said the "useful parts" of the decision:
It may take more than the equivalent of one floating point operation to decide whether to fire. For instance, if you are weighting multiple inputs to the neuron differently to decide whether to fire now, that would require multiple multiplications of those inputs. If you consider whether you have fired recently, that's more work too.
Neurons do all of these things, and more, and these things are known to be functional-- not mere implementation details. A computer cannot make an equivalent choice in one floating point operation.
Of course, this doesn't mean that the brain is optimal-- perhaps you can do far less work. But if we're going to use it as a model to estimate scale, we have to consider what actual equivalent work is.
There's basically a few axes you can view this on:
- Number of connections and complexity of connection structure: how much information is encoded about how to do the calculations.
- Mutability of those connections: these things are growing and changing -while doing the math on whether to fire-.
- How much calculation is really needed to do the computation encoded in the connection structure.
Basically, brains are doing a whole lot of math and working on a dense structure of information, but not very precisely because they're made out of meat. There's almost certainly different tradeoffs in how you'd build the system based on the precision, speed, energy, and storage that you have to work with.
I've also heard in a recent talk that the optic nerve carries about 20 Mbps of visual information. If we imagine a saturated task such as the famous gorilla walking through the people passing around a basketball, then we can arrive at some limits on the conscious brain. This does not count the autonomic, sympathetic, and parasympathetic processes, of course, but those could in theory be fairly low bandwidth.
There is also the matter of the "slow" computation in the brain that happens through neurotransmitter release. It is analog and complex, but with a slow clock speed.
My hunch is that the brain is fairly low FLOPs but highly specialized, closer to an FPGA than a million GPUs running an LLM.
And we don't know how many GPT-4 instances run on any single A100, or if it's the other way around and how many A100s are needed to run a single GPT-4 instance. We also don't know how many tokens/second any given instance produces, so multiple users may be (my guess is they are) queued on any given instance. We have a rough idea how many machines they have, but not how intensively they're being used.
> You can cut a brain open and see how many neurons it has and how often they fire. Kurzweil's 10 petaflops for the brain (100e9 neurons * 1000 connections * 200 calculations) is a bit high for me honestly. I don't think connections count as flops. If a neuron only fires 5-50 times a second then that'd put the human brain at .5 to 5 teraflops it seems to me.
You're double-counting. "If a neuron only fires 5-50 times a second" = maximum synapse firing rate * fraction of cells active at any given moment, and the 200 is what you get from assuming it could go at 1000/second (they can) but only 20% are active at any given moment (a bit on the high side, but not by much).
Total = neurons * synapses/neuron * maximum synapse firing rate * fraction of cells active at any given moment * operations per synapse firing
1e11 * 1e3 * 1e3 Hz * 10% (of your brain in use at any given moment, where the similarly phrased misconception comes from) * 1 floating point operation = 1e16/second = 10 PFLOP
It currently looks like we need more than 1 floating point operation to simulate a synapse firing.
> The other estimates like 1e28 are measuring different things.
Things which may turn out to be important for e.g. Hebbian learning. We don't know what we don't know. Our brains are much more sample-efficient than our ANNs.
Firstly, Kurzweil underestimates the number connections by order of magnitude.
Secondly, dentritic computation changes things. Individual dentrites and the dendritic tree as a whole can do multiple individual computations. logical operations low-pass filtering, coincidence detection, ... One neuronal activation is potentially thousands of operations per neuron.
Single human neuron can be equivalent of thousands of ANN's.
As long as the bottleneck is the fab capacity as wafers per hous, the number of operations per second per chip area determines who will produce more compute with best price. It's a good measure even between different technology nodes and superchips.
Nvidia is leader for a reason.
If manufacturing capacity increases to match the demand in the future, FLOPS or TOPS per Watt may become relevant, but now it's fab capacity.
I know they can provide creative new solutions to totally novel problems from firsthand experience… instead of assuming what they should be able to do, I experimented to see what they can actually do.
Focusing on the simple mechanics of training and prediction is to miss the forest for the trees. It’s as absurd as saying how can living things have any intelligence? They’re just bags of chemicals oxidizing carbon. True but irrelevant- it misses the deeper fact that solving almost any problem deeply requires understanding and modeling all of the connected problems, and so on, until you’ve pretty much encompassed everything.
Ultimately it doesn’t even matter what problem you’re training for- all predictive systems will converge on general intelligence as you keep improving predictive accuracy.
Until we get to a point where an AI has the wherewithal to create a fab to make its own chips and then do assembly w/o human intervention (something along the lines of Steve Jobs vision of a computer factory where sand goes in at one end and finished product rolls out the other) it doesn't seem likely to amount to much.
An LLM is not going to suggest a reasonable improvement to itself, except by sheerest luck.
But then next generation, where the LLM is just the language comprehension and generation model that feeds into something else yet to be invented, I have no guarantees about whether that will be able to improve itself. Depends on what it is.
It would be left to human researchers to investigate them and find out if any work. If they succeed, the LLM will get all the credit for the idea, if they fail, it's them who will have wasted their time.