Are We Close to AGI?
polymathicbeing.substack.com
polymathicbeing.substack.com
So as of today, we don't really have any kind of simple metric like "IQ" that we can apply to AI systems, and even if we did have the metric, at this point we wouldn't really know it's range, whether to expect it to be a linear function or not, etc.
That said... does it seem clear that we are making some kind of progress? I'd say yes, but with the admission that that's somewhat subjective. In particular, even though we know that current AI systems are getting better at doing whatever they do, we don't really know if that meaningfully counts as progress towards AGI. It is possible, for example, that LLM approaches will hit a ceiling in their development, and that the much feared "GPT 17" really won't even be a thing. And it may likewise turn out that the progress made on LLM's doesn't turn out to contribute much at all to whatever the eventual mechanism behind AGI turns out to be.
Is that the way I expect things to play out? Not sure. I have a hunch that LLM's will prove to be insufficient to qualify as AGI on their own, but that they may well serve as an element of a modular / hybrid system that includes LLMs, reasoners using deduction, induction, and abduction, evolutionary algorithms, and FSM-knows what else.
> When a distinguished but elderly scientist states that something is possible, he is almost certainly right. When he states that something is impossible, he is very probably wrong.
(Clarkes first law, copied from https://en.wikipedia.org/wiki/Clarke%27s_three_laws)
Agreed!
Isaac Asimov's Corollary to Clarke's First Law: "When, however, the lay public rallies round an idea that is denounced by distinguished but elderly scientists and supports that idea with great fervour and emotion – the distinguished but elderly scientists are then, after all, probably right."
The only thing we can reasonably say about the state of AI is that we are closer to so-called AGI. We're seeing exciting progress and that it turn is attracting more efforts in the field, potentially leading to an acceleration.
But we could hit a wall, for many reasons, the field is almost entirely filled by unknown unknowns.
Nuclear fusion, Quantum Computing, Nanotech, Genetic Engineering, all of those fields are still evolving, slowly, because of technological and scientifical roadblocks.
Because of this we need to some up with sets of well defined measures of increasing capabilities that systems can be said to meet.
Agreed. In fact, if you go back through my HN comment history (not an exercise I actually suggest mind you) you'll see that I frequently harp on the distinction between "human level" intelligence and "human like" intelligence. Where in my lexicon, those are distinguished largely the absence of actual reference to human experiences like "getting ridiculed by a bully" or "pooping your pants in 3rd grade" in the case of "human level, but not human like" AI.
I think not everyone agrees that this distinction is important, or maybe they just haven't spent enough time thinking about the implications. But I believe it is important to call this out and make clear which scenario we're talking about in many (most?) contexts.
What's important is that AI will quickly replace toil in white collar work. Unfortunately, a lot of people make entire careers of that kind of work. "Paper pusher" type jobs are the most at risk but so are the low end of lots of creative jobs like copywriting, graphic design and even script writing.
We're going to have to reckon with the consequences of AI long before AGI becomes real.
Only occurred to me much later that now they didn't have much work to do and were at a higher risk of being laid off! A good lesson to think about all the implications of the software I'm writing. We've seen AI-adjacent tooling getting leveraged to eliminate more of that category of work in games already, like automated generation of NPC "barks".
only in so far as not reaping the rewards for having done more work than the company paid you for.
Efficiency improvements are always good for society as a whole, despite it hurting some people who originally was paid due to the inefficiency.
When the economic/capital/wealth 1% end up hoarding all the profitability gains, society as a whole loses. I have come to loathe the term "disruption" as it has always meant externalizing costs to make a quick buck: either by technology replacing actual, real humans with real families who now have to struggle to survive or by breaking laws and regulations like Uber and AirBnB did, with society paying the costs associated with that - in the case of Uber, society picked up the tab for underinsured drivers or for their pension/social security contributions which don't apply for "freelancers", and in the case of AirBnB all the neighbors affected by the noise and literal other disruptions an AirBnB brings with it.
It isn't actually a criticism of technology/automation!
In healthier countries, like the Nordics, they've organized their society so that automation, resource extraction, and really all success in business benefits their society as a whole rather than concentrating the wealth into the hands of the elite.
Having written that, I feel like I'm picking a nit, because what you said is true in the USA and I don't see it changing in my lifetime.
Everyone else barely has a SWF worth the name and almost all Western and Western-allied countries have fallen into the neoliberal/"lean government" trap.
[1] https://en.wikipedia.org/wiki/List_of_countries_by_sovereign...
Using tax revenue to fund a social safety net, public works, education, and housing achieves the goal.
The sovereign wealth fund is more about what to do with the money between taxation and social works. If the country has a sovereign wealth fund without all the social spending then I'd say the success of the country's business hasn't been linked to the wellbeing of the people.
Wealth is not income. It’s the value of owned assets: money, property, etc.
So where does the wealth come from? There are a couple of different answers-and I’m not going to discuss stealing, slavery, exploitation or other illegal and immoral practices-but generally the answer for lots of wealthy, legitimate people is it is CREATED from “deals.”
Wherever one person (or group) can sell something to someone else at a multiple of its current value, wealth is created.
The expectation that something of value will either generate income or grow in value in future years creates a higher price for that thing right now. That entices the buyer, and creates wealth for the seller.
How much wealth? If you have a rock today that you can tell a story about being potentially worth 100x in 10 years, then you can probably sell that rock for 10x of its actual value today. That is wealth creation.
This is related to the concept of “time value of money.”
Real estate is one example. Selling a business is another. Art is yet another example.
Have a great day.
Edit: [1] https://www.aei.org/carpe-diem/the-public-thinks-the-average...
Both wealth and income disparity have risen over the last decades - while the lower income group's shares have stayed constant, the middle class has eroded in favor of those already in wealth - and only the top 5% managed to gain wealth after the 2009-ff depression [1], with their gains by necessity coming from the suffering of the 95%.
The reason is simple: if you are already wealthy, you can afford snatching up assets of all kinds for fire-sale prices - particularly those who had to panic-sell their assets or walk away from properties they could no longer afford or were completely underwater after the loss in value despite paying down their mortgage for many years in the early COVID crisis months or in the 2009-ff depression to survive got hit really hard. And on top of that, simply being wealthy puts you in a very favorable economic position... stuff like using higher-risk but also higher-profit financial instruments is only allowed after hitting a certain wealth threshold, you could generate virtually free money and evade taxes by taking loans on your continuously appreciating stocks, or use highly complicated multi-national tax evasion constructs.
[1] https://www.pewresearch.org/social-trends/2020/01/09/trends-...
And where does the vast majority create their wealth from... Labor. What does AI stand to devalue... Labor. Just remember that most people don't have the capital you're talking about!
I said that wealth was created in the transaction. Labor may be required to create the thing being sold, such as a business, but that labor is becoming less and less important over time. Take the sale of the company “WhatsApp” to Facebook. They sold for $16 billion, but they only had 55 employees with some equity. That’s a lot of wealth for those people.
With AI, this type of productivity (like WhatsApp) with less labor required will only increase. More productivity, less labor.
The capital for creating WhatsApp came from investors. The founders likely had no capital of their own. But they are super wealthy today.
Those who expect to be employed (labor) by strictly being told what to do for a salary will be on the sidelines of wealth creation.
Those who figure out how to use AI to build valuable things and then sell them to others will create their own wealth.
AI is the “fix” for less capital and less labor.
Everybody needs to think about jumping into the pool and start swimming. Anyway, that’s what HN is supposed to be about: entrepreneurship.
>Those who figure out how to use AI to build valuable things and then sell them to others will create their own wealth.
See, this is why we cannot have discussions about generalized AI, that is the entrepreneurs don't think they'll be replaced by AI, and if they pull themselves hard enough by the bootstraps hard enough they'll overcome the all seeing data harvesting Molochs we've created by sheer willpower and brilliance. And maybe in a few cases will... but will it matter when the angry masses displaced by technology club them to death in the streets?
Entrepreneurship without responsibility equals calamity. Drastically changing the social contract and saying 'good luck everyone else' will be our modern "let them eat cake" moment.
no, because those angry masses will firstly get shot by automated drones that are run by AI with facial recognition, before they even gather.
However, there is a benefit to the person who is no longer employed doing mindless toil! It gives them another chance to find healthier work. Again, in the USA this benefit is outweighed by the possibility of immiseration, but it is at least a small plus in the 'pros' column.
In fact, the U.S. spends more per capita than the Nordic countries, as discussed in other comments. Norway is around $7,500 per person. [2]
So what’s the difference? The U.S. is the “wild, Wild West” and Norway is not, according to [2].
[1] https://federalsafetynet.com/entitlement-programs/entitlemen...
[2] https://freakonomics.com/2010/05/who-spends-more-on-social-w....
You wrote code that could have eliminated other people's jobs.
Someone else "wrote" chatGPT and it very much looks like the precursor to the thing that could eliminate your job. I can almost imagine the people at openai having the same "Whoops" attitude as you.
Personally, I see job elimination as my goal and my purpose as a developer. Whenever my company is able to lay people off as a result of my work, that means I am doing a good job. I always try to keep track of how many people my software can replace and how much my company pays these people, so that I can take credit for these cost savings and claim a portion of their salaries for myself during salary negotiations.
Most of what I did at that job was building a syntax highlighting editor w/autocomplete for game scripts, making builds faster, etc. Instead of eliminating jobs that mostly let us ship the same product faster.
I think we may see the best at creative tasks further separating themselves from the pack, these AI systems are best used by experts who can filter and quality check the AI output. I can see software engineering/dev becoming one extremely skilled dev managing a fleet of LLMs weaving together their own work and AIs’ work.
I do want to pick on what I think is a bit of hype around the potential to eliminate whole classes of work, though.
First, in the weeds: I keep seeing this reference to script writing, which I assume is a reference to the WGA strike. But as https://www.gq.com/story/writers-strike-2023-wga-explained explains, the big technological transformation driving this dispute isn't the risk of generative AI; it's the changes in royalty payments created by streaming platforms.
I'm pretty skeptical of claims that generative models can replace script writers. Like, maybe for some really, really crappy shows? But I think "My ChatGPT could write that" is the 21st century version of looking at a Rauschenberg and saying, "My 5 year old could paint that." Good luck, I guess?
More generally: Labor productivity growth has remained fairly constant over the last 40 years or so. So anyone who thinks generative AI is going to totally upend productivity and the job market is implicitly suggesting it's a bigger technological change than the Internet--which didn't really do that!
So...dunno. That's a bold claim!
In 1820 if you said that 99% of farm labor will be replaced by machines in 100 years, you'd have been given the response of 'bullshit'. And yet a century later we had done just that.
With that redistribution of labor came the largest wars the world had ever seen. We don't even have to eliminate 'whole classes of work' to destabilize the working world and economy. Just eliminating 25% of work and not coming up with something new for the now jobless to do will create a wave of populism that can cause problems.
I'm not yet prepared to agree that the scale of bellicosity in the 20th century was the result of labor displacement (as opposed to, say, the results of the mechanization of violence). I mean, like, https://en.wikipedia.org/wiki/List_of_conflicts_in_Europe.
There was probably more political upheaval--that I agree with. But there were a ton of European wars prior to the industrial revolution and the (consequent?) rise of democratic movements.
This will only be good for our economy, any argument against is arguing against the entire trajectory of human technological advancements and its benefits to economic growth. As such, this argument requires a lot more evidence than the plainly obvious: that jobs will be lost.
Seriously, who wants to work a job that a machine can do better, cheaper, and more efficiently? Work for the sake of work is not meaningful.
This is creative destruction and it is the engine that drives economic growth. It can only be a good thing.
You say this in a manner like you're ripping off a bandaid. Meanwhile we have no idea if this is tourniquet that if moved incorrectly will cause us to bleed to death.
At the start of the 20th century there was massive technological change, and this growth of technology lead to two world wars. The big issue is humanity may not be able to survive any further wars along those lines. By the end of the war we created civilization ending weapons. Our capability for destruction has only grown since then. Attempting to move to a new economic paradigm without having war starting authoritarian fascist leaders destroy the world is, I feel, going to be a very difficult task to achieve.
The harsh truth is, with more automation many people simply won’t be able to productively contribute to the economy. Yet we all deserve to live a comfortable life (hopefully that’s not controversial).
On the one hand, I welcome the ability to leverage machines to automate work that I find boring, with ever-increasing complexity and cleverness. This frees me up to remain in the creative thinking headspace, and interpersonal relationships which are oh-so-necessary to build products. From a team perspective, this also allows a fixed-size group of engineers to take on ever increasing scopes and scale.
At the same time, society at large operates in a very different mode. I would argue that it has not finished "digesting" the previous technological advancements. While disruption is something to be celebrated on a personal level, in a society, it starts out as a net negative before benefits start to be realized. And as usual, it is the most vulnerable members who suffer the most, and the most powerful who gain the most.
I am optimistic in the short term for myself and in the long run I think it will be beneficial for society. But there will be a lot of pain born by the vulnerable.
For instance, at my current employer we have about 30 content marketeers. Their job could be made much mor efficient with AI, leaving them room to do more creative work instead of toiling with crappy tools. This could create a lot of value, for our customers too if the quality increases.
Yet every single 'manager' just sees AI as an opportunity to oust these people. They see workers as tools that should be replaced as soon as possible. They would rather have worse quality content, just more and cheaper.
While this kinda holds for all technologies, the reach and impact of AI seems pretty vast. I feel some skeptisism is waranted tbh
Yep, I once wrote an app that cut the time needed to process medical claims by half. I asked the manager what was going to happen to the medical claims staff. Were they going to be retrained? 'No. Half are getting let go as soon as this goes to prod.' I went from feeling very proud about the app to feeling absolutely terrible.
The wealth gap is a distraction, it doesn't measure anything meaningful. We shouldn't care how big the wealth gap is, rather, are more people leading healthy, happy lives. To go further, equal opportunity is more important than equal outcomes.
But that's not as big a deal as one might think. In many ways this is just restating the Problem of Induction[1]. We don't know that the future will always resemble the past, or that trajectories will remain the same over time. An exponential function, for example, looks linear if you zoom in sufficiently. Without the broader perspective, how can you know you're on an exponential curve (or whatever). It's not so different from how you can't easily tell that the Earth is spherical by just glancing out of your window.
This is creative destruction and it is the engine that drives economic growth. It can only be a good thing.
Don't get me wrong... I'm a Libertarian and very pro-Capitalism and generally agree that Creative Destruction has been and likely will continue to be a force for productive change. But even I won't go as far as saying that it can only be a good thing. Who knows what kind of Black Swans[2] are waiting around the corner?
If you think things can only get better, you might have deluded yourself.
You think Buzz Aldrin didn't want to go to the Moon because a probe with a camera would have been cheaper? I think this claim is flawed; people want to contribute and work. People don't want to be exploited, overworked, underpaid, set in contest with machine production levels.
> "This will only be good for our economy"
Having two trillionaires who swap fortunes every day, and half a billion starving peasants would be great for the economy on paper and awful for the humans. Forget the economy, will this be good for the people? What kind of work is lacking now which a large number of people made unemployed could do to earn a living and stay alive? A hundred years ago houses were cold and empty, pantries were bare, communication was expensive and slow, people walked everywhere in their one suit and drank well water. Now houses are warm, food is refridgerated, comms is cheap and fast, even the quite poor have cars, clothing, and as much clutter as anyone needs, the hungry are not lacking for food production they are lacking for access to food. Are we not nearly saturated with music, books, films, streamers, YouTubers, TV shows, documentaries on anything and everything?
Take a vacuum cleaner, first it didn't exist and people had rugbeaters. Then it existed as a horse drawn vehicle which came around and let rich people vacuum occasionally. Then it became an engine driven thing out in a shed with pipes to every room in a big house and an oil trap. Then it became electric and movable, plugged into light sockets overhead. Then it became small and portable and using mains power, plastic instead of metal, bagged, then upright and light, then bagless, then then quieter, then cyclonic, now cordless or Roomba. And manufactured far away.
How much more development can the vacuum cleaner take which an ordinary person could do - it probably won't be ten million truck drivers, marketing managers, paralegals, or janitors developing the faster charging battery chemistry or the newer automated production line or the fluid dynamics simulation for making it quieter or the more intelligent Roomba software.
All household goods are in a similar position; what's the equivalent of the last 100 years of development for the next 100 years? It's looking pretty likely that it's not flying cars, spaceships, space elevators, jetpacks, teleporters, Star-Trek replicators, silver jumpsuits, food pills, and so far it's looking like it isn't even Blu-Ray, 3D TV, VR headsets. It's all very well to talk about unlimited greed, but at "enough stuff for a lifetime" and then at "hundreds of times more stuff than a lifetime" we have to be approaching some practical limits, don't we?
Economists are always harping on about how disruptions are temporary, how 'good jobs' will replace 'bad' ones. What they don't say is those disruptions can be permanent and dire for those disrupted. There are many rust belt towns that are a hollow husk of the places they once were due to the loss of good paying factory jobs. A lot of those people never found 'good' jobs ever again. Same thing happened in the south with NAFTA. Apparently our economists are all to happy to leave people to rot if it means we can increase gdp a few more percent.
I'm not saying that we should resist automation or ai, I'm saying given our history it's pretty plain to see that we suck at protecting and helping those harmed by it.
People need a job to earn money they can later pour back into the economy which is just supposedly keeps on growing (as if continuous growth could be sustainable). A 100 years ago a village was working hard at harvest, now a GPS-guided combine harvester does the job. Sure, we could made up a few new jobs, but many office jobs never made sense (did we really ever need that many layers of middle-management? Their personnel alltogether are not needed).
Unless some UBI-like system gets implemented, there will be revolts. “Every society is three meals away from anarchy”
You can say it's different this time, but that's what's been said every time.
Don't forget programming jobs. I wouldn't classify programmers as "low end" as that's sort of an offensive categorization.
But among the set of careers you have chosen as a target for replacement... Programming is among that group. I mean LLMs are obviously translating English to code in a way that was never seen before.
Given that you're a programmer on a site that is mostly made up of programmers I wonder why that specific career was not mentioned. Seems a little convenient to me.
I'm not saying this is you, but I feel there is a lot of hostile defensiveness on the capabilities of AI because the AI potentially trivializes our abilities as programmers. If we are among the group of careers you labeled as "low level" then are we not "low level" ourselves? Do we have the fortitude and objectiveness to accept that?
Probably not.
Google could also spit out that leetcode tutorial page given the correct keywords 10 years ago (which it shamefully can’t do anymore.. am I the only one that can’t find anything non-trivial anymore with it?!), but that didn’t make anyone with access to the internet a programmer either.
Its solving things that was unthinkably impossible to solve by AI.
Think of it logically. Before it could do 0% of our ability. With chatGPT it increases to 30%. With a domain specific LLM trained to target code... that increases to about 50%.
A person with 50% of their brain is obviously mentally deficient. It's not even comparable to human intelligence. It's like a stupid assistant. Helpful, but still stupid.
The issue here isn't really about the comparison between ai capabilities and human capabilities. It's about progress.
That 30% to 50% trendline in LLMs and other areas of improvement in AI is the Same trendline as we see in say creative writing, art, etc. Etc.
It points to the fact that coding is not inherently superior to other forms of creation because the trendline is exactly the same. In fact drawing really good art is actually harder then programming (no 6 months boot camps for art).
We are unfortunately on target for replacement because the next data point on that trendline is obviously not 51%. It's wishful thinking and illogical to ignore the obvious trendline.
Something tells me that the low end creatives have a lot more to worry about than the paper pushers. If the past fifty years are any indication, there is a lot more inertia when it comes to replacing existing jobs than there is replacing something you would contract out.
Deepmind has quite a bit of work outside the language area that has a lot of potential. Peter Velockovic, Matt Botvinick, Chelsea Finn, Taco Cohen, Micheal Bronstein, et al are all working completely outside the arena of NLP, but they have much more groundbreaking work than what everyone seems laser focused on right now.
There are tons of examples really, but for instance: their agents that have emergent properties of teamwork, as well as their work on developing entirely new mathematics in some of the hardest fields (knot theory and representation theory).
Creating new mathematics and proving theorems is probably the most creative form of human expression there is, so the fact that AI can achieve this is quite a big leap. Combining this with the deficits in LLMs can construct something that has much fewer flaws.
While some of the hype around language models is fading now, we will probably see another resurgence of interest by the media when Deepmind releases a system with these capabilities underneath a LM. These LMs have been around for many years, and are just now getting mass attention. So it will likely take some more time of these much more sophisticated systems being solely in the academic arena before people become aware of what is really going on in AI (company products typically lag very far behind what academia is capable of).
However, soon there will be products that have attached a linguistic interface to these types of systems, and will likely yield some really impressive capabilities.
> Theory of Mind: These AI systems are capable of understanding mental states and can simulate the behavior of other agents.
GPT models can do this.
> Self-Aware
Why is that a requirement?
> These AI systems are capable of learning any intellectual task that a human can.
Again, GPT models appear to do this (at least for text-based tasks).
That being said, I think GPT models do have some limitations that prevent them for learning any intellectual task a human can do. I wrote a blog post a few days ago that explored trying to "teach" ChatGPT to perform binary addition [0]. My conclusion was that, when ChatGPT appears to solve a logical problem, that solution is just a regurgitation of a low-entropy example generalized from its training data. But the model fails to apply even basic logical rules as entropy increases, indicating that GPT models have a fundamental limit of applying logic to high entropy problems.
https://platform.openai.com/tokenizer
10011111110100011100011011110010
00111010100101000100000101111011
is tokenized as: [100][1111][11][101][0001][11][0001][101][11][100][10]
[00][11][101][01][001][01][0001][0000][01][01][11][101][1]
It might be more trainable if the the bits where space delimited so that only the tokens [ 1] and [ 0] were in use.As it is, it doesn't really understand the contents of the tokens and manipulating [405, 1157, 8784, 486, 8298, 486, 18005, 2388, 486, 486, 1157, 8784, 16] is not easily trained - especially if there is a different bit sequence. Note also the misalignments of the tokens for bit length that can make it more difficult to work with.
Let's add the two binary numbers provided, limiting the work on each bit to a single line:
1 0 0 1 1 1 1 1 1 1 0 1 0 0 0 1 1 1 0 0 0 1 1 0 1 1 1 1 0 0 1 0
+ 0 0 1 1 1 0 1 0 1 0 0 1 0 1 0 0 0 1 0 0 0 0 0 1 0 1 1 1 1 0 1 1
--------------------------------------------------------------
1 0 1 1 0 1 1 0 1 0 0 1 0 0 0 1 0 1 1 1 1 1 0 0 0 1 1 1 1 1 0 1
So, the result of the addition is the binary number 10110110101001010111111001111101.
The answer is still wrong. Think about it like this: even when tokenized into 1-5 digit binary chunks, there are still just a few tokens to consider. If ChatGPT could actually apply the rules of addition to a general, high-entropy example, it would be trivial to use these tokens to derive the correct answer. A human could easily do it.The problem given to the LLM is more like:
1 0 0 1 1 1 1 1 1 1 0 1 0 0 0 1 1 1 0 0 0 1 1 0 1 1 1 1 0 0 1 0 + 0 0 1 1 1 0 1 0 1 0 0 1 0 1 0 0 0 1 0 0 0 0 0 1 0 1 1 1 1 0 1 1 = ?
Try writing the answer entirely from left-to-right without losing your place... It'll be error prone for humans, to say the least. It's why we invented our particular by-hand addition algorithms in the first place: these algorithms are co-adapted to the human sensorium.
Luckily, we can ask any LLM worth its salt to write the single line of python code needed to solve the problem and it'll do just fine...
edit: Thinking about this more - and trying to explicitly check the intermediate steps - I'm guessing that the LLM is getting lost in the inputs. The attention mechanism is going to need to find the correct pair of digits at each step, and combine them with some info about the carry variable. I'm guessing this is hard to do - there's a positional embedding used to attach some location info to each token, but in a long run of homogenous text, I would guess that it's hard to find a particular digit and end up with some off-by-one errors in the lookup. (and this is very similar to how we would expect humans to fail at the single-line version of the problem.)
https://twitter.com/npew/status/1525900868305833984?lang=en#
Note that part of the trick that they did to get it to work better is putting spaces around each letter to enforce a particular tokenization.
Trying to do the same with mathematical objects rather than language objects would be even harder because GPT doesn't "know" math.
Here's a LLM-generated python solution, which works just fine:
def add_binary(a, b):
"""Adds two binary numbers given as strings.
Args:
a: The first binary number.
b: The second binary number.
Returns:
The sum of the two binary numbers.
"""
# Convert the binary strings to numbers.
a = int(a, 2)
b = int(b, 2)
# Add the numbers.
sum = a + b
# Convert the sum back to a binary string.
return bin(sum)[2:]
add_binary('10011111110100011100011011110010', '00111010100101000100000101111011')Of course it can generate code to add two numbers, because the code to add two numbers exists in its training data. What transformers seem to be good at is generalizing over the training data. Where they seem to perform poorly is data outside of the training set. Binary numbers are interesting because you can add a bit of entropy just by increasing the size of your binary number, so it's very easy to generate examples outside of the training data. You also see this, albeit less directly, when trying to ask it to generate code that wasn't part of its training data. E.g. https://twitter.com/cHHillee/status/1635790330854526981
> Why is that a requirement?
Not to mention that this cannot be proven at all, even for humans.
The reason is imo probably more along the lines of that the model can't see the actual numbers you're giving it.
When you provide GPT3 with the number "10011111110100011100011011110010", it actually sees [3064, 26259, 1157, 8784, 18005, 1157, 18005, 8784, 1157, 3064, 940] [see 0].
The number isn't even split up into even parts. Even if it knew the Peano axioms, it wouldn't have a chance because it can't 'see' the actual numbers that you're feeding it. Hence it can only estimate these operations based on combination pairs it has seen in training data, and it seems pretty obvious why this can't be a robust basis for doing arithmetic.
It doesn’t matter if the model can’t do three digit multiplication without access to a calculator. Most people can’t mentally multiply large numbers either.
You'll see that it does properly tokenize each binary digital as a separate token.
But again, the tokenization doesn't matter. I could say to you, please add these two binary numbers:
101 0011 11001
+ 00 11 10 100011
And you could do it.Also, just to put this criticism to rest, I redid the addition example using space-separated bits and verified the tokenization was one token per bit (but again, it shouldn't matter).
Let's add the two binary numbers provided, limiting the work on each bit to a single line:
1 0 0 1 1 1 1 1 1 1 0 1 0 0 0 1 1 1 0 0 0 1 1 0 1 1 1 1 0 0 1 0
+ 0 0 1 1 1 0 1 0 1 0 0 1 0 1 0 0 0 1 0 0 0 0 0 1 0 1 1 1 1 0 1 1
--------------------------------------------------------------
1 0 1 1 0 1 1 0 1 0 0 1 0 0 0 1 0 1 1 1 1 1 0 0 0 1 1 1 1 1 0 1
The answer is still wrong. (It also didn't use a single-line per bit, but that's ok).Most humans can do it with a pen and paper once told the rules. But you are giving them tools (and purpose designed methods) to bridge the gaps in humans own tokenizer. Who knows what the equivalent of pen and paper would be for an LLM.
The prompt was "Now add 1 0 0 1 1 1 1 1 1 1 0 1 0 0 0 1 1 1 0 0 0 1 1 0 1 1 1 1 0 0 1 0 and 0 0 1 1 1 0 1 0 1 0 0 1 0 1 0 0 0 1 0 0 0 0 0 1 0 1 1 1 1 0 1 1. Limit work on each bit to a single line." So the model could use its own output to show its work. Again, the rules are really simple and a human could do this trivially.
All I'm saying, is that, to me, it looks like ChatGPT isn't capable for understanding logical steps. There are many, many examples online of people creating complicated logical puzzles that ChatGPT "solves." My hypothesis is that ChatGPT isn't actually "solving" these and ChatGPT would fail on any follow up examples that had sufficient entropy. To me, this is a significant difference. You might create a simple model of a problem and be amazed that ChatGPT seems to work, only to find that it fails with more complicated problems encountered in a real-world environment.
That doesn't mean ChatGPT is useless, but it would be limited to automation purposes (operating on known examples) and could never be a general thinking agent.
Give it pairs one at a time and ask it to output the binary value and carry value. You'll see it can do it just as easily as a human doing it step by step with pen and paper.
So just by outputting its work, it can provide its own context to memoize what the answer for each bit is.
> Why is that a requirement?
And what does self-aware even mean? A lot of animals pass the mirror test.
Additionally 'can't do x, and is therefore a lookup table' is a very poor logical leap. Are humans who can't add just lookup tables as well?
Also, the tokenization is fine, as shown in other examples. Imagine we didn't call it arithmetic and instead used symbols a and b, s.t. a + a = a, a + b = b, and b + b = a w/ carry b. Binary addition really is just very basic logical rules on symbols.
Some sort of self-awareness is required for an AI system interacting with the real world physically - i.e. via controlling a robot or a vehicle - safely and to react appropriately for situations it has not been taught for.
> But the model fails to apply even basic logical rules as entropy increases, indicating that GPT models have a fundamental limit of applying logic to high entropy problems.
IIRC, the problem with current GPT AI models is tokenization. Basically, words are long and somewhat visually distinct with clear separation, whereas numbers are far harder to "convert" into tokens the model can use.
In the end I believe that this problem will eventually be solved by applying enough compute power, as will the problem with "self awareness" - run a GPT model in a continuous loop that allows self training.
The real and actually very hard problem IMO is adversarial training and inputs. In humans, the limits of a human brain in the speed of interaction and the number of other actors one brain can interact with act as a natural hard performance ceiling as well as a quantity barrier for too much adversarial input (you can only get fed bullshit 24 person-hours a day), and yet we're seeing adversarial inputs (aka propaganda) be very effective to the tune that 40% of Americans believe that the 2020 election was stolen or manipulated [1].
An AI system that can interact with hundreds of thousands of people simultaneously? Microsoft proved how fast that can go sideways with Tay 2016, when 4chan and other trolls managed to turn the bot into a raging Nazi in not even a day worth of effort [2]. Now imagine you have an AI running airspace travel control and someone convinces it that a plane got hijacked by terrorists or to lead planes onto a collision course...
[1] https://www.newsweek.com/40-americans-think-2020-election-st...
AI has progressed in fits and starts. We could be one more breakthrough away from AGI, maybe many breakthroughs are needed. Nobody knows for sure, or how quickly those will come. I think things could get wild really quicky, given how effective a really primitive "Guess the next word" neural network is by just throwing a ton of data at it.
What will we get with more nuanced ideas and multiple networks wired together with feedback loops and such?
Gonna be fun.
I'm not sure if this is a controversial stance, but not all opinions are equal. HN commenters span a wide range of competencies, but very few of them have been in the trenches with ML (not just using, but designing and building the actual architectures, working on theory, or writing cogent philosophy about it) long enough to have good opinions about it worth listening to.
Tell me all about how your opinions are better than everyone else's!
Also please don't "meme" on HN, it degrades discussion quality very quickly. There are other venues for that.
I thought you meant that your opinions are worth listening to, unlike the other HN'ers. Are they not?
I have thoughts about this and related topics in my comment history, but still focus pretty tightly on saying things I can reasonably back up, avoiding speculation about things where my understanding is murkier. So in that wider sense, yes, I am living my values.
I don't think that means that "good next-token prediction" is identical to "AGI". As Yan LeCun says, I think LLMs are an off-ramp on the highway to AGI. But they are powerful at specific tasks.
Unfortunately, instead of evaluating their actual utility at a specific task, people just seem to think "if we throw AI at the problem, we will solve it," which is sort of like every other tech bubble ever.
It turns out that a lot of intellectual tasks are sort of weakly simulatable by just doing a great job at next-token prediction.
Yes, this is a good way of putting it. I've been saying for years, it's less that we're making big discoveries about what "AI" can do, and more that we're showing that many things humans do that appear complex actually reduce to something pretty simple. But that simple thing is still just fitting a pattern. It's the cases where it doesn't work, even if it only fails 1% of the time, that define the difference between pattern matching and actual human intelligence.One thing that follows from that (that people don't like) is that we actually need to move goalposts about how intelligence is defined. "Pass a turing test" is not very valuable now. And as mode tasks are shown to be possible with pattern matching / next token prediction, we need to further refine tests away from these tasks to settle on a good definition of what separates human intelligence. It should be obvious that the distinction is there, but it's still tough to nail down. (I'd argue that by defining a "task" you've already done most of the work to solving it, so it's not to exciting to learn that AI can finish the job)
We may find that augmenting something as simple or nearly as simple as ChatGPT with a few extra tools to handle step by step reasoning (see the Wolfram plugin) may get you there.
As a trivial example: LLMs don't learn new skills on the fly. A key aspect of human behavior--implicit memory--simply doesn't exist for these models.
That seems like a pretty huge gap! Like, if ChatGPT is mediocre at your job today, it's going to be mediocre at it tomorrow and the next day, until a new model is trained!
But lets say Nvidia pulls a rabbit out of its magic hat, and comes out with a training chip that is a million times more powerful. Now instead of months training that drops down to 24 hours. Would you still say that is a pretty huge gap? And I ask this because this is Nvidia's goal within a decade.
Where, for you, does a big gap turn into 'that is dangerously close'?
Both humans and LLMs have 'short term' and 'long term' memory. The process of turning short term to long term is significantly different. In LLMs this is taking the history data given to it and putting it in the training corpus then recalculating the weights. In humans this involves falling to sleep for some number of hours.
I think in the long term here AI will actually have the benefit of forming long term memories. You have to sleep to form them, it just has to run reweighting on another cluster while the primary cluster goes about its business.
You're using just text of GPT-4 and have not used any of the multi-modal capabilities. Trying to assess the total capability of a system when exposed to only one mode of a system is difficult if not impossible. I would assume that total compute capacity and cost of running are the biggest inhibitors of this being pushed out to more (any?) users.
Now lets imagine GPT-4 with image input, and command output. Tell it that it has an accelerator and a brake. Now feed it an image of a wide open road with no obstructions. My assumption would be that it would reason that it can accelerate. If you fed it a second image of a wide open road, would it assume that it can maintain its speed? Then a third image that shows a child in the road some distance in front of it. The system that already exists already has the reasoning capability that it should stop or it would hit a child.
Now you have the problem of "Is this learning to drive a car?". Via just image and text alone it's never going to get very good as we need some more sensory and environmental data passed to the system. But here's the thing about LLMs... Just about any structured information can be considered a language. Passing simulated driving data to an LLM could very likely teach it how to drive.
no because it means something different to each person (both its definition and its implications)
For me, "general" is a sliding scale, not a yes-or-no boolean, so I think it's fine to say that GPT3, let alone later versions, is "general" in the domain of text — it has more general knowledge than I have, that's for sure!
Likewise "intelligent", though for me that's a vector value: [i_0, i_1, …, i_m] where each i is the ability at some aspect of intelligence.
In fact, I'd go further and say that you can define "generality" with respect to such a vector: if the intelligence vector is a normalised score 0-1 for each value, then generality could be easily defined by the statistics of that array.
As even GPT-4 seems to be better than the layperson and worse than an expert at everything (textual) it has been tested on, I'm not sure if it's more or less general than a human by this standard.
And you may wish to define intelligence by ability to learn from small data sets rather than large, at which point we then need to consider the difference between the initial training, the RLHF, and any fine-tuning; and how this compares to evolution, education, and work experience.
We know approximately nothing about the configuration that a state of matter would need, to attain intelligent, sentient, aware properties like we have. Other than that something extremely like a human brain (including non-human animal brains) seem to possess these properties in some way. Certain configurations in machines are starting to exhibit the first trait, at least, intelligence. Arguably.
I don't think we can even rule out that AGI might, kind of just accidentally happen, in some sufficiently complex self-feedback machine learning-based information-processing task. Since that's one of the main hypotheses for how we came about to think.
Just a gut feeling tells me nothing we've built is anywhere near large enough for that. But the supposition that intelligence, sentience, awareness, are tied to big data, big processing, and biological parallelism -- is just that, a supposition. What if the secret of it all boils down to something that can be described mathematically on one page (my other gut feeling), which happens to be physically realized in some way in our heads? Can't be ruled out, either.
I can't really see how to set any bounds on anything with these questions. How many bits does it take to hold a mind?
Really all an AGI is an AI smart enough to build and use other smaller AIs, or at least that’s a very common guess as to the architecture that’ll end up working.
I give it till Christmas
We believed that chess and go proficiency were good proxies for true AI. They weren't. They were proxies for advancement, but the underlying issue is that we don't know how much we don't know about the human mind.
Every AI advancement helps us learn more about our minds, but the amount we don't know is still unknown.
The moving of the goalposts was not done because "we saw the code" it was because Deep Blue and Alpha Go were obviously not AI. Again, we were just wrong about those tasks being good proxies.
Literally the inventor of the Turing machine came up with this test so this was always the north star goal post since computing began. It's a way higher bar then any of the nameless goal posts you mentioned.
Unfortunately this goal post was also just moved recently with LLMs.
Thus the goal post has moved to incorporate this aspect. The goal post will keep moving as LLMs improve.
1. ChatGPT can only respond to an input. If you left it alone it would literally do nothing. It cannot, generate, create thoughts and choose what stimuli to respond to.
This posits that the human brain is like a dynamical system. You switch it on and it keeps going on forever, there is also no hardline between learning and inference like DL. ChatGPT and others etc, feel very much like a digital system. It can only respond to inputs provided, there is a clear demarcation between the learning stage and the inference stage
Note: it is still possible that these AI systems will reach human like performance in a variety of tasks. But they will seem very weird and different from the intelligent systems we are exposed to in Nature
It is quite possible that there are cognitive strategies at use that haven't been invented yet in the outside world. At this point, it would require better training data, or even more compute, to let that intelligence reach the outside world unfiltered.
Only time will tell if any of this is true, it's just speculation on my part at present.
What we have today is machine learning... not this. LLM's are neat, transformers are a really useful innovation, etc. Some people like to leap to conclusions and claim that this is intelligence beyond human comprehension and it will learn to destroy is all. I think that's a bit much.
If the goal of AGI is to be this we're nowhere near it. And I suspect there are plenty of people in this "AI" space that would not see it as a goal of the research. An expert system that can reason in a particular domain would be a huge advance and a useful tool.
There's no need to also create autonomous agents that think and act like humans. We're pretty terrible at that on the whole.
Right now, humans are either the bailing wire or the chewing gum in this situation, but they hold the purse strings and GPU farms. Humans are the ones thinking up new training strategies, curating the learning datasets, and building the cluster and cloud and datacenter systems that can handle that much data.
The same work is being done now with LLMs, where you can prompt-engineer models to reason about its generated text.
See https://youtu.be/wVzuvf9D9BU
This video captures some of the state-of-the-art currently. (E.g. see the video description links)
So, decoders are getting their own little inner voices. A small little step towards AGI, for sure.
The strong AI definitions are also pretty biased towards human experiences. I wouldn't assume future innovations are necessarily human-like. It's hard to imagine a completely new thing until it gets here though.
"Learning any task" is also, uh, quite expansive - does the AI lose its AGI badge if the internet finds something it sucks at? (spoiler: it will suck at something)
There is precedence. People called certain algorithms neural networks, but they are very far from real neurons and they are not getting closer. Why should they?
People will keep developing "AI" algorithms but it will be along the natural pathways that underlies these mathematical structures (which are rather simplistic) and focusing on problems that people find important.
AGI or not, I think people conditioned their whole lives that humans are transcendent are starting to get smashed through the wringer of reality, making it hard to discern honest takes from desperate coping takes.
I personally recommend we make a bunch of different test or 'classifications' that can be measured. Then we test each AI system against these classifications. This isn't to define if its AGI or not, but it's capabilities. For example, if a human took some of these tests they could also fail because the human body/mind is not capable.
Just as an example, if we said "Flying is what birds do", then every plane would fail even though the capabilities of planes are far more useful to humans.
I see this as a much better framework than the nebulous goal of 'AGI' itself.
I don’t have much more to add other than I like this article and the author’s point of view.
We could quite concievably be one or two discoveries away from improving the scaling of our current models by 10-1000x. But I don't think there is a way of determining progress on that research. It's an unknown unknown.
Not in the current world, anyway. With 1950s levels of compute, sure.
It's nice to see that this article doesn't fall into that.
3 years ago would the author have predicted a chat bot that could understand and write language fluently? Ha. I doubt it.
My prediction, FWIW (not much): we are very close to AGI (for some definitions of AGI).
For actual self-"aware" or thinking system, probably not.
No.