Scoring well in a benchmark that's called AGI does not make an LLM AGI.
Scoring well in a benchmark that's called AGI does not make an LLM AGI.
If so I'm hoping we can track them down and have them tell us if they think this is AGI.
If I can't give it an arbitrary task and have it solve that task eventually, it's not a general intelligence.
(obviously it might take years for me to get good enough at something, or if you set the "arbitrary" task as something ridiculous, but lets work in good faith here and think of something the average human could do after learning about it)
If we progress to the point where an LLM instance can meaningfully learn to get better at something overtime without retraining, then I will accept that is basically AGI. Right now, they still seem to be pretty boxed into their training, even if you can prompt them to act differently.
Of course, it's a moving goal post, because we have no clue what general intelligence is exactly. But it's definitely not general yet. Now the goalpost is to achieve that kind of level of thinking which I did in that 4 hours. When it reaches it, we will find something else it clearly lacks. Until we can't. Then, and only then we reached AGI. Until you see comments, reviews, etc about things which it cannot do, until then it's not general.
Once we have 1000 tps, i am sure robots etc.. will also start working like magic.
It's magical to me as well, but I don't feel like it's AGI.
Because in my experience a Senior Programmer does not need the right prompts to deliver the right outcome! :-)
For example it should be easy to tell it to not write a paper in the style of a clickbait SEO article or use all of its stupid hallmark AI writing patterns “it’s A, not B!” And a smart human that would be told that would be easily able to comply with that but the model needs to be told in a very detailed way and it seems to lack even basic capabilities to reflect on this, when explicitly given a sentence it will be able to rewrite it but otherwise it’s mostly blind to it. That’s to me a hallmark of it being overtrained on the specific tasks or problems so it appears very smart but once you go off script it still shows that it’s not a “real” mind.
Of course it’s amazing and has super human capabilities in many areas but if you honestly think it’s better than Einstein like some people suggest why can’t it write a simple “good” academic paper even after giving it specific examples and instructions.
Maybe that’s what makes these things dangerous, they have super human capabilities in some areas but apparently lack self awareness, taste and meta reflection abilities. The only reason people aren’t afraid more is that they don’t act in the physical world yet, imagine giving it a body, superhuman strength and letting it care for your child when it has a strong “urge” to comply with your exact request and little to no self awareness and human basic instincts.
"You're absolutely right to call me out on that. I shouldn't have stopped the baby crying by killing it, that's on me."
An AI model that’s human-level at programming is an incredible achievement. But it isn’t general intelligence. It’s highly specified intelligence.
Or you'd ask it to add a new page to your website and shout at it to use your existing brand colours instead of inventing some and realise it's not AGI at all...
This honestly doesn’t happen to me much anymore. In what areas do you find LLMs routinely make stupid mistakes?
It's objectively very difficult and technical, it's spatiovisual, it's artistic, learning resources for it are sparse and most just learn by the FAFO method, current AI sucks terribly at it, and it's not likely to ever be specifically targeted by benchmaxxers.
Or, as someone else points out in another thread here, academic writing. It's one of the things newer models seem to have actually gotten worse at. Even when you give them detailed instructions on how to write and what to avoid, the "load-bearing", "A but not B" and journal-like writing make it in anyway, with the supposed AGI having no ability to reflect on how blatantly unacademic (and often unreadable) its writing is.
- Created useless pydantic schemas with all fields Optional[Any]
- Created a REST endpoint that silently mutated on GET (unsubscribed users from a mailing list)
- Failed to log costs in my app so users could have bankrupted me, etc, etc.
Good job I actually review its code.
Only the translation and language understanding capabilities are enough to be impressed, and they are 2 year old already. Now, the AI do see, draw, speak, listen, think, work, etc.
Someone from the 90's would simply not believe that the AI would be a machine but would think for sure that a human is behind. The only odd thing would be that this human would both exhibit high intelligence and stupidity at the same time.
If anything, the fact that it is so powerful is almost a concern, because I think we are still way underestimating what these systems will be able to do when we give them more cognitive capabilities.
At the moment we are something like, having had great success with propellers and have promised we will fly to the stars.
People love to say 'this is the worse they will ever be', then extrapolate to conclusion that they will continue to accelerate at the same rate of progress of last few years .. it may, maybe, or we will hit a ceiling, might be a temporary one, could be 5 years or 50 years ..
"Homer, you can't just declare Artifical General Intelligence; you need to like, make something or something...mmmmrrrhh"
In a closed a press briefing earlier today, OpenAI co-founder and president Greg Brockman offered an unusually direct formulation of that message, ending the session with: “Welcome to the AGI era.”
"""
But if you’re asking when a model has a sustainable general intelligence, for me, it’s pretty easy…
When it makes financial sense to run it 24 hours a day.
It makes either position pointless to argue.
Aren't we way way past that already? QPS to any of the frontier models for a given point in time is most likely (far) greater than zero.
Directly - something can be useful without being AGI.