We've been here before: AI promised humanlike machines – in 1958
theconversation.com
theconversation.com
So I feel outside the overhyping (some hype is warranted but this much?) could be less, this is different; it’s not a promise, I am chatting to my computer on my lap as if it was my human assistant; it gives me answers faster and usually better than most humans would. That’s simply not the same as we had in the previous instances. Knowledge systems did this too, but were encyclopedia’s; they couldn’t come up with stuff that wasn’t explicitly put in and categorized. You often controlled them by answering a tree structure of multiple choice screens (there are some nice episodes on the Computer Cronicles about this on Archive.org). Now you can have this knowledge base in your (vector) db, ask open questions in any language about it; it will answer human-like, lies/bluff and all.
"I would tell Marvin this whole AI thing is just horrible. Why are we doing it? And he would say it will be effective for getting our grants, so shut up and just play along and in fact, it was true. It was very good for getting grants back in the early days, in the 70's was when I started at this, you would go to a grant making organization such as the Defense Department, you'd say we're going to build this super smart thing and if we don't do it, our enemies might and it'll get smarter than people and it was Okay, here's your money, here's your money, oh my God, you better... and so it was very effective. And actually, the whole thing started off as storytelling to get grants"
Currently this doesn't get as much media hype for some reason. But definitely the military is all in on the current hype-wave.
Because is it really 'hype' if an AI fighter pilot can beat human F-16 fighter pilots? This has already been done and old news.
Games, poker, go, StarCraft, -- all have 'game' logic that can also be used on battlefields.
Robots, logistics, fighter plane drones, battlefield 'perceptual' analysis, etc...
There are companies building these things now, they aren't hype.
Just like Darpa Funded the internet, in the future when the military has huge AI components, we'll be looking back on today fondly like the good old days before there were thousands of autonomous drones tracking us all the time..
I'm not sure speed running the internet forum experience by using data centers with 10x the energy requirements of traditional ones is a net gain.
Google already is at 90%.
The margin and market motivation to do so is high and those few companies have a lot of money.
Nonetheless, we talk about an industry / society change; A change how we will write and operate and talk to computers.
I woudl say its worth it
Energy is fungible. Exploding marginal energy requirements put enormous pressure on the grid tho. In the US several gas burning plants are being planned in the last few months given AI data center growth projections. This will have a definite impact in total emissions and push back any goals of holding on to the current 2C increase in global temperature.
The percentage of datacenters around the globe is very small (like 1%). In contrast this saves us a tremendes amount of co2 due to optimization of logistics, doing your bank business at home, calling someone instead of meeting them etc.
Datacenters are in my opinion the biggest net positives for co2, are easy to make green, have the most money behind them (which means faster and better investments) and are ooperated by our leading tech companies, who will use this to push further the industry of green energy.
Those datacenters and especially AI energy investments are also the biggest research advantage we have. Optimizing solar energy gains and storage capacity. We need them for simulating/generating new materials, production processies, etc.
We need to do a LOT in regards of heating. Heating is critical.
I think its critical for our society to put more into R&D of materials and others for more and faster optimization of solar and batteries.
Nvidia for example does a lot now in omniverse, simulating the real world. Its also potentially co2 saving if you simulate your full car, warehouse etc. digital and iterate over it super fast and opitmize it before it ever creates any co2 in real life.
You make a fair point. To the extent these AI tools enable better, more efficient production systems in the real world there is a case to be made it could be a net gain for society. Arguably it could also increase the carbon intensity of the economy in the short term. While renewable sources are gaining ground, most of the bulk and marginal energy demand is met by carbon heavy sources now and in the foreseeable future, and deployment of renewables also requires a lot of energy and by extension, for now, carbon emissions.
The EV consumes more energy at the beginning, solar panels and wind turbines too. Unfortunate its hard for people to get 'economy of scale' and its super frustrating that we have the investment<>expensive<>benefit hen<>egg issue.
Heat-pumps, EVs, solar and batteries could become even cheaper even faster if we would invest faster and more. In 10 years those have eclipsed every current alternative.
What i think is a good example is Alpha Fold: The graph on this page https://www.moltenventures.com/insights/a-breakthrough-in-pr... shows the jump alphafold provided.
Now tx to alphafold2 a huge library exists for all researchers. And i have seen many other breakthroughs.
Segment anything from facebook is a LOT better in image segmentation than what we had before. This makes it much faster for everyone having segmentation tasks to segment faster.
Wispher is really good in speech to text. It basically beats a lot of old school software on the market.
> we have the investment<>expensive<>benefit hen<>egg issue.
Maybe it's easy for people to get economies of scale AND path dependence. Unfortunately society has been put in a trajectory that maximized the profits of minerals rights holders in the industrializing US.
And yes, we are in a sub optimal local minimum in terms of efficiency and need to go over a hump to get to better minimums and going over that hump may mean increasing carbon intensity of the economy for a while.
Do we have time to do that though? Should we carbon de-intensify the economy and aim for a trajectory that minimizes climate related shocks or should we go all out on an accelerationist hope that we can bootstrap a better system by running the current one red hot? These seem to be the two sides of the debate we're in.
The main reason why i get paid well is, that everything i do, i normally not do for one person or a single company.
Energy which could be used for other things. I don't say it's bad to use power for AI, but just saying "well, it's green energy anyway" is short-sighted imho. At least as long as there are still any things not yet powered by green energy.
The investment required for those data centers will act as stabilization of green energy investment and because known software companies are involved, for a better support on the software side
I think in the long run, LLMs will be packaged with the OS (because only organizations like Microsoft and Apple have the resources to burn on them) and will run on local hardware (the consumer bears the cost of the computation).
I am interested on if it will have diminishing returns for some roles, if the outputs aren't novel enough over a long period of time. Eg: it's used to write marketing copy, but I already see the "feel" of it getting criticized and a lack of novel tone when used daily.
It’s obvious that modern AI is far, far superior to ELIZA or 50s perceptrons. It reminds me of people who said space travel will never be possible, or heavier than air flight might be developed in ten million years.
I think intelligence is similar; the work of yesteryear pointed towards the possibilities of "flight", LLM's provided the same sort of boost as the space race of the 50's and 60's. I am not sure where the analogy ends and if a "moonshot" will happen - or if things will stall at some point. At the moment I am thinking that a stall might happen. Auto regressive techniques have given us a very very clever hans, reinforcement learning seems to have disappeared from sight, graph networks are giving us fantastic tools (weather, biology) but nothing that appears or claims to have cognition.
Keep in mind LLMs are explicitly trained to avoid claiming they have cognition.
It isn't a technology problem, it is a funding problem.
I do believe artificial self-awareness is possible but I doubt we're at that point already. And until it reaches a certain level it will remain indistinguishable from patterns in training data.
There aren't even objective, empirical definitions of “self-awareness” or “reasoning” that would support claims of examples. Those things are subjective internal experiences, not verifiable phenomena.
Goal-post moving? Sure. But I think it just betrays a deep discomfort with what has been achieved. Kind of like the meme that it's all just a giant plagiarism machine.
And there's still a lot of opportunities to get more human like today: it can't empty the dishwasher or drive a car yet even if you'd give it arms, legs and eyes.
It doesnt help that the automated future imagined in the past would "free up" people so that they could spend all of their time writing songs and poetry.
But yes; I remember Ruby and the Galactic Gumshoe, the Andorrians were a race of artists who insisted that automation-induced unemployment was not a disease, it was the cure. When I think of the most secure jobs of the future, these days I think, toilet plumber and janitor.
There is always 1000 things you can do for other people. The question is just do those other people have the money to pay for it after rent and bills.
Increasingly they do not.
This video [0] at 1:25 suggests that it can. At least in some near future.
Of course, they don't show it moving around. But arms and hands seem like better than other demo videos.
Looks to me to have a good enough understanding of the "real progress since then".
>the Perceptron would lead to machines that “will be able to walk, talk, see, write, reproduce itself and be conscious of its existence.” More than six decades later, similar claims are being made about current artificial intelligence. So, what’s changed in the intervening years? In some ways, not much.
Which shows the author's complete lack of understanding of the basics. In ML reasearch, it has become very clear that compute is everything and architectures are actually not that important. So what has changed? And why is today fundamentally different, despite us still using the same perceptron architecture with a few tweaks? Look at the Kurzweil curve [1]. The available compute power has allowed existing architectures to flourish beyond what was possible in the 60s. Do you think it's a coincidence that computers in the early 2010s were suddenly able to distinguish pictures of cats and dogs reliably, when that was utterly impossible just a decade earlier? Or perhaps you think it's weird that GPT3/4 capability level models arrived in the 2020s? Just look at the curve (which was made in the early 2000s btw) to find out why.
[1] https://images.squarespace-cdn.com/content/v1/5ec0712f88e816...
Hate to say "no, you're wrong", but… this is just not true. The latest LLM stuff only came after the invention of the transformer architecture. After the development of a new architecture, we always have a "hey, if you throw a supercomputer at it, it works EVEN BETTER!" phase: think Deep Blue, which was just alpha-beta search with a fitted evaluation function on a really powerful machine.
I see no reason to believe that transformers doing text prediction are a generally-applicable AI architecture. I suspect they're even incapable of many things, no matter how big they get: GPT-4 remains incapable of the things I claimed GPT models were incapable of back in the GPT-2 days.
>GPT-4 remains incapable of the things I claimed GPT models were incapable of back in the GPT-2 days.
Which are...? Anything we have seen so far from scaling LLMs is that capability is purely an issue of training/fine-tuning. It's not a question of the architecture.
Quantitative improvement over state-of-the-art is nigh-irrelevant. You can drive that up arbitrarily high, for almost any given metric, just by throwing more compute at it (with a few architecture / dataset decisions to push it in a particular direction). That's not the kind of thing I consider an improvement: it doesn't make anything possible that wasn't already possible.
The measure of an AI system is fundamentally about doing stuff, but primarily about doing stuff efficiently. Computers have always been able to play chess by exhaustive brute-force search, given a galaxy-sized GPU and a few million years.
Thank you for confirming my point. Because this is what it's all about. It might not seem like much if you only think on timeframes of startups, but in the long run this will beat any architecture improvement. Yeah, maybe you can come up with a fancy revolutionary design and achieve instant 50% improvement over SOTA. You will get all the attention of the media and prizes and stuff. But how often does that happen? On the other hand, you could also just silently wait one or two generations of Nvidia cards and get the same thing. And even better, this also works when there is absolutely no improvement in architectures by anyone. This is only highlighted by transformers, because they came out in 2017, but only now have they been able to solve e.g. the bar exam.
Scale > architecture is the uncomfortable reality that all ML researchers eventually come to terms with.
> Scale > architecture is the uncomfortable reality that all ML researchers eventually come to terms with.
Architecture -> scale. Current LLM architecture doesn't scale very well, you need massive increase in compute, memory and data for tiny improvements. Likely there are much more efficient ways to train these models, it shouldn't be this hard for them to learn arithmetics.
Yeah, "scale > architecture" assumes that those are indepedent, whereas architecture affects scaling reach, needs, and options.
Also would a Perceptron with the original architecture do what an LLM does just given scale?
Yes. The answer is yes. This is exactly what the universal approximation theorem [1] says, which was proven decades ago. It is also just another line of reasoning that compute matters more than architecture, but fewer people outside of ML research know about it. And it's also less practical, since the formal theorem has few empirics to go by for really large models. But mathematically, a scaled perceptron is 100% able to replicate any LLM.
[1] https://en.wikipedia.org/wiki/Universal_approximation_theore...
What the universal approximation theorem tells you is that there's some algorithm you could use to train a (sufficiently-large) perceptron model to behave arbitrarily-closely to the equivalently-trained LLM – which, yeah, of course there is. There's a way of training an array of floating point numbers to do that, too. It's called the transformer architecture.
(Your misunderstandings are similar to the sorts of misunderstandings I had before I got in the habit of trying to (abstractly) implement every CS concept I came across. Kudos for trying to reason about AI from first principles, but you have to do a lot of reasoning before you stop being much wronger than the empiricists.)
Oh really? And what would that look like? People have literally been looking at this for decades. And yet we're still using gradient descent with a few tweaks. So I would not bet on anyone coming up with something revolutionary. Especially not while hardware regularly revolutionizes the capabilities of models trained in this ancient way. Chipmakers have always allowed us to overcome the hurdle created by our lack of advances on the training architecture front. No matter whether you look at convnets or transformers.
I'd expect better from them than from technologists who actually understand the technology, becuase the latter typically have unreserved praise for their creations and complete blindness to the negative effects of technology on society.
Edit: actually read the article, and it just seems to note the cyclical nature of AI, which the lay person may not be aware of.
Outside our own field, they look amazing and insightful; within our field, we focus on the "wet pavements cause rain" level of mistakes that they make.
I am one of those. However, distinguishing between truly autonomous AI systems and those that are simply high-performing can become challenging. The line between the two might become blurred when performance reaches a level that seems autonomous. In this context I believe that another "AI winter" will be hard to identify in the future.
There is some tendency to discount the current state of AI because it doesn't match the level of hype - ignoring the fact that hype is non-linear and it makes no sense to speak of an appropriate level of hype anyway. All that stuff is just pop culture. If you can look past it and objectively view the current state of the art through the lens of 10 or even 5 years ago (gpt-2 era) IMHO you'll see we have enough progress to last 10 or 20 years without declaring a "winter" even if from here on out progress will be slow.
It didn't even get to its own point (which is also wrong - we definitely haven't been here before).
But this time is for real. /s