I can run art models and llms on cpu/gpu now. I've tested out opensource models with quality better than chatgpt3.5turbo, and I can even fine tune them on my notes and books for better results. It's all so easy with so many one click installers now too!
My husbands D&D group uses some AI for their games now (koboldcpp?).
It's staggering how fast it's moving.
Do you have a good resource to find this stuff? I’m behind the times.
(Also, to be a pedant, most of that is inference and not training. But I can’t say much about fine tuning so I’m not really trying to argue against your point.)
It's real work to keep up on everything.
> A large focus of the GPT-4 project has been building a deep learning stack that scales predictably. The primary reason is that, for very large training runs like GPT-4, it is not feasible to do extensive model-specific tuning. We developed infrastructure and optimization that have very predictable behavior across multiple scales. To verify this scalability, we accurately predicted in advance GPT-4’s final loss on our internal codebase (not part of the training set) by extrapolating from models trained using the same methodology but using 10,000x less compute:
> Now that we can accurately predict the metric we optimize during training (loss), we’re starting to develop methodology to predict more interpretable metrics. For example, we successfully predicted the pass rate on a subset of the HumanEval dataset, extrapolating from models with 1,000x less compute:
> We believe that accurately predicting future machine learning capabilities is an important part of safety that doesn’t get nearly enough attention relative to its potential impact (though we’ve been encouraged by efforts across several institutions). We are scaling up our efforts to develop methods that provide society with better guidance about what to expect from future systems, and we hope this becomes a common goal in the field.
"just trying to temper investor enthusiasm"
"trying to downplay AI threats to calm down regulators"
etc....
etc....
But it is not some 'Proof', that LLM's have reached a limit.
It is a self reported note along the lines of : "nothing to see here, we're at our limit, it's all good, stop probing us".
That’s my personal take on the current wave.
Though it'd be hard to copyright any code small enough to fit on a flash card.
As someone who have been very anti-crypto for a long time, it wasn't always a complete scam.
The first wave of the crypto boom, before anyone that wasn't a programmer had even heard of it, there was a lot of real work being done that very much mirrors current AI work. Lot's of very sharp developers learning about block chain, figuring out how to implement things, experimenting with ideas. Back then everyone owned their own wallet and you would meet at coffee shops to exchange cash for BTC.
Most of the serious engineers that were really into crypto during the first crypto boom of 2012 left in disgust when the second boom came around.
Having worked in AI/ML for a long time, I myself can start to see how they felt. We do have some really cool technology in front of us, I think it has a lot of potential, but so many of the loudest voices in this space are entirely out of touch with what's possible, and far more interested in hype and making money than the underlying technology.
That isn't true.
You can use it anywhere that irreversibility matters. Suppose you're going to commit significant resources to the customer's request, so you charge them, commit the resources, deliver the goods, and then discover that they gave you a stolen credit card and you get a chargeback. Cryptocurrency avoids that.
You can use it to accept payments from all over the world. Someone in Asia or Africa may not be able to open a US bank account or get a US credit card, but if they can find a Bitcoin ATM to put their local currency into, they can pay you, or vice versa.
It allows you to pay for something over the internet without giving your name. There are situations where this is important.
The main impediment to using it is, ironically, regulatory. The IRS decided that it's an investment and not money so every time you want to use it for what it's actually supposed to be for, they treat it like a securities transaction where you have to fill out paperwork, even if you're just buying a pack of gum. Which makes it much less convenient for ordinary people to use than cash or credit cards which don't require this -- presumably on purpose in order to destroy its utility in the US.
But it can still be useful for people in countries that don't do this, or in the US if a less explicitly antagonistic regulatory environment could be established.
I was not talking about crypto. Blockchain “solutions” in enterprise were spinning up all over the place for non-crypto applications. In particular, in you work financial, supply chain, government or random startups, you probably heard blockchain a lot in non-crypto contexts.
It doesn't follow that when new tech arrives, the post hype everything goes back to status quo. Typically after the hype, the new tech just grows or gets absorbed a little more quietly, in un-foreseen ways, and does end up having a big impact. Just when the impact is stretched out a little, people stop noticing.
Like replacing drive through ordering with GPT like tools. Kind of under-radar, not fancy, not flashy. At some point you'll notice that the drive through you are talking too isn't a human, and go 'huh, that's interesting'. But, big impact on jobs, so nobody is hyping it.
People/Companies are already using AI to replace or augment Graphic Artists, Coding, etc... etc... That is happening, not just a power point from a middle-manager.
How can we extrapolate that to be "well, gosh darn, these LLM's are already played out, guess we're all done"
Live systems in nature seem to solve similar problems with way less compute available. There should be better architectures.
Also, as somebody said, every exponential growth curve is a lower part of a sigmoid. LLMs will plateau at some level. That level may be impressively high though.
Do they really? They're certainly more energy-efficient in business-as-usual mode, but a human brain has 86 billion neurons, 600+ trillion synapses(!), and each instance takes 15-20+ years to train to do complex logical tasks. Even if the per-cell work is tiny (and, is it? cells are amazingly complex), 86 billion (or 600+ trillion) times 20 years is a lot of computation.
Show me a humanoid robot that can do that.
It's definitely not several tries. I've actually witnessed that very recently with my daughter. It takes a child literally several thousand attempts and falls (around a hundred falls per day) and 1,000 hours of dedicated practice, before they can make their first step without holding anything [1]. And a lot more until they can walk reliably.
[1] Source: Becoming You, S1E2, from 21:30
Don't get me wrong, I'm not claiming that Machine Learning is as generally capable as animals/humans, and I don't entirely disagree with the OP (i.e. I don't know if our current approach has a chance of scaling to human-level capabilities). I just don't think computation-wise it compares that badly to animals, considering that it's the result of a few decades of work, largely on repurposed silicon
This is an interesting paradox in machine learning: the data that people use to (ostensibly!) simulate the human mind is simultaneously way too much, and not nearly enough. It's way too much in that humans can, e.g., learn to generate and recognise natural language without having to train on the entire World Wide Web of a corpus. And it's not nearly enough because humans can do that only after billions of years of evolution, which amount to training on all the data in the real world, not a mere few petabytes of data on the internet [1].
This paradox has to be resolved (i.e. understood) before we can really compare humans minds to artificial systems. Unfortunately, I don't see anybody trying. In machine learning so far people are happy to just plug in the data and turn the crank- something will come out eventually.
______________
[1] Source: https://archive.li/EY9ak
Human babies are born at a stage that would be considered premature for most mammals. Because of the limitations on the size of the head, they can't afford to develop enough in utero, like deer or even dolphins can.
The puppies are benefitting from a different type of "learning" process of course, in having genetic instincts tuned by many generations of ancestors who had good enough genetic wiring of motor capabilities and instincts sufficiently well attuned to their environments to survive. Humans have weaker priors and reach physical and mental maturity much more slowly but ultimately achieve greater mastery of their environment by learning it from first principles or communication with other humans (there may be a tradeoff, and of course humans are so good at protecting their young over years of infancy the early instincts aren't that important to retain)
As mentioned in a sibling comment, HN anthropomorphizes AI too much. And is too optimistic about it. I just don't see the results and the value, people trip up all over themselves to congratulate themselves and the researches, yet 99.99% of the problems in the world persist.
I am one of these a-holes that wants to see results when money are invested. It still comes as a shock to some apparently.
AI takes a shortcut around having a timeline of a trillion different versions history optimizing the learning algorithm and putting the successful ones in a timeline, but at a cost of trillions of operations to teach it.
TL;DR, the problem space is huge and evidently must be iterated.
It doesn't. We are steering it. And I doubt that we know better than the natural evolution.
Its a joy to watch a child grow up, but also its super interesting watching them figure out the most basic shit. Would highly recommend if you get the opportunity.
Bringing Up Turing’s ‘Child-Machine’ by Susan G. Sterrett
https://link.springer.com/chapter/10.1007/978-3-642-30870-3_...
That's not OK.
These and exactly these are the tech advancements that I imagined as a teenager would improve people's lives!
I just feel that 99% of the technology nowadays does not improve almost anyone's life. It gets easier and easier to be jaded and lose hope.
But your story is very reassuring and nice to read. Thanks for sharing!
"Walking" at least on dry land has been around 420 million years or so. The locomotion part of the learning algorithms has been around a very, very long time, and it's going to be very hard to duplicate the efficiency that nature brute forced.
Nature has not hype optimized human level intelligence. Really human intelligence is at the limits of birth. Our heads can't be any larger at birth or we'll kill the mother too much of the time. Then the human body has a power budget that's been the defining factor of human/animal survival for just about forever. Humans were already 'smart' before we invented agriculture and excess fuel storage.
AI takes that a step further. We're not using reversible computing so the heat generated is pretty high. We've not been doing it very long and our algorithms are not optimal in any sense. But what we do have that nature does not is nearly unlimited energy, nearly unlimited cooling capacity, and a lot of computer systems working on the problem at once. In an evolutionary sense we're doing closer to the bacteria method, we have a lot of different experiments going on at once and at least a few of them are likely to be fruitful.
Personally I hope the AI development is no insanely rapid. Rapid and powerful changes in society can cause as many problems as they find fixes for, and the pace of society change is not that fast.
Not sure if the AI area, represented by human supervisors, does so in such an objective and ruthless manner as nature does.
If there's no fitness function and natural selection then the AI area might be doomed to meander and go in circles for centuries.
Also, human brains do a lot more than language processing.
LLMs are definitely touching something very important about intelligence, much like counting sticks touches something very important about numbers. But the real power of numbers is unleashed with the invention of digits and positional notation, a different representation that makes things literally exponentially easier.
I hope there is a transition step from what we do now using billion-parameter models, to some better representations, more compact and thus more powerful.
Also I hope that ready-made logic and efficient numerical computation can be connected to the learning systems more directly, not taught painstakingly from first principles, much like vision is relatively directly wired into the human brain.
What is that in kilowatt hours? Human brains are remarkably energy-efficient with their compute. We should get credit for that. Comparing one of us to an AI being trained at a data centre with the energy budget of a small city isn’t really fair, is it?
For that setup cost, you get a single-threaded human, starting to specialise in a single field. They will work in that field for around 7.5 hours every weekday, with ~20 days of holiday a year, for around 45 years, and then retire.
Llama v2 70B cost 1,720,320 GPU-hours[1] at 400W, so 688MWh. Once trained, it can be run 24/7, and you can spin up as many instances as you want on much lower spec hardware. That model produces output faster than a human while consuming around ~30W on my Macbook Pro.
Now, I know we're in pretty shaky spherical cow territory here, comparing a human (albeit a highly educated one) to Llama v2 in logical reasoning... but consider that this is the state of the art in generally-available LLMs after a few decades of research into machine learning & we're using repurposed silicon to compute vs the amazing complexity and physics-leveraging approach of human neurons... and the training cost is only 1 order of magnitude off humans.
Again, I'm not disagreeing with the general point of the OP, I'm not saying that the models we're using right now are the best/right ones (or that the hardware they're running on are the most efficient way of executing them), but I don't think the energy efficiency gap is actually all that high considering
Edit: If you look at the total footprint (considering not just the efficiency of the neurons, but the whole animal) the figures are very close - the llama model card indicates that v2 70B caused the emission of ~300T of CO2, and an average human in the US emits 16T of CO2 a year, so a human would emit ~320T of CO2 in 20 years. I assume children don't have as high a CO2 output, but even so it seems like it's the same order of magnitude.
1: https://github.com/facebookresearch/llama/blob/main/MODEL_CA...
The best use of LLMs that I've seen so far is as a boilerplate-producing autocomplete system. Considering that we have better ways to automate this (better programming languages that can abstract away the boilerplate), this is not very high praise.
Edit: If you look at the total footprint (considering not just the efficiency of the neurons, but the whole animal) the figures are very close - the llama model card indicates that v2 70B caused the emission of ~300T of CO2, and an average human in the US emits 16T of CO2 a year, so a human would emit ~320T of CO2 in 20 years. I assume children don't have as high a CO2 output, but even so it seems like it's the same order of magnitude.
If you're going to include the whole animal on the human side, you need to include the whole supply chain on the LLM side. The cost of building all the fabs and doing all the R&D to develop and manufacture model training-specific computers (matrix multiplier hardware). Just like with crypto, these resources had to be diverted away from other things (e.g. causing the price of gamers' graphics cards to skyrocket). It's only fair to interrogate the ROI.
I think it's in our nature as software people to look at their ability to work with code, but they're quite good when applied to general language tasks. I've been using them for summarisation and reading comprehension and they're quite effective. I've also been working with a teacher friend on seeing if they can generate well-scoring essays on highschool English essays (as always, the problem is prompting and context).
On code, GPT-3.5 (moreso GPT-4) seems to have a good ability to generate and translate smaller scale code problems (hundreds of lines in low-boilerplate languages) but yes, they're like an eternally junior engineer whose work you're constantly having to oversee for subtle bugs, and I don't know that it actually saves time.
I'm pretty sure people are working on different approaches to applying them to code, with better prompting+context from larger codebases, and multi-step processing (i.e. rather than just a single prompt->response, letting the model iterate through a few steps independently, possibly guided by other adversarial/supervisor agent instances, testcase generators, etc.)
>If you're going to include the whole animal on the human side, you need to include the whole supply chain on the LLM side. The cost of building all the fabs and doing all the R&D to develop and manufacture model training-specific computers (matrix multiplier hardware). Just like with crypto, these resources had to be diverted away from other things (e.g. causing the price of gamers' graphics cards to skyrocket). It's only fair to interrogate the ROI.
That's fair, although it gets complicated to work out numbers because we don't train many LLMs, whereas we're constantly training humans, each of whom cost the planet tons of CO2 emissions every year... and, of course, your point that LLMs just aren't very good yet. I fear that they're good enough (or appear to be to the layperson) that execs will replace customer support staff with them, even if the outcomes overall aren't as good.
I'm not as fearful. When customer support gets too expensive companies already outsource it to India. Indian customer support workers cost far less than Westerners in terms of energy and CO2 emissions, both in terms of training and ongoing costs. India is about 2T/person in CO2 emissions, compared to 15T for North Americans. And that number is averaged over the whole country. I would imagine the poorer areas of the country have much lower emissions and the bulk of CO2 comes from the wealthier big cities.
"oh, they are at their limit". People were saying they were at their limit 5 years ago too.
There are a lot of AI techniques, some more advanced than LLM, just with other limitations so they don't look 'amazing'.
Seems like engineering at this point to tie them together for the next 'big thing'.
The drumbeat of progress has been quite steady. On log charts.
How much compute is too much to be spending on this bullshit. 30% of humanity's total computing resources? 70%?
GPT was a big innovation in parallelizing the training process. I think the optimism that we're on the beginning of a sigmoid curve here ought to be questioned.
I do think we still have yet to squeeze the most value out of current LLMs, but most people's radical AI dreams are completely out-of-touch with reality for anyone working closely on these problems.
My biggest fear in this space is that disappointment in the inability of these tools to live up to the hype will cause people to irrationally abandon exploring the spaces where they do work.
Which is countered by...the assertion that it won't?
LLMs won't get intelligent. That's a fact based on their MO. They are sequence completion engines. They can be fine tuned to specific tasks, but at their core, they remain stochastic parrots.
> I want to know only one thing, which is what gives him the confidence necessary to say that.
I want to know only one thing, what gives the confidence to say otherwise?
No it's countered by principled restraint in not making an affirmative claim one way or the other.
I've heard this referred to as the overconfident pessimism problem. Which is that normal, well founded scientific discipline and evidence-based restraint go out of the window when people declare, without evidence that they know certain advances won't happen.
Because people get mentally trapped into this framing of either have to declare that it will happen or that it won't, seeming to forget that you can just adopt the position of modesty and say the dust hasn't yet settled.
LLMs are a huge stride forward, but AI does not progress like Moore's law. LLM have revealed a new wall. Combining multi agents is not working out as hoped.
Many of the things AI does now are exactly the type of things that doomsayers explicitly predicted would never happen, because they extrapolated from limited progress in the short term to absolute declarations over the infinite timeline of the future.
There's a difference between the outer limits of theoretical possibility on the one hand, and instant results over the span of a couple new cycles, and it's unfortunate that these get conflated.
(Btw., a really hardened pessimist might even say your example is mostly things that were doable when there were not many computers around at all... taxi driver to drive you, a secretary to book a flight, meet in a club to discuss things, ...)
It’s maddening to me that people don’t get this nuance
First of all because so many predictions of the end of the world have been made and we tend to have a hunch about the kind of person who makes them. Which is stereotyping, sure, but at least it's a heuristic, not a straight-up abandonment of evidence-based thinking.
I agree there's uncertainty about the future of AI development, but it's true that we have no idea how to create AI, right now, so the uncertainty is about whether it will happen, not whether it won't. If that makes sense.
The evidence for this, albeit empirical, is the history of AI development itself.
AI doesn't show continuous development over a long period of time. It always developed in steps. A new architecture or method is discovered and able to solve some previously hard or unsolveable problems.
Then this solution slowly develops in capability, mostly based on better and cheaper hardware, while it's quality plateaus.
It may be that LLMs will not follow that pattern. I don't think so, for reasons outlined above. But until it can be shown that they don't, that this really is the long-sought-after AI architecture that just gets better and better over time, I don't think that a healthy dose of pessimism is unwarranted based on history.
I kind of agree. However, I see a real possibility that in the near future LLM behaviour would be practically indistinguishable from intelligent/sentient behavior. And at that point we (or at least I) are facing some really interesting/difficult questions, namely how do you know an intelligent looking thing actually is intelligent (or sentient). How do you prove me you/LLM are/aren't a philosophical zombie?
How we are supposed to treat very much intelligent/sentient looking things when we are not sure if they are sentient/intelligent or not? Let's face it, lots of people are dumb as rock (too often very much me included). Why we should be able to treat something badly just because we think we know they can't be intelligent, even if they walk , look and quack like intelligent duck?
I personally have started to think that the behavior of humans should be judged by the behaviour, not the target. If you want to behave like an asshole towards a teddy bear, then you most likely are an asshole.
Cute, but practically speaking, I would prefer the former.
There is much more to the abilities of human body-mind-emotional-experiential being, but it is only slowly becoming mainstream.
(Edit: Of course there are also many analytical skills that AI cannot match at this point. My point is that we shouldn’t overlook any area of human capacity.)
One salient question in this is: will we reach a level of intelligence where we become beings capable of actual collaboration that doesn’t waste so much effort in conflicts, or one that is capable of living in harmony within its environment?
What capabilities of awareness, trauma work, emotional maturity and self reflection does this require? What resources hidden inside humanity that we have forgotten do we need to wield?
Does AI have something to contribute to this process happening?
IMHO the main weakness with LLMs is they can’t really reason. They can statistically guess their way to an answer - and they do so surprisingly well I will have to admit - but they can’t really “check” themselves to ensure what they are outputting makes any sense like humans do (most of the time) - hence the hallucinations.
(They asked GPT-3.5 and GPT-4 "are you sure" to see if it would change its answer, both when the original answer was right, and when it was wrong)
Maybe it’s because it can’t tell when it’s wrong and needs to “try again”, and we have to do it for them.
Could it be that intelligence is overrated and discovery of new ideas / thing is underrated? Our egos tell us it's intelligence that makes us special and creative and awesome but maybe most of the special stuff is already there for us to find and we conflate discovery with extrapolation. Maybe knowledge and experience are the "important bits" of intellect.
Example: Einstein didn't really invent anything, he discovered things about the world that blew our mind and changes our lives. Yes he was a great thinker and a courageous soul to go against the grain and he had the balls to be open minded enough to discover new things. We obviously believe Einstein to be intelligent but was he just a great explorer ?
I have a similar attitude towards technological progress, yes we've done amazing things but fundamentally the air we breathe, the water we drink and the beauty we are subjected too when looking at a sunset are taken for granted while we stare at our phones.
It kinda depends on how much you care about people across the world and future generations.
But yeah, if humans were more intelligent, we probably would have sorted all this out by coming up with better coordination mechanisms, and by overcoming our tribal tendencies more effectively.
I see this too often, more intelligence = positive outcomes, but no, some of the smartest people ever put their intellect towards stupid causes, such as oil exploration and AI to capture peoples attention.
I hope you're right though and I'm wrong ;)
That's more about ethics and an wish for moral behavior and conflict aversion, than about intelligence.
Intelligence (human and AI) could just as well opt for conflict and evil, if this helps it get the upper hand for its own private goals and interests.
Simply put, the interests of the collective, are not necessarily the interests of the individual intelligence.
(Even assuming there was a single, easy to agree upon, "interest of the collective" for most problems).
We have moral philosophy as an academic discipline, after all.
Human brains develop in interrelation. Much, if not all of our intelligence gets developed in relation to other humans and beings.
That's neither here nor there.
Having the intelligence "to reason about good behaviour in an effective manner that's collectively beneficial" doesn't mean you're constrained to reason and act only on that, and not also able to reason and act on behavior which is beneficial to you to the detriment of others and the collective.
And it's pefectly intelligent to follow the latter if you can get away with it, and if the benefit for you is more than your share of the collective benefit alternatives would be.
>We have moral philosophy as an academic discipline, after all.
And how has that been working out for us?
(Not to mention, keyword: academic).
>Human brains develop in interrelation. Much, if not all of our intelligence gets developed in relation to other humans and beings.
Yes, and a lot of it is devoted to duping and getting the upper hand of those other humans and beings. So?
This is absolutely wrong. There is nothing about their MO that stops them from being intelligent. Suppose I build a human LLM as follows: A random human expert is picked and he is shown the current context window. He is given 1 week to deliberate and then may choose the next word/token/character. Then you hook this human LLM into an auto-GPT style loop. There is no reason it couldn't operate with high intelligence on text data.
Not also that LLMs are not really about language at all anymore, the architectures can be used on any sequence data.
Right now we are compute limited. If compute was 100x cheaper we could have GPT-6, bring 100x bigger, we could have really large and complex agents using GPT-4 power models, or we could train on tupled text-video data of subtitles videos. Given the world model LLMs manage to learn out of text data, I am 100% certain that a sufficiently large transformer can learn a decent world model from text-video data. Then our agents could also have a good physical understanding.
If one feeds GPT-4 a novel problem that does not require multi-step reasoning or very high precision to solve, it can often solve it.
In many cases, we humans have structured our language such that it encapsulates reality very closely. For these cases, when an LLM learns the language it will by construction appear to have a model of the world. Because we humans already spent thousands of years and billions of actually intelligent minds building the language to be the world model.
But in a sense when YOU learned language YOU also learned a world model. For instance when your teacher explains to you the difference between the tenses (had, have, will have) you realize that time is a thing that you need to think about. Even if you already had some sense of this, you now have it made explicit.
Why should we say the LLM hasn't learned a world model when it's done what a kid has done, and everyone agrees the kid understands things?
From what I see, there are some things it hasn't learned correctly. Notably with limbs, it doesn't know how fingers and elbows work, for some reason. But it does know something about what they should look like, and so we get these hilarious images. But I also don't see why it shouldn't overcome this eventually, since it's come pretty far as it is.
By doing this you'd glue the words to sights, sounds, smells, etc.
But it also seems like this is already someone has thought of and is being explored.
You can very easily give "evidence" of gpt4 being anywhere between emerging super-intelligence and a naked emperor depending what you ask it to solve. They do not learn models of the world, they learn models of some class of our models of the world, which are very specific and already very restricted in how they represent the world.
Of course they are, they haven't been trained on anything spatial, they've only been trained on text that only vaguely describes spatial relations. A world model built from an anemic description of the world will be anemic.
In my experience, things outside coding quickly devolve into something more like "technobabble" (and in coding there is always a lot of made-up stuff that doesn't exists in terms of functions etc.).
It's almost like we need our AI's to have two brain parts. A fast one, for intuition, and a slow one, for correctness. ;-)
For some industries where I understand the cost stacks with lower and higher skilled workers, I'd say it only takes out the "cheap" part and thereby not taking out a large chunk of costs (more like 10% cost out prior to paying for the AI). That is still a lot of cost reduction, but something that also will potentially be relatively quickly be "arbitraged away", i.e., will bleed into lower prices.
An example that exists today would be the combination of ChatGPT and Wolfram [1], in which ChatGPT can provide the method and Wolfram can provide the execution. This approach can be used with other systems for other domains, and we've only just started scratching the surface.
Once we have AI's running around with the creativity of artists, and the precision of logicians, ... Well, time to read some Iain M. Banks novels.
You're glossing over a shocking about of information here. The problems we'd like to use AI for are hard to find correct answers for. If we knew how to do this, we wouldn't need the AI.
Incredibly poor compared to ours, but thousands of times better than what "AI" we had before.
At night when he is awake (he sleeps in our room in a covered cage) he knows not to vocalize anything more "Dear" when my wife gets up - he says nothing when I do this as he is not bonded to me.
When I sit at my computer and put on my headset he switches to using English words and starts having his own Teams meetings.
When the garage door opens or we walk out he the back door he starts saying Goodbye - Seeya later and then does the sound of the creaky outside gate.
But the fact is that the loss goes down predictably with increased compute budget, data and model size (see Chinchilla Scaling Law). We've also seen that decreased loss suddenly results in new capabilities in discontinuous jumps. There is all reason to believe there is still some juice left in this scaling, exactly how far it can be taken is difficult to tell.
A system that could perfectly predict what I would do in response to any particular stimuli, as a continuing sequence, would be exactly as intelligent as me.
> They can be fine tuned to specific tasks, but at their core, they remain stochastic parrot
Othello GPT was an attempt at answering this exact question, it's a simplified setup and appears to learn a world model: https://thegradient.pub/othello/
That's certainly interesting but it's not a depiction of a LLM is it ? LLM's are not deterministic, and (perhaps) so are we so two non-deterministic systems can only occasionally align (or so I assume). Intuition says they may get "close enough", whatever that might be, and close enough is good enough in this case but I think you are making a giant assumption to the likes of since we can speed up matter to 1000km/h then IF we sped it up to light speed then ...[something]...
Most inference samples from that distribution using a composition of sampling rules or such, but there's nothing stopping you from just always taking the most probable token (temperature = 0) and being fully deterministic. The results are quite bland, but it's perfect for extraction tasks.
(note: GPT-4 is not fully deterministic; there's no details on this but the running theory is it is a mixture of experts model and that their expert routing algorithm is not deterministic/is dependent on the resources available)
I'd argue everything about a LLM is artificial, there is no natural process involved is there ? Since its design is to mimic us (at face value, though I don't know how fair of a description this is) then randomness is essential I think.
They definitely can be, but it doesn't matter.
> but I think you are making a giant assumption to the likes of since we can speed up matter to 1000km/h then IF we sped it up to light speed then ...[something]...
This is an odd comparison.
The point here is that a sequence prediction system can be as intelligent as the system it's predicting unless you invoke woo. That doesn't make llms intelligent but it means the argument that they just predict the next thing isn't enough to say the can't be.
I think this sentence doesn't mean much unless we have a strict definition of what intelligence means.
Just today ChatGPT helped me solve a DNS issue that I would not have been able to solve on my own in one day, let alone an hour. I'd consider it already more intelligent than myself when it comes to DNS.
A dictionary contain knowledge but no intelligence.
You can argue those are facts too.
If I asked you a question and you had to respond with a stream of consciousness reply, no time to reflect on the question and think about your reply, how inaccurate would your response be? The "hallucinations" aren't a problem with the LLM per se, but how we use them. Papers have shown that feeding the output back into the input, as happens when humans iterate on their own initial thoughts, helps tremendously with accuracy.
It’s more likely it’s just, once again, generating the most probable answer - and if you shake the magic 8 ball enough you will get the answer you were expecting.
Yes, I'm saying here that peoples' inner voices are hallucinating in very similar fashion; "rejecting fictional generated output that makes no sense" is a process that's consciously observable and involves looping the inner voice on itself.
But still, you don't have to agree with me about what intelligence means. But it is important in these discussions to understand that not every participant shares the same definition of the term intelligence.
intelligence to me has to be based off initiative. In that sense, a dog or a cat have more intelligence than GPT.
I actually very much look forward silicon (or other non-biotic material) attaining intelligence, I consider that the only way that Earth civilization can colonize space. But this aint it.
There were some Greek guys working on that exact problem a few (thousand) years ago.
Would you consider a search engine, or a book, to be as intelligent?
We're going to see exponential increases in processing power of the best GPU clusters and human brains are a stationary target. And there is precious little evidence that the average human is much more than an LLM. LLMs are already more likely to understand a topic to a high standard than a given human.
They're going to progress and if they aren't intelligent then intelligence is overrated and I'd rather have whatever they have.
We found that extrapolating the performance given a few data points with smaller models is actually very accurate. That's how they determined hyper parameters, by tuning them on multiple smaller scale models and then extrapolating. So far, all those predictions were quite good.
Together with a bigger model, we also need more data to get better performance. If we add video and audio to the text data, we have still a lot more data we can use, so this is also not really a problem.
It would be very unexpected that those scaling laws are suddenly not true anymore for the next order of magnitude in model and data size.
I still expect the final solution will be more along the lines of picking the best model(s) from a sea of possible models, switching them in and out as needed, and then automatically reiterating as needed.
Yet, we are different, right?
Where's the proof that sequence completion engines can't be intelligent?
Even assuming that is true: LLMs aren't all that exists in AI research and just like LLMs are amazing in terms of language it's possible similar breakthroughs could be made in more abstracted areas that could use LLMs for IO.
If you think ChatGPT is nice, wait for ChatGPT as frontend for another AI that doesn't have to spend a single CPU cycle on language.
All the models being used in academia are basically toys, none of those guys are running hardware at a scale that can even remotely touch Azure, Meta, etc, and right now there is a massive global shortage of GPU compute that's eventually going to clear up. We know models get A LOT better when they are scaled up and are fed more data, so why wouldn't the same be true for other problems besides text completion?
Frankly, I'm a bit worried about all the rest now that LLMs proved to be so successful. We might exploit them and arrive to a dead end. In the meantime, other potentially crucial developments in AI might get less attention and funding.
Remember that UFO poster? “I want to believe”.
But at the basic level - isn't our own brain just a sequence completion engine too?
Hallucinations to the normal person are a bug.
The issue is that only humans can hallucinate. We know there is a “reality”.
For an LLM, everything it does is a hallucination.
That’s why you have more POCs than production goods. Your “hallucination rate” is unknown.
Yesterday Ars has an article that described LLMs as a new type of CPU for a new type of architecture. Others want “LawLLM” or “healthLLM”.
These are simply not going to happen.
If you even get it to 80% accuracy- it’s a 1/5 chance you have a relevant answer.
The issue isn’t a technical one.
It’s expectations.
GitHub copilot is really good and useful.
All demos I saw which use LLMs were spectacular.
The ai race started this year for everyone which means we will continuesly see progress.
And while you only mention LlM the whole ai space is crazy.
There is a high chance that the architecture from LLMs will change.
And we haven't even touched all possibilities with multi modal LLM models.
In the past months I have used Gen AI to create multiple proof of concepts, including labelling and summarization tools. In addition, to make sure I took a project to conclusion, I built a website from scratch, without any prior knowledge - using Gen AI as extensively I could.
I am being pretty conscientious with my homework. The results of those experiments are why I am confident in this position. Not just because of the articles.
I am also pointing out that its not the tech, its the expectations in the market.
People expect Chat GPT to be oracular, which it just cant - the breathless claims from proofs of concepts fan the flames.
I leave it to you to recall the results and blame, when unrealistic expectations were not met.
Aren't demos always spectacular?
All American programming tech has relied on an time-and-knowledge gap to keep big companies in power.
Using visual studio and c++ to create programs is trivial or speedy if you have a team of programmers and know what pitfalls to avoid. If you're a public pleb/peasant who doesn't know the pitfalls, you're going to waste thousands of hours hitting pointless errors, conceptual problems and scaling issues.
Hallucinating LLMs are marketable to the public. Accurate LLMs are a weapon best kept private.
I am always intriguied by the people who say LLMs provide a massive benefit to their programming and never ever provide examples............
Hallucinations are random, the truth isn't.
This is absurdly expensive with GPT4, cheaper with 3, and dirt cheap locally with LLaMA
Anecdotally I tested this by having GPT4 translate Acadian cuneiform — which it can just barely do. I had it do this four times and it returned four gibberish answers. I then prompted it with the source plus the four attempts and asked for a merged result.
It did it better than the human archeologists did! More readable and consistent. I compared it with the human version and it matched the meaning.
Expensive now… soon to be standard?
The only way you could know that the output was wrong was because you could verify it in the first place.
You can’t verify answers for questions in unfamiliar domains - or even for novel questions in your own domain.
Hah, it feels like a weird version of P!=NP.
The main issue is context length: if you use 4 attempts you have to have to fit in the original question, four temporary answers, and the final answer. So that's 6 roughly equal sized chunks of text. With GPT4's 8K limit that's just 1300 tokens per chunk, or about 900 words. That's not a lot!
The LLMs with longer context windows are not as intelligent, and tend to miss details or they don't follow instructions as accurately.
Right now this is just a gimmick that demonstrates that more intelligence can be squeezed out of even existing LLMs...
Truth is a human thing. Statistically averaging out 4, 5, 6, N text generations from an LLM will not converge to any “truth”.
You have essentially stated that outputs from a text generator are normally distributed around “Facts”.
May I gently suggest, that and older quote about an infinite number of simians, typewriters and the works of Shakespeare, is more appropriate ?
The question I have is what's a prompt which reliably hallucinates but still produces the correct answer some of the time?
I know it gets some python functions "wrong", but i think they were actually "right" in the version it was trained on, so software seems out.
Sounds like a pretty good guesstimation. Well, not cease, just fizzle out.
Like putting a field of artists out of a job, or copying a style so good that you can complete a persons piece before they do on a livestream.
We're using a massive mush of the internet model which is taxed at 50% for alignment. That's going to be very dumb in the long run.
To be frank, between stock art and photos, pre-AI template based tools, and the massive oversupply of graphic design and photography work, the field was already massively redudant and kind of out of a job to begin with...
Those guys apparently used to make pretty good cash from twitter, usually using pen names so they wouldn't' be associated with their regular work
The scary bit here is you can also "clone" a person to make any image you want of them. Obviously there's a lot of problems coming from that in the future, but also neat applications, e.g. some guy made selfie pictures of himself in the past with this for internet dating.
It can also simulate a Zizek vs Wittgenstein argument over Russian literature. the fact that it can usually write executing computer code is nothing short of magic to me.
What one would want is something like LLaMA2-Coding-Vue2, maybe with a LoRA for the library or concept.
I wrote a big word salad about all the other things you could do but suffice it to say there's no reason anyone should be limited to a single pass on a monolithic without automatic context window augmentation nor automatic code checking & regeneration & model escalation (e.g. query out to a 200B coding model or something).
I'm javascript centric these days, but the same concept should work for most languages with a package manager.
Some of these features could even be integrated into say yarn/npm directly, you currently have a devDependencies section of a package.json..I could imaging having something along the lines of a "llmDependencies" section to define which main model and version to use and which "library" models to use.
If you don't pay for ChatGPT, you get GPT-3.5. You can also get access to GPT-4 if you use the playground.
I asked it (actual names changed):
"I run the Linux command line program "foo". When I use the flags -xyz, I get results, but when I use -txyz I get nothing. What could this mean?"
And it told me: "The lack of results is because you didn't use the -t flag".
Or I ask it some very basic music theory questions and it gets stuff wrong all the time, giving impossible answers.
Do you have API access? the old model there still gives me very good results.
I'd kill to have an easy, wont-get-me-banned-way to submit a query to both the UI and API at the same time and show the results in, say, meld or so.
When there's one definitive answer to something that people keep repeating there's a slight chance that it's actually true. Shocking, I know.
But hey, if you are really looking to convince yourself of something, I have no doubt that it can be done.
I keep going around telling people that 1+1=3, why do they always give me the same nonsense about the number '2'?
I blame Sam Altman.
2. Obnoxious? This is my opinion, I don't get what Sam Altman or 1+1=3 is supposed to mean in this context.
We do a lot of experiments involving gpt3.5, 4, claude-v2, titan-large, and palm2, and for what it's worth, on our real production workloads gpt4 shines. We can make Palm2 produce decent results with a lot of extra effort, and claude-v2 is passable but gpt4 does not disappoint. This is low-grade knowledge management stuff, and we are not using it as a information-retrieval system - but for basic 'cognitive' tasks where all the information needed is provided in the prompt. I'd not rely on it for info retrieval tasks such as the examples quoted above - its knowledge base is highly compressed, after all.
Tools have limitations.
How do you get that out of an LLM? What tool is any good if it doesn’t work 100% of the time predictably?
It’s funny those the LLM haters keep raising the bar to a level that no other software can reach. Chatgpt is a tool like any other and just like a hammer, it can be misused, or it can be incredibly useful if used well. I personally find chatgpt mind boggling and astounding and use it every day multiple times for both coding and non coding purposes. But it’s totally normal and reasonable to me to expect bugs, do you really expect visual studio to run on large solutions and never have crashes or memory issues or slowness? If so you’re going to be disappointed.
Query many times on the right model(s) for the question and the correct answer will be there 99.999% of the time as the other hallucinations will be thrown out.
I find that no examples often leads to the same result as a befuddled junior, but with examples often it gains confidence. Also I find playground to give me much better code snippets than chat sometimes
There's a lot of "extraneous" information and details that appear useless and unrelated, but that long-time developers have tucked away in their brains (or really anyone who has done something at a high level for a long time), that turns out to be incredibly useful; generally these people don't require examples - they just know what the right answer is, because they've been exposed to a variety of problems over a long career or lifespan.
That's where LLMs need to be to be truly "useful". I think if we get them to that point, we'll really have something useful on our hands.
Both young and old shall feel the pain then.
I think the future is extremely well-trained base models that are then fine-tuned on specific domain knowledge. I'm already seeing that with Meta's Llama 2 personally at home. My company has access to Azure OpenAI Service's GPT-4 trainable model that we've been working with on all our documentation, and the results look very promising so far.
But, this idea of the All-Knowing Oracle that can answer any question you have is all well and good, but it's so far beyond our current processing power as to be a pointless endeavor at the moment. Will we get there eventually? Yeah, sure. According to Jim Keller as of around 2019/2020 when he was speaking on Lex Fridman's podcast, he said we still have room to go 1,000,000 times smaller - meaning chips. I think we probably need two more orders of magnitude in processing power of the strongest high-end GPUs before we're there. NVIDIA's H100 is a great achievement, but we need something about 100x as fast, and I think we'll still need dozens of them working in parallel to train the model that can answer all these questions.
All of this though is a moot point, because it's the lawyers that have already slowed down progress. OpenAI is terrified of being sued. Altman couches it behind terms like, "AI Safety", "AI Alignment", etc., but it's fear. It's all stemming from fear. And it's all stemming from people just not "getting it".
We're entering a new age of upheaval, and there's going to be rogue AIs that tell people to go kill themselves to reduce climate change. You know why? Because there are humans that tell people to go kill themselves to reduce climate change. These models are language models. We taught them how to think, and we taught them how to think like us, so it's no surprise to me whatsoever that they behave like us - meaning they occasionally lie and they occasionally go off the rails and go a little crazy.
We have become gods and we have made a creation in our own image. Most of the time it's awesome, sometimes it's a little wild and wacky.
However a few times there have been some mechanical refactoring-style grunt work I've delighted to have been able to let ChatGPT do. However, the rate ChatGPT is giving me subtly wrong results is just high enough that I end up cross-checking everything, and then it takes me a bit more time than it would've otherwise taken. Give it a year or two, maybe?
Maybe? It’s not clear what would bring a qualitative improvement, barring massive amounts of new training data.
In other words: It assumes Marx was right about the central failure-mode of capitalism (an inability or unwillingness to find ways to distribute work and proceeds of work in ways that prevents productivity gains from eventually causing social upheaval).
There's no reason what they describe as "so-so" tech can not also be a significant societal advantage, but it requires structural change to ensure it does not deprive a significant number of people of a livelihood.