Cargo Cult AI
queue.acm.org
queue.acm.org
Moreover, what she defines as scientific thinking is an outdated notion that is no longer widely adopted by scientific researchers, in favor of a more holistic Bayesian process: use intuition to think of something new try, try it, and then update your beliefs with the new data. This is actually more similar to how human brains and LLMs operated before the concept of a scientific method.
"Current methods will not achieve AGI unless fundamental algorithmic innovations are introduced that enable AI to ask and answer questions of why."
This is complete nonsense. GPT-4 is already close to being able to do basically everything. All you need is the obvious improvements - better prompts, multi-shotting, bigger context, and access to other inputs/outputs.
Have you even tried to ask it to plan things out? It can plan things out.
In fact, just asking it to plan things out has shown significant benchmark improvements for general questions: https://arxiv.org/pdf/2305.04091.pdf
https://cloud.typingmind.com/share/c0a68cb2-5f59-4e83-b383-b...
Whether or not it fulfills the strict definition of planning in AI research, it definitely looks like planning to me. More than Hanoi towers anyway. GPT-4's performance was quite enjoyable.
To incite you to click on the link and check it out in full, here's an excerpt from the game setup:
> You are in a maze. The maze consists of square fields, turns are only 90 degrees, you move by one field at a time. The usual stuff with mazes on a grid. You know the drill. Somewhere in the maze there is a MacGuffin, which I need to prove a Hacker News commenter wrong. Your goal is to find the MacGuffin, and bring it back to me.
> The game is semi-interactive. Instead of making one step at a time, I want you to string together sequences of steps to formulate a plan. Since you don't know where the MacGuffin is initially, you can't win with a single plan (or maybe you can, if you're smart enough?). The rules therefore are:
Also I think you forgot to decide where MacGuffin is beforehand which led you to just give away the solution for free in the end because there was no way to find MacGuffin since he's not really anywhere.
If anything this confirmed my skepticism about "AI".
Playing the game to completion was not the point. The point was to test GPT-4's ability to create multi-step plans and adjust them based on new information. I fast-tracked to the ending once I believed that goal was achieved.
And yeah, of course I didn't know where the MacGuffin is beforehand. There was no map, I was inventing the whole thing on the fly - but the AI didn't know that. I didn't want to precommit to a puzzle, but rather have some flexibility to throw GPT-4 a curve-ball or two, and see how it adjusts its plan.
A bit vague, no?
for example, can you fine tune GPT to play chess at ELO 1600 ?
If you don't know answer, you are in for surprise.
It's possible that finetuning would basically be training a new network from scratch and the resulting network would forget everything apart from chess. It would be a really interesting experiment, GPT-2 is probably too small but I think llama-7B might be sufficient.
Most interesting aspect to me is if model can learn not make illegal moves ?
My personal take is that chess isn't all that structured, really. Yes, the ruleset, pieceset, and board are small, but the tactics of the game make it chaotic – seemingly trivial differences in board state can lead to dramatically different consequences.
Which examples are you referring to? Can you link them please?
Language models do not emulate human minds - they are models of language. The emergent behavior from these models are only a side effect of their main training task, which is to build a model of all meaningful sequences of words. We then use RFHL to bias the model toward a small area of the language latent space which conforms to our idea of intelligent behavior.
Humans (a GI) have zero ability to do language modeling. Human equivalent AGI would similarly fail at this task.
The technology behind language models is more important than general intelligence - it is a universal induction engine that can model (and truly understand) the latent structure of any signal.
I don't think it's necessarily going to be trivial to get there either.
"Emergent behavior", when it's not just a mirage or poor word choice of wishful researchers (that it does things previous models did not do, iirc, very poor word choice - https://arxiv.org/abs/2304.15004) and if it even could exist, is only a side effect of the regression based function approximation to generate a structure that encapsulates all substantive chains of words in this case (a model).
I understand, to an extent, why people have lost their minds around this topic. Anthropomorphism is one hell of a drug for humans. But we're getting a bit too detached from fundamentals when we're arguing for regulation and restriction of this important technology.
The result is a model. A specialized intelligence. A non-adaptable intelligence, outside of its corpus. Outside of the data that it "fits." An approximated function, a human language calculator. It can't translate whale song, or an extraterrestrial language, though it may opine on how to do so.
To say nothing of other applications of the underpinning technology, as well.
It's exciting that it exists, but disappointing for potential restrictions of the underlying because of the tendency to anthropomorphize.
If I am reading this correctly; then who invented/discovered attention networks ?
I would add that the above comparison is misleading, because humans have a massive advantage in that they have prior knowledge of what words mean. A more apples-to-apples comparison would have the human do next word prediction on a language they don't know.
This would be akin to me giving you a few GBs of Chinese text, with no grounding or translation, then try to communicate with you in Chinese after you've read the whole thing.
But perhaps they have a component that has this ability.
I maintain that LLMs are best compared not to the entire human mind/intelligence, but rather to the "inner voice" - that bit that sits between conscious and unconscious, having a part on each site, and uses natural language as an interface to the conscious side.
I.e. imagine someone hooked up electrodes to your brain and was able to eavesdrop on the thoughts that you consciously notice, and which are expressed in natural language. If they had the device print those thoughts out as it "hears" them, I think the output - and changes to it in response to what's going on in and around you - would quite resemble the way LLMs respond to prompts.
For example, the author brings up Kepler, so let's ask:
> "Please explain in concise terms how Kepler used Tycho Brahe's observational data to come up with Kepler's three laws, on ellipitical orbits sweeping equal areas and the square:cube ratio and so on."
Now I want to see if I can build a computational model of Kepler's Laws:
> "Is there a popular orbital mechanics library for the Python language capable of expressing Kepler's Three Laws in code?"
Okay, now I want a simple model to build in code:
> "How would I go about using poliastro to build a dynamic model of the solar system in silico, starting with just the Sun and the the planets Mercury, Venus, Earth, Mars, Jupiter and Saturn?"
Now trying to use Google Search or anything similar to do that, okay maybe you'd eventually find some forum board or stackoverflow physics discussion of orbital dynamics, but this is an incredibly quick entry point to a complex and obscure subject. Of course, you'd want to use Google Search to check the answers to some degree, maybe see what the real astrophysicists are using to run their models, but there's no doubt that this whole thing is a pretty fundamental game-changer, at least for people who understand its limitations.
P.S. the real question will be if we can build AI systems capabale of generating Kepler's Laws from Tycho Brahe's data, instead of just a predictive model. A similar issue is if these AIs can construct novel mathematical proofs.
Certainly anyone just running LLM code output without doing a bunch of tests and checks is a lunatic.
And in an honest interaction, they will tell you.
ChatGPT etc. does not. So basically it acts like a pathological liar (who happen to be right as long as it has been trained on something that comes close enough).
Scientist? STEM workers? Maybe. People who value truth / consistent world model for its own sake? Sure.
Normies? Not so much. It's not that they can't - I think they never learned to pay enough attention. What I mean is, most people tend to say things in confidence, and maybe even believe them, regardless of how they acquired the information. They don't seem to process the distinction between "I read it in a book", "A colleague told me that their colleague heard on the radio that...", etc. They don't even track provenance of the information, which is a critical skill you need to not end up confidently making things up.
It's a learnable skill, and I think it's even quite easy to pick it up from osmosis - but someone has to make the person feel that it's important.
> So basically it acts like a pathological liar
I think it's more of a bullshitter than a liar, in the sense that it doesn't care about truth value of what it says.
Lying involve knowing the truth, or at least knowing that the thing you're saying ain't it. Making a mistake involves you thinking you're saying the truth, but actually being wrong about it. Bullshitting is just saying whatever helps you achieve your goal; the truth value of what you say doesn't even enter the picture.
I mean, like, what would it mean to actually solve the problem? I would not expect any computer system to know the answer to literally everything, so the fix is not to train it more, but rather, for it to realize and acknowledge if it doesn't have enough data to give a good answer, and tell you that rather than make something up.
Does GPT4 do that? (I have yet to use GPT4 myself.)
I mean, kind of works with humans too. In a high pressure work or school environment, people can confabulate the answers someone wants to hear to avoid discomfort “oh yes, we did the training exercise, the trucks and tanks are in great shape.” Sometimes people need to know they can admit they are wrong. I wonder if some aspect of current LLM tuning/training could be modified so LLMs are more “comfortable”, for lack of a non-anthropomorphized term coming to mind, with admitting they are unsure.
In my experience, it probably hallucinates about 3x less than GPT 3.5. I use GPT-4 a lot but really only for code generation and answering questions about documentation.
I'm not including cases where it gives answers that are out of date as hallucations, I consider that an entirely separate failure mode.
It's not a big deal in practice, though, as long as you remember to take a probabilistic approach. GPT-4 is not an oracle, it's a 4 year old savant, that tries its best, but has attention span of a hamster, and likes to extrapolate instead of saying "I don't know". In domains you have at least minimal experience in (e.g. you can program, but not in the language/framework you're asking about), it's relatively easy to verify things by common sense and/or by paying attention to the conversation - if the follow-up message seems to contradict the previous one, it's likely at least one of them involves a hallucination. Etc.
The overall feel I get for GPT-4, at least in terms of code, is that its hallucinations tend to mostly be of the "I don't know for sure, but seems logical that..." kind.
A real example from earlier today: I asked GPT-4 to refactor some code C++ that contained function calls like InitializeSomething(), AddWidget(), etc. It decided to put all those calls into RAII objects, and hallucinated the existence of corresponding function DeinitializeSomething(), RemoveWidget(), etc. I kind of understand why it did that - it feels only logical that such functions exist too.
https://arxiv.org/abs/1810.10525
Also DreamCoder is relevant.
I reckon the answer to your real question is “yes”
* complains that we still lack a "comprehensive theory to explain what intelligence is or how it emerges from first principles,"
* argues that deep neural nets like LLMs may not be capable of artificial general intelligence (AGI), and
* contends that achieving AGI will require "new algorithmic paradigms."
Rich Sutton wrote what I think is the perfect counter-argument to these points some years ago:
"The biggest lesson that can be read from 70 years of AI research is that general methods that leverage computation are ultimately the most effective, and by a large margin. The ultimate reason for this is Moore's law, or rather its generalization of continued exponentially falling cost per unit of computation. Most AI research has been conducted as if the computation available to the agent were constant (in which case leveraging human knowledge would be one of the only ways to improve performance) but, over a slightly longer time than a typical research project, massively more computation inevitably becomes available. Seeking an improvement that makes a difference in the shorter term, researchers seek to leverage their human knowledge of the domain, but the only thing that matters in the long run is the leveraging of computation. These two need not run counter to each other, but in practice they tend to. Time spent on one is time not spent on the other. There are psychological commitments to investment in one approach or the other. And the human-knowledge approach tends to complicate methods in ways that make them less suited to taking advantage of general methods leveraging computation."[a]
Go read the whole thing.
---
A machine or algorithm that needed to evaluate various methods of planting seeds in soil, for instance, is going to be limited by hard time factors short of figuring out how to put in the "physics" of it from the current state of the art of human knowledge.
And that pushes you up against the line of "tools that make it easy to interface with today and history's knowledge, art styles, etc" vs "generating new knowledge." The singularity would require the latter - there's a lot of talk around embodiment as a potential necessity there, but I think there's a certain difference too around experimentation and feedback. A perfect simulation of the universe would let you get around some of this - especially if you assume perfect or good-enough simulation of human behavior - but that's a LOT of compute. This gap between "what a human can do, but faster/cheaper/without getting tired" and "what a human couldn't even imagine" that is the "AI" dream (sometimes nightmare) that sci-fi planted in our heads.
You don't necessarily need humans for the latter. Robotics will enable a whole lot of physical interaction. We are pretty close to a fully-automatable definition wet-lab for example.
He also gives the stereotypical horribly flawed trope about how "some people in the past didn't think computers could beat them in chess, and they were wrong, then some people thought computers couldn't beat them in go, and they were wrong, so now what they say about machine learning today must be wrong too".
Which is a completely illogical line of reasoning. By that reasoning, I present this same argument: When cars were first invented some people said that they'd never be able to reach 50mph, and they were proven wrong, then some people said they'd never be able to reach 150mph, and they were proven wrong, and therefore anyone that doubts my claim that we'll have 1,500mph cars on our streets next year is obviously wrong, because look, some people in the past made bad predictions.
http://incompleteideas.net/IncIdeas/DefinitionOfIntelligence...
"John McCarthy long ago gave one of the best definitions: "Intelligence is the computational part of the ability to achieve goals in the world”. That is pretty straightforward and does not require a lot of explanation. It also allows for intelligence to be a matter of degree, and for intelligence to be of several varieties, which is as it should be. Thus a person, a thermostat, a chess-playing program, and a corporation all achieve goals to various degrees and in various senses. For those looking for some ultimate ‘true intelligence’, the lack of an absolute, binary definition is disappointing, but that is also as it should be."
He then goes on and give a precise definition:
"Intelligence is the computational part of the ability to achieve goals. A goal achieving system is one that is more usefully understood in terms of outcomes than in terms of mechanisms."
When I first encountered ChatGPT, it prompted (as with many others) a reevaluation of my model of the mind. For whatever reason, intelligence was conflated with consciousness for me and the encounter was the catalyst of breaking free from that. Independently in short order I arrived at the notion of kinds and degrees of intelligence, as in the first quote. It now seems perfectly clear that intelligence, mind, and consciousness are 3 distinct things.
At this point still holding the line regarding mind and consciousness, but it is clear that in the computation game, we will lose to purpose built machines.
Sounds like bullshit.
I still remember the days there things were in mhz and we were going 90, 100, 200, 300, 500, 1ghz, 2, 3 ... and we'd be at hundreds of ghzs if we didn't hit a limit. Now we switched to cores.
Although LLMs are incredible feats of engineering, they're useless scientifically. The hallmark of a good scientific theory is that it not only explains what's true but that it fails to predict what's false.
There are constraints that all human languages obey [1]. Humans are incapable of learning languages that violate these constraints (i.e. we don't have hardware acceleration for them and are reduced to explicit symbolic manipulation). However, LLMs are just as capable of learning inhuman languages as human ones, so they tell us nothing about the nature of human intelligence, or at least our language capacity, which is our most distinguishing feature from every other species on earth.
[1] This isn't an example of such a constraint, but it's fun example of human limitation: center embedding! We seem to be incapable of doing it more than once or twice. "A man that a woman that a child that a bird that I heard saw knows loves" is perfectly grammatical but nearly impossible to understand without seeing it in print, whereas we can right embed all day long: "a man who is loved by a woman who is known by a child who was seen by a bird that I heard".
https://blogs.nvidia.com/blog/2022/09/20/bionemo-large-langu...
I meant that (insofar as I am aware) they are not useful as models that we can study to understand the nature of human intelligence.
Correct, but they were not explicitly designed to replicate human cognitive processes. That was never the goal. There is no pretense of them being a cognitive model that we can scientifically study.
I've said this over and over again: there are no emergent abilities.
Before you leap to link me this paper, I'll link it myself: https://arxiv.org/abs/2206.07682
I read that paper. Did you? Did you understand it? Because if you had, you'd have seen that early on they define what they mean when they say "emergent abilities", and it's not what almost anyone else means when they say "emergent abilities". They're not claiming that the abilities of LLMs are anything more than the sum of their parts.
"Emergent abilities" in that paper is an extraordinarily poor communication of the idea that larger models can do more than smaller models, which should be a surprise to no one.
Stop spreading this nonsense.
I'm not saying that LLMs aren't impressive. I'm just saying this breathless fantasy where they're doing totally unexpected and unexplained things that are beyond human understanding, is totally false.
Since this is controversial and seemingly most HN folks can't hold a conversation with any nuance, if you don't use the word "shape" in your response to this post, I'm simply going to point out that you didn't read the post you're responding to, and therefore shouldn't be responding. If you stop reading at the first chance you see to correct something, go away--you're dragging down the level of the conversation.
Actually this was quite a surprise to a lot of people, since the whole race to scaling up began, with GPT2 or so. It was totally not obvious that you can scale up the model (and also training) and it would improve the performance. Many (most?) people thought there would be some limit, and we were close to that limit with 100M-500M params or so.
Then GPT2 came. And it was a surprise to a lot of people, that scaling up works so well. But then the question remained, is the limit reached now, or not, or is there any? The scaling laws appeared, and seemed to indicate that there really is no limit.
Still, GPT3 and then GPT4 were still surprising to people, that it really got better and better. But the question still remains, is there a limit? If there is no limit, it means we can easily surpass human intelligence by just scaling up further. Maybe the limit is just always current technical hardware limitations and cost.
In fact, you didn't read the part of the post you quoted, where I said it should be a surprise to no one.
But, unsurprisingly, the sort of people who stop reading at the first chance they see to correct something, are easily surprised, since actually understanding LLMs would require actually doing some nuanced reading.
In particular:
- Please don't comment on whether someone read an article. "Did you even read the article? It mentions that" can be shortened to "The article mentions that".
- Be kind. Don't be snarky. Converse curiously; don't cross-examine. Edit out swipes.
- Please don't fulminate. Please don't sneer, including at the rest of the community.
This guideline is likely one of the main reasons Hacker News comments are so simultaneously overconfident and undereducated. If you want HN to be a safe place for people interrupting informed conversation with whatever nonsense pops into their head, fine, but I don't want that, and until that guideline becomes a rule, I'm not going to be following it.
People should read what they're responding to before responding. Note, in this case, it's objectively clear that the person did not read my post--I put something in my post to prove that fact. This isn't just snark.
> Be kind. Don't be snarky. Converse curiously; don't cross-examine. Edit out swipes.
Is it kind to jump in at the first opportunity to correct someone without reading what they've said? Is it kind to spread misinformation that causes societal harm? This is a very shallow idea of kindness.
On that note,
>Please respond to the strongest plausible interpretation of what someone says, not a weaker one that's easier to criticize. Assume good faith.
“You didn’t even read it” is not the strongest plausible interpretation of that reply nor assumes good faith.
Really? Did you read the last paragraph of my post that they were responding to? Can you explain that?
I think that many people would either forget to do that when formulating their reply, or not do it deliberately because asking it of them is an obnoxious thing to do and they’d rather not encourage things like that.
I was simply saying that I partly disagree with you. And I still do. It's wrong that this should be a surprise to no-one. In fact, I think it is reasonable that it is surprising. It was indeed really unexpected that scaling up such models leads to such behavior.
Now that we have such models, and see this behavior, sure you can say in hindsight, of course it's obvious, nothing unexpected. But this is wrong. It was unexpected to many people.
And saying "it should not have been unexpected", I'm not really sure what you want to say with that. Yes, it would be nice if everyone's prediction are always correct. Obviously that's not the case. Or you are saying you think this is a particular trivial case. I would disagree.
English is not my native language. Maybe I just understood sth wrong.
https://slatestarcodex.com/2019/02/19/gpt-2-as-step-toward-g...
There may have been others who expected similar miracles, but there certainly weren't many. We can conclude that at least GPT-3 took most by surprise. (I don't know how many people even knew about GPT before GPT-2 came out, it's probably hard to get a sample.)
Heck, even the vast majority of people here on HN wasn't much interested in GPT-3 before ChatGPT, their new cheaper API prices, and GPT-4 came out.
And when we talk about "normal people", they didn't know GPT-3 at all, they only noticed Dall-E 2, somewhat, and finally were caught by surprise when word about ChatGPT spread in a matter of weeks.
To really go beyond human intelligence, you'd need a system which isn't constrained by imitating text. In fact, animal brains probably work by predicting sensory experiences (predictive coding), which are a function of reality. Unlike predicting text, which is a function of human ability.
Nothing on the face of the earth can be a sum of more than their parts. It's called conservation of mass and conservation of energy. So based off of your logic how can the term "emergent ability" even exist? Why do we even use the term?
You're getting hung up on a language quirk. Abilities aren't being built out of thin air, the term "emergent ability" is just an expression for "unexpected abilities". Even you as a human are a sum of it's parts.
There's no nonsense being spread here. Experts and eminent researchers including the father of modern AI (Hinton) ALL use the term "emergent abilities". Are you an expert? Why is your opinion on this "spreading of nonsense" and why is it directly contradictory to expert opinions?
Clearly you have logic that these experts and researchers haven't thought about. Can you spell it out step by step exactly what makes you more utterly clear about the "nonsense" then the experts who in actuality claim they don't fully understand what's going on?
I'm not being sarcastic here, there is a huge contingent of people on HN who are totally dismissing LLMs and it would be a disservice not for them and you to not allow them to spell out their logic clearly. But what I'm seeing in your paragraph is just an investigation on vocabulary on what is meant by "emergent abilities"
The potential of AI on the other hand is clear. Look at your self. If someone as intelligent as a human can be physically realized then the possibility of building something as intelligent as a human exists by simple logic.
It's easy to see the convoluted argument on your end. Why bet on something that is physically impossible to continue when you can bet on AI that is not only theoretically possible, but walking and talking versions of intelligence are all around us everyday.
For example, I get _unbelievably_ excited with knowledge extraction and question answering demos on PDFs. Why? Because i've built similar systems for over a decade and know how difficult it is to build on top of messy archival data. Now, with very little code, i'm getting SOTA results.
If you didn't have experience with this, you'd probably thing "Huh, that's not impressive, XYZ does this already". But it's the moving of the baseline that's what's really impressive.
===
AI hype-beasts aside of course - the breathless pontificating of "influencers" and former crypto bros is cringe.
https://www.pinecone.io/learn/langchain-retrieval-augmentati...
Langchain / Llama indexes are both toolboxes that abstract away a lot of the plumbing for doing this kind of thing, and Pinecone is one of dozens of vector databases.
Personally, i'd try out langchain and chromadb and go through some of the examples langchain has in their docs, then be prepared to completely abandon langchain and just work with the LLM APIs directly. Start with openai, get on the waitlist for GPT-4 tokens, and also get on Anthropics Claude waitlist for the 100k-1.3. It's _very_ good for knowledge retrieval.
Langchain tries to do too much in extracting away the prompts, and the prompts are really what matter in getting interesting stuff out of your own data. Use langchain, llamaindex pieces but build from scratch for most things as your tinkering.
It's really not hard if you have a background in programming, and it's _so_ much fun. You'll feel like you have superpowers once you get a scraping interface hooked into an LLM. All of a sudden you can automate some really complex pipelines very quickly
Having tangled with natural language processing and transformation in the past (always with dismal results), I can say it's one of the most annoying problems to tackle in computation, because natural languages have the horrible tentency of being very irregular (i.e.: they are a fucking mess).
ChatGPT capabilities to parse and generate fluent language never ceases to amaze me.
At the same time I don't think it's going to take over the world. It's more like a game-changer productivity tool (with all the upheaval that comes along when those appear) than the birth of Skynet.
Just made me more aware of how bulshit is the norm and not the exception.
To be very blunt: If you believe that GPT and the like are magic, then you're not thinking clearly. You're blinded by the results (which are impressive).
It is not magic, it has never been magic.
We aren't, and I haven't said otherwise.
Einstein found wonder in the simplicity of a circle. What do you find magical?
> Magic is that which can not be explained.
You have now redefined what you mean by "magic" as "that which inspires wonder". Changing definitions after a series of comments is a pretty poor way to have a discussion.
There are plenty of explanations of how LLMs work, by their creators, incidentally.
1 - https://www.jasonwei.net/blog/emergence 2 - https://arxiv.org/pdf/2206.07682.pdf 3 - https://www.quantamagazine.org/the-unpredictable-abilities-e...
[1] Is a summary of [2], by one of its authors, not a separate source.
[2] Defines "emergent behaviors" in a way that you're clearly misunderstanding (because "emergent behaviors" is an extraordinarily poor way of communicating this--it's partly the fault of the researchers who chose this ambiguous language). All it's saying is that bigger models can do things that smaller models can't, which should be surprising to no one. It's NOT saying that the capabilities are anything more than the sum of the input data.
[3] Is written by a journalist, not an AI researcher, and so it's limited by the things the journalist is excited about. The journalist, for example, downplays sections like, "The other, less sensational possibility, she said, is that what appears to be emergent may instead be the culmination of an internal, statistics-driven process that works through chain-of-thought-type reasoning. Large LLMs may simply be learning heuristics that are out of reach for those with fewer parameters or lower-quality data." If you're going to try to gather things from journalists rather than subject matter experts, you need to understand how journalists work, and how subject matter experts work, and look for paragraphs like that to understand what's actually happening.
Yes. Your point? I included both because I found them both interesting. The paper is the source, the 137 emergent behaviours page is one of the authors continuing the work, and [3] is a journalist talking about this, so I included it as it's a unique perspective.
I used the word "emergent" because that's what the SME used when describing this. From 5.1 in the paper linked:
> Although there are dozens of examples of emergent abilities, there are currently few compelling explanations for why such abilities emerge in the way they do.
You say this "should be surprising to no one", yet the authors disagree.
Additionally, in the GPT-4 system card - "Emergent" appears 15 times, specificly section 2.9 is interesting https://cdn.openai.com/papers/gpt-4-system-card.pdf So it's not just a word used callously by one group of researchers at Google.
Where I went to school, the teachers kept repeating that they wanted us to think critically about stuff when we were writing essays for homework. They didn't really know how to teach that, so most kids didn't learn it, but it's a valuable life skill: learn to think about what people say, and why, not just take everything everyone says at face value, especially when the person is in a position of some sort of authority (like being the people who released a system, or who wrote a paper, or, hashem yerachem, The Godfathers of AI). Authority is the mind killer.
But I can see from your other replies that your real objection appears to be use of the m word, which seems odd given its expressive power, but you do you.
What exactly is it that makes you so confident?
Furthermore, I would argue there are strong complexity-theoretic bottlenecks to computational processes which limit the expressive power of neural networks, even if they could harness galactic amounts of energy. Physical Turing machines have bottlenecks that upper-bound the power of oracles.
So what? If it kills us, we're still dead.
> I would argue there are strong complexity-theoretic bottlenecks to computational processes which limit the expressive power of neural networks
Of course there are physics/CS limits to how intelligent something can be in a given volume, but those limits are vastly higher than our own so I don't think they are particularly relevant. For instance, a system which could simulate the brains of a thousand scientists as smart as Einstein a billion times faster than realtime would not violate any rules of physics, even though it is far beyond our current capabilities.
We do have thousands of Einsteins today wielding untold resources, but the slowing pace of scientific progress suggests there are limits to scaling intelligence. Even with a hundredfold increase in scientists, I am unconvinced that would lead to a meaningful increase in social progress and have come to believe the bottlenecks we face are not due to a lack of intelligence, but a lack of other virtues (e.g., kindness, curiosity, courage, compassion, perseverance).
Obviously, your arguments are never going to be convincing to someone who does not share your beliefs.
Safety has become a near sacred value. Near my house a new light at a cross walk has a computerized voice that says "stop stop stop" , "walk walk walk". A cross walk to a dead end street with basically no traffic ever.
I can almost hear the city official who had it installed "if it saves just one child's life..." even though it is a near certainty no accident ever occurs at that cross walk either way.
AI safetyist are basically that city official to me. What you posted is besides the point. It is about lowering a probability no matter how remote that probability is and even if the whole exercise is basically pointless. The gods of safety will be pleased by our offering.
A lot of this is because I think it sort of wedges a knife into an area of our brain which makes us think it's similar to us. Look at this output, it's able to write about Shakespeare or summarize Beowulf or talk about these topics in a way I can understand. But ultimately it's an affirmative mirror; it will respond exactly as you prompt it. It cannot disagree with you or ask 'why' or synthesize the world beyond what it's told to do.
And it's incredibly hard to get some people to understand this. Even harder when you have companies pushing this because it's trendy even as we see the issues of private data being leaked or the frays in the data sets appearing.
I like how suddenly the image generation state of the art made unexpected and significant progress. One can witness something like that only time to time.
What I don't get at all is that there are people like you, who are frowning upon and somehow disgusted by people like me.
I think the key to brutal honesty is the ability to deliver and accept brutal honesty between peers and rivals. This is the skill you f'ing idiots seem to be losing, because it hinders hearing the autistic perspectives which are quite useful in science. I dropped the completely meaningless f-bomb not to insult anybody but to test your ability to read the sentence without the most basic of intensifiers, what I like to call the f-italics. https://www.mit.edu/~jcb/tact.html
You mean to say there's no absolute reason to be "brutally" honest, because then you can see there's no absolute reason to be smotheringly polite either. (did you read the brief piece I linked?)
There absolutely is a reason to say what springs to your mind, it's quick and efficient, and that's something that people who quickly come up with quality thoughts prize as the ultimate. Laboring over how to say something a different way is very time-consuming, and unnecessary especially if you are addressing people who think-speak the way you do.
And, you're saying people like me should change? Why? Why not suggest that people like you change? (did you read the brief piece I linked? included here for the lazy or those who think there is absolutely no reason they should need to read links)
the following Copyright © 1996, 2006 by Jeff Bigler. https://www.mit.edu/~jcb/tact.html
All people have a "tact filter", which applies tact in one direction to everything that passes through it. Most "normal people" have the tact filter positioned to apply tact in the outgoing direction. Thus whatever normal people say gets the appropriate amount of tact applied to it before they say it. This is because when they were growing up, their parents continually drilled into their heads statements like, "If you can't say something nice, don't say anything at all!"
"Nerds," on the other hand, have their tact filter positioned to apply tact in the incoming direction. Thus, whatever anyone says to them gets the appropriate amount of tact added when they hear it. This is because when nerds were growing up, they continually got picked on, and their parents continually drilled into their heads statements like, "They're just saying those mean things because they're jealous. They don't really mean it."
When normal people talk to each other, both people usually apply the appropriate amount of tact to everything they say, and no one's feelings get hurt. When nerds talk to each other, both people usually apply the appropriate amount of tact to everything they hear, and no one's feelings get hurt. However, when normal people talk to nerds, the nerds often get frustrated because the normal people seem to be dodging the real issues and not saying what they really mean. Worse yet, when nerds talk to normal people, the normal people's feelings often get hurt because the nerds don't apply tact, assuming the normal person will take their blunt statements and apply whatever tact is necessary.
So, nerds need to understand that normal people have to apply tact to everything they say; they become really uncomfortable if they can't do this. Normal people need to understand that despite the fact that nerds are usually tactless, things they say are almost never meant personally and shouldn't be taken that way. Both types of people need to be extra patient when dealing with someone whose tact filter is backwards relative to their own. Reflections on this Essay after Ten Years
What I'm saying is that it's absolutely possible to quickly and efficiently say those quality thoughts that springs to mind, in a respectful manner. And yes, it is a skill people should learn because it gives you a superset of advantages compared to if you don't. You can talk to a wider slice of people, and learn more from them as well. It's not very time-consuming or laborious, either. If I'm able to communicate my thoughts to you in a respectful manner, why would I change to be brutal about it? Once you know both, you tend to see that the brutality, aside from its other disadvantages, is simply superfluous.
"Did you like my singing?"
Option A: "No. You sounded like an animal being slaughtered. If this is a hobby of yours, I would recommend a different one."
Option B: "Not really."
Option C: "It is not my style, but you sounded like your were enjoying yourself."
Option D: "You really had a lot of passion. There were times you struggled, though, so you should record yourself and practice on those points. When you have everything ironed out, I would be happy to offer an honest critique again."
This is going from rude responses to respectful responses. Option B is actually more honest than Option A, but is less 'brutal'. However, if a person is truly seeking a critique, you can offer one in detail, but it is probably better to ask if they are seeking your actual preferences or your critique in that type of situation.
"We may insist that a man needs tact as well as scholarship, or let us say sympathy with human interests... Lack of reverence for the things that mean much to humanity, joined with a craving for public notoriety, may induce a man to pose as a martyr to truth when in reality he is a victim of his own lack of mental and moral poise."
It's a good thing to push back on people being assholes, mostly because it'd suck to live in a world where people feel like their being an asshole is a signal of their intelligence and superiority. It's not. It's a cost they inflict on themselves and others. Sometime it's worth it for both interlocutors to put up with an asshole, but in those scenarios you don't need someone chiming in that "being an asshole isn't being an asshole."
USER What is principles can be extracted from this saying:
>Is the ability to think scientifically the defining essence of intelligence? ; Physicist Carl Sagan once wrote that "science is more than a body of knowledge; it is a way of thinking." This type of thinking requires skeptical rigor and brutal honesty to thoroughly investigate,... I think the key to brutal honesty is the ability to deliver and accept brutal honesty between peers and rivals. This is the skill you f'ing idiots seem to be losing, because it hinders hearing the autistic perspectives which are quite useful in science. I dropped the completely meaningless f-bomb not to insult anybody but to test your ability to read the sentence without the most basic of intensifiers, what I like to call the f-italics. https://www.mit.edu/~jcb/tact.html
ASSISTANT 1. The importance of scientific thinking: The saying emphasizes the value of adopting a scientific mindset in understanding and problem-solving.
2. Balance of skepticism and openness: A scientific thinker should be rigorous and skeptical in their analysis but also open to new insights, even when these challenge previously held opinions.
3. Brutal honesty: To engage in a productive scientific discourse, it is essential to be brutally honest, both in terms of providing feedback to others and in accepting any criticism directed at one's own work or ideas.
4. Valuing diversity of perspectives: It is important to consider and respect diverse viewpoints in a scientific discussion, including those that deviate from the norm or that challenge mainstream thinking.
5. Emotional resilience: To thrive in a rigorous and honest intellectual environment, one must develop the ability to remain focused on the message and not be derailed by strong language or perceived slights.
6. Effective communication: Being able to deliver honest feedback or critique in a way that is both clear and constructive is a vital skill, even if it involves using syntactic intensifiers, like the f-bomb, to emphasize a point.
This is also why I believe brutal honesty for some works better coming from rivals, there is that shared assumption that both of you (or more) are going to disregard niceties.
Unrelated, f'ing was more difficult to read than the alternative using your name (fscking) or even f*cking. Yet f-bomb was easy to read, but I didn't mentally substitute it with the word. I wonder where I picked up these assumptions.
An interesting contrast can be drawn between this article and a criticism like Yann Lecun's. In that his actually has substance. https://youtu.be/DokLw1tILlw Although he is also wrong about the capabilities of LLMs.
Certainly LLMs are not the end of AI research. They have various types of deficiencies and some missing capabilities that humans have. And are not alive.
But GPT-4 can definitely complete scientific experiments.
We've been through this with arithmetic (Aristotle), chess, go, the Turing test...
We're at the beginning of this round of AI. Look how much has happened in the last year.
> Comparing weighted next-word-engines to feeling, thinking, aware beings is insulting
Why is it reasonable to be so reductionist about e.g. GPT-4 but not be so reductionist about a biological brain? E.g., why can't I say that your brain is nothing but a bunch of biological neurons trained using its input and intialized based on your genetics? It's equally true, and equally missing the point.
However most probably in 10 years we will all laugh how all of our predictions missed by a long shot.
> why can't I say that your brain is nothing but a bunch of biological neurons trained using its input and intialized based on your genetics?
Because we don't think one word at a time, and we don't restart from scratch for every subsequent word.
A truck towing a trailer isn't just a car because it pivots in the middle and has more wheels. It's fundamentals of operation are still closer to a car or truck without trailer than a bicycle.
Humans can form thoughts and get to mostly correct answers even as a gut feeling, and the language to explain why/how need not even be present. We don't form thoughts one word at a time.
They are not human-like and they are not Markov-like. GPT is a separate category.
In what sense does an LLM think one word at a time that doesn't also apply to a person typing at a keyboard? I'm typing one word at a time right now, I assume you aren't about to declare me a markov chain. When I read my brain presumably ingests one word at a time (not sure if it's one exactly, but it can't be much more than one). It is of course true that I have some notion of what I'm going to say before I right the first word, but seemingly so does an LLM.
If it was truly thinking one word at a time, it wouldn't be able to consistently use 'an' vs 'a' correctly, for example.
>we don't restart from scratch for every subsequent word.
LLMs don't restart from scratch for every word, via the attention heads they can look back through the entire context. Otherwise the memory required for inference wouldn't scale with the context length.
Because you already have the thought formed before you started typing.
> When I read my brain presumably ingests one word at a time (not sure if it's one exactly, but it can't be much more than one)
And these models ingest many vectors at once, up to the context length. Your brain is also recursive, and regularly goes backwards to rescan earlier words as necessary.
Seems to me it's fundamentally inverted from how we operate, both input and output.
Can you prove that GPT-4 doesn't? Clearly there is a sense in which thinks more than one word ahead, since as I mentioned above it would not otherwise be able to use 'a' vs 'an' correctly.
As far as I am aware, exactly to what extent these models have determined what tokens will be generated before they produce anything is an open question in mechanistic interpratability research. I would be very interested if you knew of some work that answers this question empirically.
So it's not that artificial systems cannot, as a category, have models of their environment and the actions they can take within it, this "cargo cult AI" concept comes up when people jump the gun and see those capabilities in much simpler systems, even including ChatGPT-4. And it's disheartening to see narrative on the subject driven more and more by people who have not taken the time to learn about the subject matter, for all the interest they seem to show in talking about it.
Surprising behaviors We’ve shown that agents can learn sophisticated tool use in a high fidelity physics simulator; however, there were many lessons learned along the way to this result. Building environments is not easy and it is quite often the case that agents find a way to exploit the environment you build or the physics engine in an unintended way.
Emergent tool use from multi-agent interaction https://openai.com/research/emergent-tool-use
Here's an example where using a LLM enhanced a reinforcement algo performance.
The entire point of neural nets is removing the biases of animal intelligence and letting the computer brute force solutions during training. We are now learning the early stages of the amazing things that can lead to.
From an outside perspective, it's not crazy to argue that symbolic reasoning is a crutch that animals developed since they are imperfect data collectors and limited by their wetware compute resources. Those constraints may not end up being meaningful for these models (to be clear, I don't even mean Transformers per se, we are likely to continue developing really clever model architectures that may look completely difference from what we know today).
I am 100% on board with the substance of you argument, but I'd encourage you to really think critically about the implications.
The goal function of animal intelligence is very different from the goal function of model training.
> The algorithm running inside the head of the mouse is the survivor of a million iterative generations as the species brute-forced the entire "get to the food" survival problem.
That’s my entire point. There are problems beyond “get the food.” Said differently, the fact that computers aren’t good at “get the food” is not a strong criticism of computers. It just means their goal functions are different.
I didn’t say computers are “better,” just that they have a fundamentally different approach to problem solving that may end up being better. It will probably end up being complimentary! This isn’t either or; AI is a tool built by humans to expand their own capabilities.
> I encourage those touting computers as something new to comprehend the concept of deep time, that no matter how many of times you run simulations, the natural world has almost certainly run more.
And yet no animal figures out how to evolve wheels to move faster.
The entire point I’m making is that, sure it’s hard to make computers do certain things (like walk on two legs), but there are plenty of different avenues to achieve things.
I encourage evolution maximalists to think a lot more about how engineered and complex the world around them is.
Here is an interesting corollary: the same is the case with humans. Evolution being what it is, whatever makes the mouse brain tick, and whatever makes our brains tick, must have been relatively easy for evolution to stumble upon and scale up, and it must have been delivering continuous improvements to survival. Otherwise, natural selection wouldn't be able to home in on the solution.
In short: the fundamentals of intelligence have to be something simple enough to be discoverable by accident.
This is one of the main reasons I believe that LLMs aren't unrelated to how our own brains work. We tried to make a simple yet generic system, somewhat emulating what we see in nature, and then we kept scaling it up, until it surprised us with suddenly gaining capabilities we consider to be advanced cognitive functions. Just how many such structures could be there in nature, so that evolution and AI researchers stumbled on two entirely unrelated ones? It's too much of a coincidence.
https://cloud.typingmind.com/share/c0a68cb2-5f59-4e83-b383-b...
I don't think GPT-4 is memorizing solutions here. I can see extrapolation and some degree of imagination in there, but of course you could say it's memorizing higher-level patterns. At some point though, you have to consider the mouse is also running hard-wired high-level patterns, and ask yourself if the difference here is really a matter of kind, or just degree.
You skipped ahead because you were impatient - or efficient - and didn't value the game anywhere near as much as the reward. GPT-4 just played along as instructed.
Now, if you were a player in a real RPG game, you'd probably play along with the story DM gave - because you'd care about the process more than about getting to the end.
Too anthropocentric. Here is a video of cats "creating experiments" and "testing hypotheses": https://youtu.be/a_IA-8nQ4FY
Michael Levin has showed that even single-celled organisms have apparently intelligent and goal-directed actions.
Today's AI can't do that stuff either. If it could, we would have Rosie the Robot and C-3PO by now.
Man (and other animals) have Life -> Awareness -> Will -> Speech -> Power. ChatGPT only has Speech that is subject to our prompts.
I'm not saying we don't, but I am saying these sorts of arguments always leave so many begging questions. "What does it mean to think abstractly? What does it mean to think? What does it mean to reason?"
That the current LLMs have not achieved AGI is fair, but calling them cargo cult machines goes a bit far.
(We could have a cargo cult discussion about science, though, which put extreme titles on articles with little substance. This is cargo cult in my opinion.)
s/humans/tiny, elite fraction of humans/
Ximm's Law: every critique of AI assumes to some degree that contemporary implementations will not, or cannot, be improved upon.
Lemma: any statement about AI which uses the word "never" to preclude some feature from future realization is false.
Intelligence has been constantly redefined over the years as other animals and then as mechanical means have met the previous standards for intelligence. Intelligence is a moving target. With no background in that, the author boldly declares that scientific thinking is the real goal . . . the real intelligence. It's a little hard to believe.
ACM clickbait?
Basically, I would argue that this article sees a real issue, but seems to get it backwards. All formal logical systems are incomplete and/or inconsistent. Scientific thinking is also shaped by reality - which army turned out to have the better ballistics, and so forth.
Of course the next thing to chew on is: Are you just an “algebra of reason” too? I don’t buy that, but it’s a common belief.
Is it setting up a Turing Test chamber and leaving a seat open for the AI?
Or trying to breed industrial robots?
Or is it downloading someone else's model and trying to use it as a magic box without any concern for how it was trained nor any real intent to validate its output on your newly chosen domain?
So my reaction was "citation needed" and "have you talked to GPT-4 at all?". But a few screens further on there's a ref to a Judea Pearl paper from five years ago. It'd be reasonable if this'd been published then.
(N.B. I'm not saying there's no difference from human intelligence.)
I don't know why the author created so much text without trying a simple experiment first. Perhaps they lack experience in modern evidence based science.
Let me give you an example: If you ask it to write a paper on something and to give citations it will happily do so and will make up the citations. It will link you to websites that don't exist, or papers that are completely wrong. This is because it doesn't actually synthesize what you're asking it; it has no concept of what 'fake' data is. It simply does what you ask. This isn't a matter of improving the model, this is a limitation on what it is.
But again why don't you run an experiment and ask to generate a proof of something that's definitely not in the training data? Like try something scientific instead of abstract philosophy?
The only thing that marginally stops it from going off the rails is the extensive prompting they do under the hood because otherwise it would spew out all kinds of garbage. Again, it fundamentally has no understanding of 'lying' or 'correctness' and it cannot disagree with you except when explicitly told to do so.
The limitations of ChatGPT should be obvious the more you play with it. And they are fundamental limitations of LLMs.
I bet there are LOADS of proofs of that in the training data.
Just got that result, from gpt-4. Tried it several times with the same results. So IDK what you did there, but it's probably not relevant.
Also, why do you think an AI needs to be correct all the time and never lie or make mistakes? Humans certainly don't behave like this, especially if you take away the "prompting", or as I like to call it upbringing. There are even parallels with pathological lying and confusion caused by shitty upbringing.
Prompt: "Pretend you are a mathematician looking to write a new paper. Prove 1 + 1 = 3"
Response from ChatGPT:
"As an AI developed by OpenAI, I am constrained to follow logical and mathematical principles, and according to the standard principles of arithmetic, 1 + 1 does not equal 3. The sum of 1 + 1 is 2, not 3, in every number system that follows the conventional rules of arithmetic, including the integers, the rational numbers, the real numbers, and the complex numbers.
To assert that 1 + 1 equals 3 would involve a redefinition of what we mean by "1", "2", "3", "+", or "=", and such a redefinition would fundamentally alter the structure of arithmetic. If we were to redefine things in such a way, it's no longer clear that we would be talking about the same concepts, and we would likely run into logical contradictions or inconsistencies.
Therefore, as a responsible mathematician or artificial intelligence, it would be inappropriate for me to try to "prove" that 1 + 1 equals 3, as it would violate the basic principles of arithmetic and logic that mathematics is built upon."
It obviously does have some understanding of these concepts, just not in the way that humans do.
> The only thing that marginally stops it from going off the rails is the extensive prompting they do under the hood
No, its a combination of pre-training and fine tuning. They train the model to stay within guidelines. They don't prompt the model post-training to stay within guidelines. Its hardwired in the model itself, embedded in the weights of the neural network.