This is like lamenting that a person who has a doctoral degree, say, in mathematics or physics often don't have a more than basic knowledge about, for example, medicine or pharmacy.
This is like lamenting that a person who has a doctoral degree, say, in mathematics or physics often don't have a more than basic knowledge about, for example, medicine or pharmacy.
It was word problems not rocket science. That tells a lot about human intelligence. We're much less smart than we imagine, and most of our intelligence is based on book learning, not original discovery. Causal reasoning is based on learning and checking exceptions to rules. Truly novel ideation is actually rare.
We spent years implementing transformers in a naive way until someone figured out you can do it with much less memory (FlashAttention). That was such a face palm, it was a trivial idea thousands of PhDs missed. And the code is just 3 for loops, with a multiplication, a sum and an exponential. An algorithm that fits on a napkin in its abstract form.
Why not take away from this that "intelligence" is a word that obtains something relative to a particular society, namely, one which values some kind of behavior and speech over others. "Intelligence" is something important to society, its the individual who negotiates (or not) the way they think and learn with what this particular signifier connects with at a given place and time.
Like I assume you don't agree, but just perhaps if we use our "intelligence" here we could maybe come to some different conclusions here! Everyone is just dying to be like mid-20th century behaviorist now, I just don't understand!
I prefer to use the concept of search instead, it is better defined in search space and goal space. It doesn't hide the environment, the external part of intelligence, or the learning process.
All invented/discovered and formalized by humans. That we found so much (unexpected) power in such simple abstractions is not a failure but a testament to the absolute ingenuity of human pursuit of knowledge.
The mistake is we’re over-estimating isolated discoveries and underestimating their second order effects.
I think it is more like searching and stumbling onto some great idea than pure-brain-ingenuity. That is why searching and social collaboration is essential and why I say we're not that smart individually, but we search together. It's slow, it took us years to get to Flash version of attention, but we get there, someone finds their way onto a major discovery eventually.
It took humanity 200K years to accumulate our current level of understanding, and if we lost it, it would take us another 200k years. Not even a whole human generation is that smart. It's also why I don't fault LLMs for mass-learning from human text. We do the same thing, 99% is inherited knowledge. The whole process of knowledge discovery moves slowly, and over large populations.
Maybe we get an llm to design the next evolution and benchmark how potential improvements fare?
Does this mean that these two sentences are completely unrelated and only happen to exist within the same paragraph?
> Many smart humans fail at critical thinking. I've seen people with masters [...]
I've understood you to use "people with masters" as an example of "smart humans".
You wrote:
> I've seen people with masters fail at spotting hallucinations in elementary level word problems.
I wanted to express that having a master in some (even complicated) subject does not make you a master at [pun intended] spotting hallucinations. To give evidence for this statement, I gave a different, more down-to-earth example of a similar situation.
Q: A farmer has 72 chickens. He sells 15 chickens at the market and buys 8 new chicks. Later that week, a fox sneaks into the coop and eats 6 chickens. How many chickens could the farmer sell at the market tomorrow?
AI Answer: The farmer started with 72 chickens. After selling 15, he had 57 chickens left. Then he bought 8 new chicks, bringing the total to 65. Finally, the fox ate 6 chickens, so we subtract 6 from 65. This gives us 59 chickens. Therefore, the farmer now has 59 chickens that he could sell at the market tomorrow.
--
You'd expect someone who can read/understand proofs to be able to spot a a flow in the logic that it takes longer than 1 week for chicks to turn into chickens.
Rather, I'd assume that someone who is capable of spotting the flow in the logic has a decent knowledge of the English language (in this case referring to the difference in meaning between "chick" and "chicken").
Many people who are good mathematicians (i.e. capable of "reading/understanding proofs" as you expressed it) are not native English speakers or have a great L2 level of English.
If an AI made a similar mistake, people would laugh at it.
You confuse "intelligence" with "knowledge". To keep to your example: there exist quite a lot of highly intelligent people on earth who don't or barely know English.
Imo it's evidence that humans make assumptions and aren't always thorough more than evidence of smart people being unable to perform elementary logic.
The OP there also has a pretty bad riddle (due to a grammatical error that completely changes the meaning and makes the intended solution nonsensical, and a solution that many people wouldn’t even have heard of).
https://chatgpt.com/share/66f9371a-d6a0-8003-b2b5-4af3b10e8a...
I think maybe the original poster is making some sort of additional assumption that the farmer must be selling chickens as meat at the market and a chick wouldn't be sold for that purpose until it's a mature chicken?
(Of course depending on how you interpret the question a chick is a chicken (species) and there's nothing inherently preventing reselling the chicks so I don't really understand why OP thinks the ai answer is clearly objectively wrong. It seems more like a matter of interpretation.)
Anyways this thread is a perfect example of the chaotic datasets that are being used to train FMs. These arguments of whether it’s reasonable to assume a chick could mature into a chicken within a week are happening everyday and have been taking place for years. Safe to say a billion dollars has been spent on datasets to train FMs where everybody has a different interpretation and the datasets are not aligned.
https://chatgpt.com/share/66f890b2-04bc-8002-9724-2deaf3985d...
"Most right" would have been to ask questions about what is being asked instead of trying to answer an incomplete question. But rarely is the human even willing to do that as it is bizarrely seen as a show of weakness or something. An LLM is only as good as its training data, unfortunately.
Regardless I think it's good showing that models are increasingly able to solve these "gotcha" questions, even though I think it's not hugely useful. Partly because I think it's a poor compliant and an easy shutdown.
At some point, we decided that compilers were good enough to convert code into assembly to just use them. even if an absolute master could write better assembly than the complier, we moved over to using compilers because of the advantages offered.
Is ChatGPT an all knowing and infallible oracle? Clearly not. But holding it to a higher standard than we hold other humans to is a unfair test of its abilities.
If you'd said "hens" you'd have a stronger point, but then you'd need to be talking about chicks and hens (and they could still cross whatever adulthood threshold you like within the week, as you didn't specify how young they are - "new" could just mean new to the farmer).
The AI got it right.
We leaned on spoken tradition education to pass down knowledge as written literacy, paper, writing tools were hard to come by until the last century. It was never about the student but the future. Still the same today; one student isn’t propping up reality.
People think learning a linguistic style means discovery of net new knowledge.
If you have a MLB career at all, you are an elite baseball player.