AI tends to have superhuman pattern matching abilities with enough data
> I realized that the AI was using the smudges on the camera to help make an educated guess here.
In short, it’s still anthropomorphism and apophenia locked in a feedback loop.
I also agree with the cousin comment that (paraphrased) “reasoning is the wrong question, we should be asking about how it adapts to novelty.” But most cybernetic systems meet that bar.
Consider your typical country music enjoyer. Their fondness of the art, as it were, is far more a function of cultural coding during their formative years than a deliberate personal choice to savor the melodic twangs of a corncob banjo. The same goes for people who like classic rock, rap, etc. The people who `hate' country are likewise far more likely to do so out of oppositional cultural contempt, same as people who hate rap or those in the not so distant past who couldn't stand rock & roll.
This of course fails to account for higher-agency individuals who have developed their musical tastes, but that's a relatively small subset of the population at large.
Nope. It's not autoregressive training on examples of human inner monologue. It's reinforcement learning on the results of generated chains of thoughts.
No, that's not how LLMs work.
It's less about the definition of "reasoning" and more about what's interesting.
Maybe I'm wrong here ... but a chess bot that wins via a 100% game solution stored in exabytes of precomputed data might have an interesting internal design (at least the precomputing part), but playing against it wouldn't keep on being an interesting experience for most people because it always wins optimally and there's no real-time reasoning going on (that is, unless you're interested in the experience of playing against a perfect player). But for most people just interested in playing chess, I suspect it would get old quickly.
Now ... if someone followed up with a tool that could explain insightfully why any given move (or series) the bot played is the best, or showed when two or more moves are equally optimal and why, that would be really interesting.
I happen to do some geolocating from static images from time to time and at least most of the images provided as examples contain a lot of clues- enough that i think a semi experienced person could figure out the location although - in fairness- in a few hours not few minutes.
Second, the similar approaches were tried using CNNs and it worked (somewhat)[1].
[1]: https://huggingface.co/geolocal/StreetCLIP
EDIT: I am not talking about geoguesser - i am talking about geolocating an image with everything available (e.g. google…)
>I have repeatedly said that "can LLM reason?" was the wrong question to ask. Instead the right question is, "can they adapt to novelty?".
I have a simple question: Is text a sufficient medium to render a conclusion of reasoning? It can't be sufficient for humans and insufficient for computers - such a position is indefensible.
Do you suppose we can deduce reasoning through the medium of text?
This sort of claim always just reminds me of Lucky's monologue in Waiting for Godot.
As far as goalpost-moving goes, it's wild to me that nobody is talking about the turing test these days.
But worse, the Turing Test is not remotely intended to be an "analogy for what LLMs are doing inside" so your comparison makes no sense whatsoever, and completely fails to address the actual point--which is that, for ages the Turing Test was held out as the criterion for determining whether a system was "thinking", but that has been abandoned in the face of LLMs, which have near perfect language models and are able to closely model modes of human interaction regardless of whether they are "thinking" (and they aren't, so the TT is clearly an inadequate test, which some argued for decades before LLMs became a reality).
To be specific, in a curious quirk of fate, LLMs seem to be proving right much of what Chomsky was saying about language.
E.g. in 1996 he described the Turing test as "although highly influential, it seems to me not only foreign to the sciences but also close to senseless".
(Curious in that VC backed businesses are experimentally verifying the views of a prominent anti-capitalist socialist.)
As far as I can see all of this [he's speaking about the Loebner Prize and
the Turing test in general] is entirely pointless. It's like asking how we
can determine empirically whether an aeroplane can fly the answer being if
it can fool someone into thinking that it's an eagle under some conditions.
https://youtu.be/0hzCOsQJ8Sc?si=MUXpmIwAzcla9lvK&t=2052The analogy I used in another thread is a third grader who finds a high school algebra book. She can read the book easily, but without access to teachers or background material that she can engage with -- consciously, literately, and interactively, unlike the Chinese Room operator -- she will not be able to answer the exercises in the book correctly, the way an LLM can.
To be honest I am still not entirely convinced that current LLMs pass the turing test consistently, at least not with any reasonably skeptical tester
"Reasonably Skeptical Tester" is a bit of goalpost shifting, but... Let's be real here.
Most of these LLMs have way too much of a "customer service voice", it's not very conversational and I think it is fairly easy to identify, especially if you suspect they are an LLM and start to probe their behavior
Frankly, if the bar for passing the Turing Test is "it must fool some number of low intelligence gullible people" then we've had AI for decades, since people have been falling for scammy porno bots for a long time
And the "customer service voice" you see is one that is intentionally programmed in by the vendors via baseline rules. They can be programmed differently--or overridden by appropriate prompts--to have a very different tone.
LLMs trained on trillions of human-generated text fragments available from the internet have shown that the TT is simply not an adequate test for identifying whether a machine is "thinking"--which was Turing's original intent in his 1950 paper "Computing Machinery and Intelligence" in which he introduced the test (which he called "the imitation game").
Try to rapidly change the conversation to a wildly different subject
Humans will resist this, or say some final "closing comments"
Even the absolute best LLMs will happily go wherever they are led, without commenting remotely on topic shifts
Try it out
Edit: This isn't even a terribly contrived example by the way. It is an example of how some people with ADHD navigate normal conversations sometimes
https://aistudio.google.com/app/prompts/1dxV3NoYHo6Mv36uPRjk...
It was doing so well until the last question :rip: but it's normal that you can jailbreak a user prompt with another user prompt, I think with system prompts it would be a lot harder
Well, in this case humans has to be trained as well but now there are humans pretty good at detecting LLM slobs as well. (I'm half-joking and half-serious)
UCSD: Large Language Models Pass the Turing Test https://news.ycombinator.com/item?id=43555248
From just a month ago.
LLMs really made it clear that it's not so clear cut. And so the relevance of the test fell.
Realizing problems with previous hypotheses about what might make a good test, is not the same thing as choosing a standard and then revising it when it's met.
How is that moving the goalposts? Where did you see them set before, and where did your critics agree to that?
It did a web lookup.
It is not comparing humans and o3 with equal resources.
It used search in 2 of 5 rounds, and it already knew the correct road in one of those rounds (just look at the search terms it used).
If you read the chain of thought output, you cannot dismiss their capability that easily.
You note yourself that it was meaningful in another round.
> Also, the web search was only meaningful in the Austria round. It did use it in the Ireland round too, but as you can see by the search terms it used, it already knew the road solely from image recognition.
That's why I'm saying it's unfair to just claim it's doing a web lookup. No, it's way more capable than that.
>That’s not Earth at all—this is the floor of Jezero Crater on Mars, the dusty plain and low ridge captured by NASA’s Perseverance rover (the Mastcam-Z color cameras give away the muted tan-pink sky and the uniform basaltic rubble strewn across the regolith).
https://nssdc.gsfc.nasa.gov/planetary/mars/mars_exploration_...
That's still metadata