How much of AI's recent success is due to the Forer Effect?
shkspr.mobi
shkspr.mobi
However, what happened is that since the distance of your bullets is small, you would fly up towards them and they would 'trick' you by at first flying away then suddenly b-lining at you. and when you fired sometimes they would fire at the same time and get you.
The comments I got from people who played the game 'wow the AI is really good'. Not sure if this is a direct example of the Forer effect
hyuk hyuk hyuk
If we didn’t move the goalposts, we’d have declared Stockfish to be full AI, despite it only being a chess-playing program, long ago.
Push out a plugin that sets up a virtual board GPT-4 can read after each move and see if its any better.
My point is that the hyperbole about “we’ve cracked AI but they changed the goalposts” is self evidently not true. You just proved it right there: I need to add more plugins because ChatGPT does not understand what it’s doing.
It’s a potentially useful tool but that doesn’t make it intelligent.
Being able to properly use tools is a sign of intelligence in of itself.
Where is this threshold? I don’t think anyone really knows yet. That’s why I don’t see a problem with “moving the goalposts,” because doing so is the best way to help us truly understand what it means to be an intelligent life form (artificial or otherwise).
And again "reasonable" is a pretty useless metric. Asking a 'reasonable' person about any system that requires expert knowledge to understand is going to derive an unreasonable answer. This is because they'll conflate intelligence with human behavior.
This is just silly. Also, a machine killing me is still not enough. We have drone targeting systems in development right now but I don’t think you’d call them AI.
> And again "reasonable" is a pretty useless metric.
When you can’t define intelligence, it’s really the starting point.
The thing with LLMs is that they look almost there but, as many are pointing out, the method by which they make inferences is an analysis of how words fit together without understanding the meaning. This is why things like ChatGPT confidently spout absolute nonsense about topics they weren’t trained on (with much human intervention): it doesn’t know what it’s saying so it doesn’t realise it’s making stuff up.
I'm convinced that when the dust has settled and historians look back to decide on THE point in time at which we achieved AI or even AGI, that time will not be in the future, but in the past.
No, we didn’t. There was plenty of disagreement over it. Heck, the Chinese Room is very popular rejection fundamentally of the premise of it.
That aside, even if in a blind scenario (where the builders didn’t know the criteria used to test) being able to fool humans in linguistic interaction would be reasonably likely to be a good test of general intelligence, LLM’s are about as an obvious of a direct and deliberate attempt to Goodhart’s Law the Turing Test as one could imagine.
Using a particular capacity to test a more general capacity is obviously vulnerable to systems built to specialize in the tested capacity, as opposed to those that have it as a consequence of general ability.
If historians are looking back to decide the point, by definition doesn't it have to be in the past? Historians looking back to the future doesn't make sense...
https://arxiv.org/abs/2304.13187
Generally, chatbots function in direct communication with humans and do not function autonomously. Thus they are dependent on not just human meaning-making but human thirst for meaning which will find it even when it isn't there.
The latter doesn't follow. They may just be dependent on good prompt design.
People who aren't used to the AIs, and don't know how to use them, get worse results. This shouldn't be a surprise; it's how every tool works.
That’s why people’s first impressions are often so wildly off. People who have the fortune of asking the right prompts for their first conversations walk away believing they’ve spoken to AGI, people who ask a less than optimal prompt (or things it is bad at like math problems) walk away not seeing what the hype is about.
Same thing with copilot. The quality of the code that it suggests depend a lot on the context, so if someone isn’t writing good comments and giving their variables proper names the suggested code will be terrible.
These systems are very capable, but using them is a skill that must be learned, and there’s little material out there to teach that skill because things move so quickly.
But. Every single test I ran lead to functionally wrong designs from smallish memory errors in C (that hilariously ChatGPT was able to correct ND explain when pointed to) over misplaced/hallucinated methods in python APIs to completely hallucinated perl packages. It never gave me a useful answer.
I don't know if that tells us something about my work vs. your work or my approach to the model vs yours. But it definitely tells me that such a model cannot replace a developer.
Even Copilot, which is a smaller model, generates working code most of the time.
Certainly not an average developer (however that's decided in C's culture). Not making any statements on skill of anyone here, but I'd bet it could replace the very low ends of skill as it is (GPT-4 and maybe Bard). I doubt there's too many of those kinds of developers though.
It's not a guaranteed positive for all languages/people, but it's farther along being useful as an enhancement technology than it is as replacement technology. I think it's making progress on both axes, but one is clearly ahead of the other and it's still possible it stops progress until other innovations are made.
I was legitimately afraid for my job when GPT-4 was released. After using it, however, I sleep easy now.
Things like ChatGPT might get 90% of the way there very, very quickly for many use cases (similar to how lane keeping and adaptive cruise control is 90% of “self-driving”) but the remaining 10% of edge cases will require humans for a very long time (similar to how we never quite managed to get “100% self driving” because the vehicles still can’t respond reliably 100% of the time). But we still value it because the 90% works.
Once people understand that 10% limitation and once the 90% value has been gained/normalized/integrated, the hype will die down and we’ll be watching a slow grind to figure out how to “fix” the last 10% which is necessary to live up to what’s being hyped today.
Just wait until GPT-5, though. It will blow your mind... \s
I was using it to try and help me create Arma3 scenarios, which often involves using an ugly little DSL called SQF.
The last thing I asked it about was how to play a custom audio file sitting on my disk. It gave me a simple one-liner, that didn’t do anything. I asked it several questions to clarify, but it seemed confident that all I needed was this one liner to play an audio file sitting anywhere on my disk.
Eventually I said fuck it, and decided to google it myself, and it turns out it’s much more involved than just a single function call with a path as an argument, as suggested by ChatGPT. In fact you have to create an entirely separate config file with classes to represent the audio, and then you can reference that class to play the audio, not just a simple string path.
I’ve had it fail with a lot of other things, mostly trying to make computer controlled units perform deterministic, scripted actions, but it’s hard to tell if GPT-4 or the game’s engine was at fault here, because in my experience making Arma’s AI units do anything reliably is near impossible, no matter how much scripting you do.
Only the next day when I still couldn't wrap my head around why a real time processing system like be designed like this and did more google searching did I find it explicitly spelled out in the docs that a batch read counts as only a single read operation and can retrieve up to 10,000 records.
It's the first time I've been bit by this and it will definitely make me wary of any information it provides me in the future for technologies I am not intimately familiar with (at which point it's unlikely I'd need to consult it in the first place).
Technology is supposed to make us smarter. Blindly believing an AI that we know can hallucinate makes us dumb with confidence.
This keeps being my argument when people at work daydream about time and cost savings by offloading non-critical business functions to AI. I say, "Great, so it can produce 1000x more work than a person. But then what army of people are we planning to use to check those outputs?"
I'm super-impressed with the current crop of language models for their ability to so accurately simulate correctness, but their inability to understand what they don't know - because, in fact, they don't 'know' any of it in the sense that we do - makes them like very productive but completely untrustworthy employees. A junior dev who monopolizes his mentor's time through inconsistent performance is not a good hire.
Have you ever had a dumb/wrong thought in your head? I'm going to go ahead and answer yes for you, you do all the time. In fact you don't (hopefully) verbalize a stream of consciousness to other people around you. In general you think of something then reflect on what it is true/false.
This is not what LLMs do, they pitch back the first 'thought' they have, "correct" or not. This is why things like COT/TOT greatly increase the accuracy of LLM output. The problem? It requires at least an order of magnitude more processing to get an answer, and with GPU time already in high demand and expensive you don't see much of it happen.
Betting on LLMs commonly being wrong is not a safe bet at this point.
It's like reviewing an overconfident junior developer's code except you can't learn their particular weaknesses. If a developer is bad about memory leaks, you know to check their every PR for memory leaks. An LLM won't necessarily produce the same types of errors given similar prompts or even the same prompt with some period of time between invocations.
https://arxiv.org/pdf/2305.08291.pdf
https://github.com/jieyilong/tree-of-thought-puzzle-solver
In this paper, we introduce the Tree-of-Thought (ToT) framework, a novel approach aimed at improving the problem-solving capabilities of auto-regressive large language models (LLMs). The ToT technique is inspired by the human mind’s approach for solving complex reasoning tasks through trial and error. In this process, the human mind explores the solution space through a tree-like thought process, allowing for backtracking when necessary. To implement ToT as a software system, we augment an LLM with additional modules including a prompter agent, a checker module, a memory module, and a ToT controller. In order to solve a given problem, these modules engage in a multi-round conversation with the LLM. The memory module records the conversation and state history of the problem solving process, which allows the system to backtrack to the previous steps of the thought-process and explore other directions from there. To verify the effectiveness of the proposed technique, we implemented a ToT-based solver for the Sudoku Puzzle. Experimental results show that the ToT framework can significantly increase the success rate of Sudoku puzzle solving.
Before, I'd often end up going down the wrong rabbit hole when I would try to write out my problem in a google-able manner.
I've been using github copilot and I'm waving back and forth between impressed and unimpressed; it regularly generates syntactically invalid code for me. Eg calling lib functions that don't exist.
Thanks!
The conventional narrative is that chatGPT can code as well as a junior engineer, but I feel like it’s more like a senior engineer scribbling on a whiteboard—no, it’s not perfect code, but it’s the right idea.
We see a lot of articles that swing too far towards “AI will change everything!” just as much as “no, AI is not actually effective/meaningful!”. This is the latter. I’m surprised that anyone would genuinely try to use chatGPT/language LLMs this way.
I’m mainly enjoying chatGPT as a way to distill widespread information on the internet into a pseudo conversation. It’s great for learning new topics - my favorite is generating and explaining code snippets for languages/libraries I’m not familiar with.
For me this is one of the more dangerous uses.
Humans are already pretty bad at detecting errors in code. Bertrand Meyer, an expert with some renown in formal methods, couldn’t find an error in a one-liner of Eiffel code generated by ChatGPT. What hope do programmers with less training have to recognize when ChatGPT has given them an incorrect summary?
I work in the code security industry, and there is one truism, that is people typically write code until it compiles and or doesn't return an immediate error, they do not write code until it is 'correct'.
The title does not claim that the "entirety" of AI's success is misrepresented. It is questioning "how much" is, and I think that this is a fair question. The author does admit that the first paragraph in the example does have valid information which shows the AI with knowledge. It's the second paragraph that is just fluff.
I think that it is valid to ask how much this "fluff" impresses us and leads us to see the AI as being knowledgeable. Perhaps the author does draw too strong of a conclusion, but he still makes a good point.
Of raw single prompt output of GPT-X?
Of chain of thought output of the best latest model?
Of tree of thought of the best latest model + plugins and other models for different opinions?
"According to a study by the Pew Internet & American life project,[3] 47% of American adult Internet users have undertaken a vanity search in Google or another search engine. Some egosurf purely for entertainment, such as finding celebrities with the same name. However, many people egosurf as a means of online reputation management. Egosurfing can be used to find data spills, released information that is undesirable to have in the public eye. By searching one's own name in an online search engine, one can take on the perspective of a stranger attempting to find out personal information. Some egosurf in order to conceal personal images or information from potential employers, clients, identity thieves and the like. Similarly, some use egosurfing to maintain a positive public image and to achieve self-promotion."[0]
Clearly the answer to the title is "basically none".
In one case I ask a question about X, and ChatGPT responds with an answer: Y because Z. Well Y is wrong, so I tell ChatGPT Y is wrong and I get a new answer: Not Y because Z.
In another case I ask a question about X, and I get an answer that talks endlessly about Z without ever giving the answer Y.
The reasoning all sounds good, it's relevant and useful, but the actual answer might as well be a Mad Lib.
That's my experience with most LLMs, especially LLaMA based ones. GPT-3.5 is better on this front (RLHF?) and GPT-4 is a lot better (massive model? More training?).
It didn't know who I was at all, a great relief all told. The eye of Sauron has not fallen on me yet, it seems.
If you were asked to describe a family member, partner, or best friend, I'd bet a good chunk of it would be Forer Effect like statements.
A good way to A/B test this theory would be to ask folks to pick between AI generated statements of a famous person, and snippets from a human-written profile on them.
They produce output that looks correct and that's an accomplishment on its own. Unfortunately the output often has little bearing on the reality presented to it.
Realistically... They need to evolve much more. As an experiment I paid people on fiver to do the exact same thing with the same prompt and they get it right. These are not Americans so the norms of American municipalities are new to them.
I think this is one of those cases where I don't think it's as much of a gotcha as the author is supposing because those tests act as a kind of mirror allowing you to project yourself onto the results and then see yourself in the reflection. It's why they're useful when you take them yourself as a self-reflection tool but not at all actionable when someone else tries to estimate you with them.
wink wink nudge nudge tarot works on the same principal
As a child, I would read the horoscope of the day with the wrong sign for mother and her female friends and they always thought it was talking about them and I never had the courage to tell them I switched as a joke. I keep imagining a fine tuned GPT for insights from the "mystic realm".
Might be flattering to people but that’s never my use case.