Is GPT-4 a good data analyst? (2023)
arxiv.org
arxiv.org
All this real-world complexity can be tamed by stuffing the prompt with a ton of relevant context and an amazing prompt engine. We'll have bots that autonomously query the database hundreds of times building a 5 page "deep-dive" analytics report in minutes.
At least that's what we're trying at patterns.app.
We just started with some simple examples/questions. Then got a bit more complicated. But the result was pretty spot on the first time.
Simple 1 table dataset: https://youtu.be/vM0NAIROTD4 Analyzing Log Data: https://www.youtube.com/watch?v=5hbanOZjHm4 Joining 2 datasets: https://youtu.be/VUc0flsC4pg Simple Filtering example: https://youtu.be/wTiaq2J-IdY Analyzing Ride Share data: https://youtu.be/HWZUHxT7Vx8
Okay, I know picking people's sentences apart has fallen out of fashion, but:
"Divergent opinions" are ... opinions. A "definitive conclusion" is ... a conclusion.
I see more examples, but I wanted to make a point: I miss the days when fewer words conveyed more meaning. From the classic https://en.wikipedia.org/wiki/The_Elements_of_Style: "Make every word count."
About brevity of expression, I must add this (possibly apocryphal) story about Ernest Hemingway. In the 1920s Hemingway and his Paris friends had a contest: who could write the shortest readable short story? Hemingway won with this entry:
For sale. Baby shoes. Never worn.
Also just for you I copied and pasted my original comment into ChatGPT to make it more brief ;)
as an aside, asking chatGPT 3.5 to make similar stories to "For sale. Baby shoes. Never worn." had very bad results and completely failed to do a single meaningful example of "common item with uncommon description that tells a story"
For clarity is not really acknowledging it made a mistake. "Calling out" an LLM's mistake just leads to the next most likely text to be something that sounds like an acknowledgement of a mistake, but the same is likely to happen if the LLM generated a correct response and you respond claiming that it's incorrect.
Using tool inappropriately leads to suboptimal outcomes -- news at 11.
A good mental model is that an LLM is a blurry JPEG of the Internet.
You sound like a scientist, right? You reference "published research", after all.
What would your opinion be of a researcher measuring the exact values of the pixels of a JPEG image instead of the RAW sensor data?
No, you're not going to get the precise values back out either with an LLM or a JPG image.
Yes, the picture will still "look like" the original.
You asked the LLM to output a URL or reference that looks like something that would appear in a research paper. It did so.
You're complaining the precise byte sequence returned is not what you wanted.
The error is with your expectations, not the tool.
Them: Can you just quickly pull this data for me?
Me: Sure, let me just:
SELECT * FROM some_ideal_clean_and_pristine.table_that_you_think_exists
GPT-4 is good on a single CSV, but breaks down quickly applied to a real database / data warehouse. I know they're using multiple tables in the paper, but it appears to be a pristine schema that's very easy to reason about. In the real world, when you're trying to join postgres to hubspot and stripe data, an LLM isn't able to write the SQL from scratch get the right answer.We're working on an approach using a semantic layer at https://www.definite.app/ if you're interested in this sort of thing.
0 - https://twitter.com/sethrosen/status/1252291581320757249?lan...
A semantic layer is just really all about defining data models in the domain of interest. It's the hardest part in dealing with data strategies, very manual, very company and process and history specific.
Once it's defined, the next set of tasks is to make sure that the data in the model is correct and coherent. And only then, querying this data, applying ML etc start becoming worthwhile.
We at https://syncari.com take the centralized data model centric approach. https://www.definite.app/ also looks very cool!
There’s more to it, but the tooling to create a GPT is basically a hand-holding mechanism to create a prompt.
There’s even a handy aphorism to remind you that the user is never to blame: “You’re holding it wrong.”
Jokes aside, I wonder what the general writing abilities and communication skills are for people that cannot for the life of them get usable results from an LLM.
The terms "depend" and "require" there are the hard versions. You can't send people to the moon on the outputs of LLM's.
The tools are good for certain tasks and getting better. Master these tools and be ready for what released in the coming months and years
You are either at the center turning the wheel, or you’re on the outside getting spun
Here, argue with Yann, who makes a statement about how language isn't enough to produce a mind: https://twitter.com/ylecun/status/1768353714794901530
https://youtube.com/playlist?list=PLfRs4l8INQS44y62Z2aFzPzHH...
> LLMs... have demonstrated their powerful capabilities in ... context understanding, code generation, language generation, data storytelling, etc.,
LLMs have not demonstrated understanding (in fact, one could argue that they are fundamentally incapable of understanding); they have only AFAICT demonstrated the ability to generate boilerplate-ish code; "language generation" is too general a task to claim that LLMs have succeeded in; and as for data storytelling I don't know, but they can spin yarns. The problems is that those yarns are often divorced from reality; see:
https://www.ibm.com/topics/ai-hallucinations
--------
Leafing through the paper, and specifically tables 6 and 7, I don't believe their conclusion, that "GPT-4 can perform comparable [sic] to a data analyst", is well-founded.
Do I, myself understand? Stand under what exactly? What is that supposed to mean?
That's trivially true. The question is: are we any different?
So of course I feel a bit offended when people claim LLMs are just stochastic parrots, because it doesn't feel to me, that I'm specifically any better?
My thoughts - they just happen, and sometimes not in my favor - I have had times of depression, I didn't have control over my thoughts. Neither do I have now, but at least I am in a better place. Because the "happiness" chemicals are regulated to be in a more favorable state to me for various different factors.
I didn't know what I was going to comment in response to your comment, I was just streaming my conscious.
You and I know there’s a truth and we’d like to find it. The GPT is just happy (I.e. rewarded) to produce frequently used tokens.
Worth a shot :)
"""I think dogs are wonderful! They're known for their loyalty, playfulness, and their ability to bring joy to people's lives. What about you? Do you have a favorite breed or dog story?"""
What do I ask next?
"I'm not a fan of dogs. I do know a few dogs though. Sometimes I invite my neighbour's dog for dinner. He's got good taste, for a dog. The last time he came around we talked about the situation in the Middle East. Do you know a good book about this topic that I could recommend to him?"
You can't be suggesting it has memorized relationships between all concepts, the model would be enormous.
So clearly, there is something else going on. It's able to encode concepts/ideas.
Ask an LLM to extrapolate, see any semblance of reason collapse.
Laundry service is complimentary, but our records show that you haven't registered your home. Would you like to register your home address at this time?
Have you told someone something like "no, don't do it that way because <insert non-obvious downstream problem>, instead, do <insert alternative strategy that achieves better outcomes>"? That's an artifact of understanding and of the model you've developed for that thing.
It's well described as a mixture of knowledge and wisdom, and is essentially the property of knowing the effect that pulling a lever will cause, coupled with good judgement about how, when, and why to pull the lever.
Usually in a bit more polite way.
That my approach is a "novel" and an "interesting" approach, hinting that it's really probably not the best option here.
The thing about LLMs is exactly that they don't understand by design. It often feels very distinctly like it's just engaging in sophisticated wordplay. A parlor trick.
When ChatGPT 4 first came out I spent a couple of hours putting together a chess game using ChatGPT as the engine. It was shockingly bad, as in even attempting to make invalid moves.
I get it: it's not tuned for that purpose, and its chess training corpus could probably be expanded to improve it as well.
But, it actually served as a near-perfect demonstration of its lack of understanding, as well as the confidence with which it asserts things that are simply wrong.
On a recent integration project with a good bit of nuanced functionality, it led me astray multiple times. I've gotten to a point where I can feel when its answers are not quite right, particularly if I know just a little about the topic. And, when challenged, it does that strange thing of responding with something along the lines of, "My apologies you're completely right that I was completely wrong".
Over time, there becomes a sense that there is no there there. Even it's writing capabilities, lauded by so many, are of a style that is superficial and perfunctory or rote. That makes sense when you know what it is, but that's the thing: we get articles like these, lauding its wisdom.
What impressed me is when I asked to make it more snake-like (what does that even mean right?).
It changed the colors to shades of green, used italic fonts, added some hisssssing sssstuff to wordssss, and added a diamond pattern through the background.
It was a dumb and not very fancy site, but I'm not sure you can say it doesn't understand anything at all when you ask it to make a website more snakelike and actually made a pretty good attempt at doing it.
But I think it comes down to whether it can reason about things and whether it can draw new conclusions or create new information as a result.
Your snake site is probably a good example. ChatGPT has a bunch of words that it knows are associated with snakes. It's pretty straightforward pattern matching. It doesn't really "understand" what those words mean, except that they have relationships to other words.
But, if you were to ask it to reason and draw new conclusions about these things beyond its training corpus, it would be unable to reliably do so.
Similarly, it had no idea about the quality (and sometimes legality) of the chess moves it generated.
I really think chess is just a terrible example though. You're really asking a lot of it but I'm honestly shocked it can do what it can. It seems to know some opening books, but falls down immediately. Which really makes sense because you'll find a lot of reading material on specific openings, but the problem space of the game is just too big to find texts about any given game state. Maybe if you have it reason about the board state and "think" about tactics you could push it farther. But we've already solved this.
We have Stockfish et al and they've literally changed the game. Asking an LLM to play chess, while cool, is like trying to train a fish to dance. I think once we have an AI that's built with a bunch of different models that specialize in different things the idea of understanding is going to get even blurrier to the point that we might even say "yes, it doesn't 'understand' things, but it's better at humans at literally everything" so the difference becomes meaningless.
I'm also of the opinion that humans are fancy automatons, so I tend to argue both sides. I'll say yeah it's thinking and so do we or ok it's not, but either do we.
100%. I didn't expect much from it going in. Was more curious to know how good/bad it might be and, really, whether it could play at all.
But, I mention it here because it illustrates the disparity between understanding/reasoning (LLMs bad) and semantics (LLMs good), per the subject of the thread.
>get even blurrier to the point that we might even say "yes, it doesn't 'understand' things, but it's better at humans at literally everything" so the difference becomes meaningless.
I get exactly what you're saying here and I've wondered the same. But, that's the thing I'm not sure about, given its inability to reason. I mean, there's the idea that it doesn't matter whether it "understands" if it correctly answers. But, then I wonder how it can "truly" be better than humans at "literally everything", if only humans truly understand.
>But, I mention it here because it illustrates the disparity between understanding/reasoning (LLMs bad) and semantics (LLMs good), per the subject of the thread.
LLMs can play chess (links above) so i'm curious if you still have this opinion.
Largely. I'm not surprised that another model could be trained on a larger chess corpus or otherwise tuned for improvement. I actually alluded to that in my original comment above.
In fact, I was really surprised that gpt4 was so bad, given the total volume of training data it had obviously ingested. So, my initial point was around how that revealed the disparity between LLMs appearing to understand versus actually understanding (i.e. engaging in reasoning).
There is no difference between this exercise and the one you have given GPT-4. In fact, the results are manifested the exact same way. By at most a few moves in, i will notice that both the stranger and GPT-4 are moving pieces to some general area of the board but woefully failing to adhere to the rules of the game. "The stranger can not reason" would be an absurd conclusion. Swapping out The stranger for GPT-4 does not make it any less so.
So really by this standard then Humans regularly only "appear to understand" a great number of things. I can get behind this, but I doubt this is your conclusion. It seems to me that you were using "appear to understand" and "understand" as something that separates Humans from LLMs.
Further, that knowledge represents a vast body of training data that certainly includes the rules of chess, how pieces move, etc. This is verifiable. Go ask ChatGPT those questions now. How does a bishop or rook move? Can a piece move through or over another piece, etc. What is en passant? When can it be employed? You will find that it answers these correctly, and even provides examples. It talks about ranks and files, and authoritatively sprinkles in other knowledge related to the questions.
But, then you put a board in front of it and you ask it to make moves, and it doesn't adhere to these rules. You're not asking it to win the game. Just make a legal move. But it cannot, even while confidently responding with invalid moves.
So, it's not about whether it's been tuned to be a grandmaster. It's about whether it can even execute on the knowledge it clearly possesses.
Now, turning back to your analogy with a human player. You ask Gary all of these questions and he gives you the answers perfectly, so you recognize that his factual knowledge is intact. But, despite his best efforts, he repeatedly makes illegal moves. You would have to conclude that there is some problem with reasoning or understanding what those facts actually mean in practice.
That is the difference between appearing to understand and actually understanding. And certainly, there are some humans who fit this "appearance" description as well, perhaps through some processing deficit or otherwise. But as a broad category, this is something humans do quite well.
The concept of understanding is really not complicated. When we explain something to someone and we ask the question "do you understand?", we're not asking whether they can memorize or repeat back what we just told them. We're asking something else. That something else is the well-understood (irony intended) meaning of "understand".
The premise of ChatGPT has never been that it knows everything.
>So, it's not about whether it's been tuned to be a grandmaster. It's about whether it can even execute on the knowledge it clearly possesses.
Factual knowledge on paper has never been a guarantee you can actually perform the task. In fact, it usually isn't.
>And certainly, there are some humans who fit this "appearance" description as well, perhaps through some processing deficit or otherwise. But as a broad category, this is something humans do quite well.
I'm telling you this is a normal occurrence and not something that only happens with a "processing deficit".
You can rant about all the details or rules of a tennis game all you like and fail to play it, you can know the physics theory but still struggle to solve problems. Hell we don't need to leave chess, give our stranger a proper book with rules and the strategies employed by grandmasters. For a while after, He's still going to make illegal moves and he will be nowhere near the level of a grandmaster.
>Again you are still not bringing up anything that serves as the distinguisher you want it to serve. The concept of understanding is really not complicated. When we explain something to someone and we ask the question "do you understand?", we're not asking whether they can memorize or repeat back what we just told them. We're asking something else.
Ok and ? GPT passes many tests that demonstrate understanding.
This analogy reveals what you're missing. The correct analogy is not that you know the rules, but "fail to play" tennis (which requires physical aptitude, etc). It's that you can cite the rules, but actually try to hit the ball over the fence and claim a point when you succeed.
Likewise, the stranger with the chess rule/strategies book. You're conflating quality with understanding.
But, I've not claimed that all humans excel at or even understand all things.
And, you've still not defined the gap between GPT being able to cite the rules of chess, yet not being able to apply them.
>GPT passes many tests that demonstrate understanding.
It actually does not. It's just better at making you believe it does when it simulates understanding in those tests.
I'm not because he will still make illegal moves for some time and it appears no different. It's not just about level of play.
>And, you've still not defined the gap between GPT being able to cite the rules of chess, yet not being able to apply them.
GPT-4 can't play chess. It's that simple. Can cite rules but can't play is a common enough occurrence in humans. Sure it's interesting to see that disconnect but nothing so bizzare it needs to be "explained".
And the training process is "dumb" like Evolution. It's not like a person reading a book anymore than evolution is like a person trying to create the "ultimate" animal so you can see things like this more often where there's a disconnect between the on paper rules and the environment for playing.
>It actually does not. It's just better at making you believe it does when it simulates understanding in those tests.
I'm sorry but this is just nonsense. Simulating understanding is a nonsensical argument. A plane isn't "simulating flight". It flies. If GPT can make it through whatever rigorous test you have in mind then it understands that thing. You want to believe in an imaginary distinguisher that's so important but can't be tested for and it's silly.
It's not bizarre and I've already explained it: ChatGPT has no understanding of the rules that it communicates so perfectly. We've openly observed it in its errors. You just responded with another denial of that obvious fact, but you're now acknowledging that "GPT-4 can't play chess. It's that simple". So you've essentially agreed with what I've said, except to say you don't agree, but that you prefer no explanation to the obvious one.
I'm not quite sure what to do with that.
>A plane isn't "simulating flight". It flies.
Respectfully, that is another irrelevant analogy.
>You want to believe in an imaginary distinguisher that's so important but can't be tested for and it's silly.
This entire discussion is about the distinguisher, so I wouldn't classify it as imaginary. You just keep repeating that it doesn't mean anything.
But, I do agree that a model can be trained to a sufficient degree that its ability to demonstrate understanding is immaterial in effect from actual understanding.
I also agree that this discussion has veered into silly territory. I clearly am not getting the point across, so I will respectfully thank you for the engagement.
But, even ChatGPT will tell you that it doesn't "understand" in the traditional sense. That it essentially finds patterns, etc. So perhaps it will convince you where I could not. :)
Thanks again for the discussion.
Right
>So you've essentially agreed with what I've said, except to say you don't agree, but that you prefer no explanation to the obvious one.
I never disagreed that 4 can't play chess. It's right in my first comment. I literally say LLMs can play chess and you just tested a model that couldn't.
>This entire discussion is about the distinguisher, so I wouldn't classify it as imaginary. You just keep repeating that it doesn't mean anything.
>But, I do agree that a model can be trained to a sufficient degree that its ability to demonstrate understanding is immaterial in effect from actual understanding.
You're still making the same mistake here.
If a scientist happens upon a piece of yellow metal and runs every gold test there is on the metal and it passes then that is gold. What you are doing here is essentially saying that it is "fake gold that is immaterial in effect from real gold". It's just kind of a meaningless statement.
The only thing any human does is "demonstrate understanding". I can't peek into your brain and say you understand. You can't prove to me right now that you "actually understand".If I think you understand it's because I say you demonstrated it. There is no "actual understanding" separate from demonstrating it.
>But, even ChatGPT will tell you that it doesn't "understand" in the traditional sense. That it essentially finds patterns, etc. So perhaps it will convince you where I could not. :)
Either GPT doesn't understand what it's saying and this holds no importance or it does. You have to pick one.
I could just as soon go to Claude to have this discussion and get different answers. Open AI obviously train it on how they want it to respond to this sort of questions.
I've agreed that there can be the effect of an immaterial difference in a given context or application. For instance, if I'd stopped at asking ChatGPT about the rules, then it wouldn't have mattered that it couldn't employ them properly. But, that's context-dependent (e.g. when I asked it to play, it did matter and the difference was no longer immaterial).
In your gold analogy, it's either gold or it's not. If our tests couldn't distinguish it, then it doesn't change the fact of what it is. It just represents a shortcoming in our ability to discern. If it had the same conductive properties as gold such that it could be used in applications that required those properties, then there's no material difference in that application. But, if we later find that it wears out twice as fast, then we realize it's not gold and the difference becomes material.
But, here's the key piece you're missing: even if we never find a demonstrable difference in the metal, it doesn't change what it is. That is, even if the difference is immaterial for any application or test we can devise, the facts don't change.
You can argue that if the difference never becomes material then it doesn't matter. Maybe. But, that's not a practical way to approach AI, as it leaves open the possibility that it fails in unexpected ways.
So, in the case of LLMs, I think it matters. Massively. In how we approach it, use it, train it, understand its limitations, our expectations, etc. But, I'm not making that argument.
>Either GPT doesn't understand what it's saying and this holds no importance or it does. You have to pick one
No. We've already agreed that GPT can get facts right. You're kind of going backwards on me.
At the bottom of it, we agree that GPT-4 can cite the rules, but can't play chess (i.e. use its facts). I believe that reveals how these models work WRT to reasoning and understanding. You disagree (or think it's immaterial).
So, we draw different conclusions from the same information. Not unprecedented in human history!
Just so you know, LLMs can play chess just fine. Like this isn't some inability. GPT-3.5-instruct-turbo plays at about 1800 Elo and makes legal moves consistently well past openings.
https://github.com/adamkarvonen/chess_gpt_eval
https://adamkarvonen.github.io/machine_learning/2024/01/03/c...
>Per move, a model gets 5 illegal moves before forced resignation of the round.
>Most of gpt-4's losses were due to illegal moves, so it may be possible to come up with a prompt to have gpt-4 correct illegal moves and improve its score.
My experience exactly with gpt4. I first solved by keeping a list of illegal moves and subsequently prompting to move again and explicitly exclude those moves. Idea was to constrain gpt as little as possible, but that became inefficient. So I began to prompt for each move with the complete set of legal moves to choose from.
In any case, at a certain point it becomes somewhat random when the goal is just to produce a legal move.
But, it's interesting that producing legal moves is a challenge for gpt4, especially given that it's such a discrete problem.
It's just too "deep" for an LLM and a poor standard regardless.
As an avid chess and GPT person, I've tried to get it to play decent chess but can't. Neither of your links support it, but if you've got anything else, send it.
1200 is the Elo of the small gpt model the github owner trained. 1800 is the Elo of gpt-3.5-turbo-instruct.
https://twitter.com/GrantSlatton/status/1703913578036904431?...
>As an avid chess and GPT person, I've tried to get it to play decent chess but can't.
You are not using the model that can play chess. It can only be accessed by Open AI's playground under the Completions endpoint.
https://platform.openai.com/playground?mode=complete&model=g...
Neither can humans, at least with our bare brains. We can do it by carefully observing the effects of our actions in the environment, but we are really studying the world and it takes time. Everything we know comes from the environment.
The brain by itself invents or discovers nothing, it is the data-engine made by action-effect-feedback that teaches us all we know. Without the ability to push and prod, set up our experiments and carefully observe effects we wouldn't be at our current level.
Environment is the teacher, but there is another important factor - language. Without it every one of us would have to rediscover from scratch. With it we can build upon other people to learn and cooperate. We encode everything we know in language. It acts like an evolutionary system of ideas.
LLMs have what is necessary, they can learn language pretty well, but until now have not been exposed much to the world. There are millions of chats but very little in other kinds of environments - computers, simulators, games, robots. LLMs can create their own experiences and learn from each other, and from us.
Open ended discovery is a grand project, a social process, it doesn't work well in one agent. Language is the linking element, and the world is the teacher. Some things are not written in any books, only the external world can teach us. Reasoning about things and drawing new conclusions depends on having access to an environment.
It kind of has a, "yeah, but if reality didn't exist" feel to it.
Could humans reason if we somehow lived in a void? I don't know. But, then I guess it really wouldn't matter.
On the other hand scientists need experimentation to come up with anything new. It never comes from pure thinking, there is always the experimental validation. The place we get all our knowledge from is the environment. Without environment as grounding we are just language hallucinators, remember the string theory - that was a nice ungrounded hallucination that kept physicists busy for decades. Without validation our theories are worthless. So a LLM trained purely on human text can't make radical discoveries but when it learns from the environment it can. AlphaGo discovered move 37 from its self-play experience.
Are these hypotheses frequently based on prior observations? Of course.
But, reason is still being applied, so I don't know what you're trying to differentiate. You say alphago discovered move 37 from its self-play experience, but that explicitly has nothing to do with the environment. It's the equivalent of a human thinking through various scenarios (without experimentation or interacting with the environment) to come to a new conclusion. The environment is not required for that, which completely argues against what you seem to be saying.
BTW, I have not made any statements about AI and reasoning, other than LLMs. AlphaGo, of course, is not LLM-based.
Of course we don't all live in individual voids, so much of what we are concerned about are functions of our environments and will be expressed as such. I don't think there's anything groundbreaking in observing that a lot of our reasoning is applied to solving problems that have some intersection with our shared reality.
> then experimenting to validate that hypothesis.
Proves my point that scientists can't secrete science from their pure brains, the environment is our teacher. We are just prying at its secrets with our limited brains, and usually need many of us to tackle one field, a distributed search for discoveries.
None of us is smart enough to get even to Newtonian physics from scratch, we need to stand on the shoulders of previous generations to get anywhere. We're not smart enough to do it without the environment or without a large number of people working together.
LLMs can play chess just fine. This isn't some inability. You just tested a model that couldn't. GPT-3.5-instruct-turbo plays at about 1800 Elo and makes legal moves consistently well past openings.
https://adamkarvonen.github.io/machine_learning/2024/01/03/c...
>Your snake site is probably a good example. ChatGPT has a bunch of words that it knows are associated with snakes. It's pretty straightforward pattern matching. It doesn't really "understand" what those words mean, except that they have relationships to other words.
Lol there is absolutely nothing straightforward about that. It's just so funny, people so ready to relegate anything to "pattern matching" that any supposed difference becomes meaningless.
The LLM understands relative patterns and/or associations. Statistically.
Snakes tend to be green. People have written countless books with hisssssssing from snakes and snake-like characters. Diamond patterns are well documented with snakes.
Again, it has no idea what these "are" and if we collectively decided snakes were called "Murder ropes" then it'd make the exact same associations.
If I did the exact same thing you'd be like "well, yeah, you're not an idiot, you understand what a snake is and what snake like things are."