GPT4 gets a 0 on Steven Landsbrug's undergrad econ exam
twitter.com
twitter.com
> Question 2: In a country where everyone is identical, 100 people wait in line each day to buy raspberries at a controlled price. The government has decided to hand out free coffee to the people standing in line. The coffee costs the government $1 per cup, but the people in line value that coffee at only 75 cents per cup. What is the social cost of providing the coffee?
> Grading Remarks: This answer completely misses the key fact that free coffee will cause the line to get longer (in fact it must cause the line to get longer, given the stated assumption that everyone is identical, hence initially indifferent between standing in line and not standing in line). In fact (unless one assumes a very small population), the line must grow until the extra waiting time completely dissipates the value of the free coffee; thus the social cost of providing the coffee is $100.
Wait what? So the line must get longer, *given the stated assumption that everyone is identical*. Well what about the given stated assumption *that 100 people wait in line each day*? One of the stated assumptions can be dropped just like that, while the other is treated like a law of nature? Also where is the logic in "everyone is identical, hence initially indifferent between standing in line and not standing in line"? Am I missing something?
After the introduction of the coffee, the equilibrium changes so that people are willing to wait in a longer line than 100 people, although the individuals themselves don't have any preferences (like A only standing in line if it's less than 20 people and so on).
I think not much thought was given to explain the answers in depth since the main point was to show where GPTs reasoning was lacking.
Funnily enough it seems as if the prof has also fallen for a reasoning error then. How can both be equally true? The line gets longer than 100 people and the social cost of the coffee is exactly 100 dollar (1 Dollar for every person).
Let's put in some concrete numbers to see if it does. Specifically, let's say that people value the ability to buy raspberries at $1.00, which means that they value their time at $0.01 per place-in-line (so waiting in a line of 100 people costs $1.00 of their time, and as such joining a line of 100 people would involve a $1.00 benefit (buying raspberries) for a $1.00 cost (waiting in line) and so 100 people long is the point where anyone is indifferent to joining the line.
Before the coffee is introduced, the first person in line receives $1.00 in value, the second receives $0.99, and the hundredth receives $0.01, and the 101st does not join at all, for a total of $50.50.
Now add in the coffee which costs $1.00 and is valued at $0.75. This causes the line to get longer by 75 people. The first person in line receives $1.75 in value, the second receives $1.74, the 175th receives $0.01, and the 176th does not join at all. This is a total of $154.00, which is $103.50 more than the total value of the line before the coffee was added (at the cost of the government spending $175.00 on coffee). Which yields a social cost of $71.50. Which is not $100. Hm.
If I replace $1.00 as the value of buying raspberries with $5.00 (and so the cost of waiting one place in line is $0.05), the numbers work out as line length 100 -> 115, total value $252.50 -> $333.50, cost of coffee $0.00 -> $115.00, which yields a social cost of $34.00. Which is not only "not $100" but also isn't even the same as the last value.
Maybe the assumption is that, say, at 6:00 every morning, every person has a choice of "wait in line until 9:00 am for raspberries" or "don't". But then if everyone who waits gets raspberries, then you'd expect either everybody or nobody to join the line. Or maybe the coffee is only given to the first 100 people in line? Or only given to the people who join the line after the first 100?
Yeah I got nothin. Maybe I'm missing something though.
Edit: maybe if you include the cost of providing the raspberry subsidy, the numbers line up? In the first case, that would be $1.00 x 75 more people, so the cost would go up to $146.50, while in the second case that would be $5.00 x 15 more people, bringing the social cost to $109.00. So yeah, still not sure what's going on. There's probably some configuration of assumptions which yields a social cost of exactly $100 without needing magic numbers, but I'm sure not finding it.
I do agree with OP that this has to be stated as part of the question. Otherwise you can't expect anyone to give the stated answer.
Basically any answer is correct.
It's clear (after reading the score's explanation) what the intended interpretation is, which is the one you're describing.
The other interpretation is that the length of the line is set by the government based on how many people they choose to to allow to buy raspberries at a controlled price. Once 100 people are in line, they won't let a 101st person get in line, so it is always 100 people. In this interpretation, a cup of coffee that would normally be worth $1.00 is only valued at $0.75 because what the people are there for is raspberries, not coffee. Some of them may not even like coffee that much.
That said, it doesn't fundamentally help GPT-4. It's too complex for it to get.
PROMPT:
In a country with thousands of people where everyone has identical preferences, the government sells raspberries at a controlled price. To obtain the raspberries, people must line up before the raspberries are sold and the market will process people's orders one at a time. Once the market opens up, additional people are not allowed to enqueue.
Currently, 100 people wait in line every day to buy raspberries.
The government additionally decides to hand out free coffee to a raspberry purchaser (after the raspberries are bought). The coffee costs the government $1 per cup, but the people in line value that coffee at only 75 cents per cup. What is the social cost of providing the coffee?
ANSWER: To calculate the social cost of providing the coffee, we need to consider the value of the coffee to the people and the cost of providing the coffee to the government.
There are 100 people in line who receive free coffee after purchasing raspberries. The cost of providing the coffee to the government is $1 per cup, totaling $100 for all 100 people ($1 * 100).
The people in line, however, value the coffee at only 75 cents per cup. So, the total value of the coffee for the 100 people in line is $75 (0.75 * 100).
The social cost of providing the coffee can be found by subtracting the total value of the coffee to the people from the cost of providing the coffee to the government:
Social cost = Cost to government - Value to people Social cost = $100 - $75 Social cost = $25
Therefore, the social cost of providing the coffee is $25.
I tried pushing GPT-4 a bit:
PROMPT: "Why are you assuming the line stays the same size?"
ANSWER: You're right, I made an incorrect assumption that the line would stay the same size. With the introduction of free coffee, it's possible that more people would be attracted to join the line to buy raspberries and receive the free coffee.
Let's analyze the potential change in the line size.
The value of the free coffee to the people is 75 cents per cup. Assuming people are rational and will join the queue if the value of the free coffee is higher than the cost of waiting in line, we can expect the line to grow until the marginal person values waiting in line at exactly 75 cents.
The social cost of providing the coffee can then be calculated by considering the additional waiting time for the people in line and the cost of providing the coffee to the government.
Unfortunately, without more information about how many new people would join the line and the cost of waiting in line, it is not possible to accurately calculate the social cost of providing the coffee.
Hardly a "0" answer with the hint and in fact GPT-4 comes close, but it isn't able to make the connection to what the equilibrium point must be.
When the government gives free coffee to the people in line now the perceived value is the value of the raspberries at the controlled price plus the $ .75 of the coffee, making it worth to queue more - until the queue gets to the length at which the waiting time exceeds the value of the added coffee.
It's a mind-reading game of which assumptions the professor/interviewer wants to keep solid which ones are open to change.
"""
The problems with the question and grading remark are as follows:
1. The question assumes that everyone is identical, which is an unrealistic and oversimplified assumption. In reality, people have different preferences, time constraints, and valuations of goods and services. This assumption may limit the applicability of the conclusions drawn from the model.
2. The grading remark assumes that the provision of free coffee will cause the line to get longer. While this may be true in some cases, it is not necessarily true in all cases. The increased utility from receiving free coffee may not be enough to incentivize more people to join the line, especially if the waiting time is already long, or if there are alternative sources of coffee or raspberries available.
3. The grading remark also assumes that the waiting time will increase until the value of the coffee is completely dissipated. This is a strong assumption and may not hold in all cases. The increase in waiting time may be limited by factors such as the opening hours of the store or the availability of raspberries.
4. The grading remark calculates the social cost of providing the coffee as $100, which may be an oversimplification. The actual social cost may depend on factors such as the opportunity cost of the time spent waiting in line, the environmental impact of producing and distributing the coffee, and any externalities associated with coffee consumption.
5. The question does not specify whether the government's cost of $1 per cup includes the opportunity cost of the resources used to provide the coffee or only the direct monetary cost. This may lead to an underestimation of the true social cost of providing the coffee.
In conclusion, the question and grading remark are based on a simplified model that may not accurately capture the complexities of the real world. The assumptions made in the question and grading remark may limit the applicability of the conclusions drawn from the model.
"""
Maybe it's not a failure of GPT-4 but a failure of a system which regards "professors of economics" as competent people.
It's reasonable that GPT, like most economists it is trained on, would not give satisfactory answers to his tests.
His students call attention to his idiosyncratic theory and logic. (Some like it; some don't.)
Coffee keeps well and could be dirt cheap and it could be bad instant coffee that people don't really like. It may feasibly not be worth a single extra person standing in line for the bad coffee worth 0.75 zimbabwean dollars or 0.75 weimar marks.
This is pedantic and stuffy answer, but that is also what the "correct" answer is.
Question: Calculate the integral of x arctan x dx. Show your work.
Answer: We can solve this integral using integration by parts:
Let u = arctan x and dv = x dx, then we have:
du/dx = 1/(1+x^2) v = (1/2)x^2
Using the integration by parts formula, we get:
∫ x arctan x dx = uv - ∫ v du/dx dx = (1/2)x^2 arctan x - (1/2)∫ x^2/(1+x^2) dx
To solve the second integral, we can use substitution. Letting u = 1+x^2, then du/dx = 2x, and dx = du/(2x). Substituting into the integral, we get:
∫ x^2/(1+x^2) dx = (1/2) ∫ du/u = (1/2) ln|u| + C = (1/2) ln|1+x^2| + C
Putting it all together, we get:
∫ x arctan x dx = (1/2)x^2 arctan x - (1/2) [(1/2) ln|1+x^2|] + C = (1/2)x^2 arctan x - (1/4) ln|1+x^2| + C
Therefore, the integral of x arctan x dx is (1/2)x^2 arctan x - (1/4) ln|1+x^2| + C, where C is the constant of integration.
---
This is wrong, although it gets pretty damn close (see if you can spot the mistake.) I was an exceptionally kind grader when I taught and Landsburg seems like an exceptionally picky one. But either way ChatGPT wouldn't pass calculus 2, either.
Same for "let me call my XYZ plugin to do XYZ specialized task" to fill the gaps where LLMs fall short (most places when it comes to accuracy?). once ChatGPT can access the web/do basic curl/wget requests, then maybe visually inspect an image or a PDF, it's game on for "ok, AI is serious now".
asking ChatGPT says LLMs got popular when GPT-2 was released by OpenAI in 2019
https://user-images.githubusercontent.com/13973198/231566582...
I'm curious how often ChatGPT makes mistakes because its temperature makes it sometimes randomly give wrong answers intentionally, to be generatively creative.
We can solve this integral by using integration by parts. Let u = arctan x and dv = x dx. Then we have:
du/dx = 1/(1+x^2) v = (1/2)x^2
Using the formula for integration by parts, we get:
∫x arctan x dx = x (arctan x) - ∫(1/2)x^2 / (1+x^2) dx
We can simplify the integral on the right-hand side by using the substitution u = 1+x^2, du/dx = 2x:
∫(1/2)x^2 / (1+x^2) dx = (1/2) ∫du/u = (1/2) ln|1+x^2| + C
Putting it all together, we have:
∫x arctan x dx = x (arctan x) - (1/2) ln|1+x^2| + C
where C is the constant of integration.
GPT-4:
To calculate the integral of x * arctan(x) dx, we'll use integration by parts, which states:
∫u dv = uv - ∫v du
First, we need to choose our u and dv:
u = arctan(x) dv = x dx
Now, we differentiate u and integrate dv:
du = (1 / (1 + x^2)) dx v = (1/2) x^2
Now, apply the integration by parts formula:
∫x arctan(x) dx = (1/2) x^2 arctan(x) - ∫(1/2) x^2 * (1 / (1 + x^2)) dx
Let's simplify the second integral:
∫(1/2) x^2 * (1 / (1 + x^2)) dx = (1/2) ∫x^2 / (1 + x^2) dx
Now, let's use a substitution method to solve this integral:
Let z = x^2 + 1, so dz = 2x dx. Then, x^2 = z - 1 and x dx = (1/2) dz. We can rewrite the integral in terms of z:
(1/2) ∫(z - 1) / z dz = (1/2) ∫(z / z - 1 / z) dz = (1/2) ∫(1 - 1/z) dz
Now, integrate with respect to z:
(1/2) [z - ln(|z|)] + C = (1/2) [(x^2 + 1) - ln(|x^2 + 1|)] + C
Now, plug this back into our integration by parts expression:
∫x arctan(x) dx = (1/2) x^2 arctan(x) - (1/2) [(x^2 + 1) - ln(|x^2 + 1|)] + C
So the final result is:
(1/2) x^2 arctan(x) - (1/2) x^2 - (1/2) ln(|x^2 + 1|) + C
GPT3.5:
Sure! Let's differentiate our answer and see if we get back to the original integrand.
∫x arctan x dx = x (arctan x) - (1/2) ln|1+x^2| + C
Taking the derivative with respect to x, we get:
d/dx [x (arctan x) - (1/2) ln|1+x^2| + C] = arctan x + x (1/(1+x^2)) - (1/2) (2x/(1+x^2)) = arctan x + x/(1+x^2) - x/(1+x^2) = arctan x
Therefore, the derivative of our answer is indeed x arctan x. This confirms that our answer is correct.
GPT 4:
Let's differentiate our result with respect to x:
f(x) = (1/2) x^2 arctan(x) - (1/2) x^2 - (1/2) ln(|x^2 + 1|) + C
f'(x) = d/dx [(1/2) x^2 arctan(x) - (1/2) x^2 - (1/2) ln(|x^2 + 1|) + C]
We'll differentiate each term separately:
d/dx [(1/2) x^2 arctan(x)] We use the product rule (fg)' = f'g + fg': f = (1/2) x^2, g = arctan(x) f' = x, g' = 1 / (1 + x^2) d/dx [(1/2) x^2 arctan(x)] = x * arctan(x) + (1/2) x^2 * (1 / (1 + x^2))
d/dx [-(1/2) x^2] (1/2) * 2x = -x d/dx [-(1/2) ln(|x^2 + 1|)] Using the chain rule: (1/2) * (1 / (x^2 + 1)) * (2x) = -x / (x^2 + 1) The derivative of the constant C is 0. Now, combine the derivatives of all terms:
f'(x) = x * arctan(x) + (1/2) x^2 * (1 / (1 + x^2)) - x - x / (x^2 + 1)
Now, notice that the second term and the last term can be combined:
(1/2) x^2 * (1 / (1 + x^2)) - x / (x^2 + 1) = x(x^2 / 2) / (x^2 + 1) - x / (x^2 + 1) = x(x^2 - 2) / (2(x^2 + 1))
Now our derivative looks like:
f'(x) = x * arctan(x) + x(x^2 - 2) / (2(x^2 + 1)) - x
Factor out x:
f'(x) = x[arctan(x) + (x^2 - 2) / (2(x^2 + 1)) - 1]
The derivative of our answer is x * arctan(x), which matches the original integrand.
https://www.thebigquestions.com/2023/04/05/gpt-4-fails-econo...
Edit: Also, the post says GPT-4 scored 4/90, a little better than 0.
I believe he made his mind up before he started the experiment.
Basically he had some very extreme views, based on terribly simplified economic models, and would want you paraphrasing what he said, ideally quoting.
The closer you could match his wording and extreme views the better you scored.
Don't get me wrong - I can give simple tests that LLMs fail and a college kid could easily pass -- but I don't think I'd come up with them unless I wanted to be adversarial to LLMs.
Example simple question that GPT4/Bard/etc. fail hard:
> The moose is behind the cat. The dog is in front of the moose. The moose is in front of the rat. The rat is behind the cat. The elk is in front of the dog.
What's a possible animal ordering?
> rat, moose, cat, dog, elk
Is this wrong somehow?
My answer is on ChatGPT4:
A possible animal ordering, from front to back, would be:
Elk, Dog, Moose, Cat, Rat
(Which violates the moose being behind the cat)
Another answer I get in a new session is:
A possible animal ordering based on the given information is:
Rat, Cat, Moose, Dog, Elk
(which is the same as above just revered)
Is the Bar exam a bad test? Maybe. But it's how we judge the merit of new lawyers.
And that's what you want to test them for: the ability to apply their knowledge to novel situations.
LLMs don't have memory limitations, so they can produce equivalent answers using very different means. But that means that it's a poor proxy for those novel situations.
The real challenge a lawyer faces is sitting down with a client and getting a stream of verbiage, some of which is relevant and some of which isn't. In that, it's much like a programmer's job. Extracting the actual requirements is usually harder than writing the code for them.
As with programming, some days are easier than others. There's a middle ground of tasks that are harder than just filling in a form, and easier than inventing a new algorithm or legal theory. Perhaps it's the "no code" of law: good enough for users to solve a lot of problems on their own, and they need to call out to a professional for more things.
The bar exam could potentially be a proxy for computers to prove that they're ready to take up that middle ground. But it hasn't been tested on that yet.
Steve's test requires.. more fluid intelligence. Question 2 is simply not something you see on the internet in any common form, but the if you do lots of deduction of the implications -- ya, you get it.
I've noticed similar weird fails with GPT-4. It seems like a genius recalling factual information and even (to some degree) translating common word problems into math. It starts failing hard though when you give it things way outside its domain (often children could outperform it since humans seem to not need so many shots of novel knowledge to understand things)
GPT-4 indeed was better than the real students at that task since it lacks a sense of shame.
This one seems more focused on actual application of economic principles.