Large language models as simulated economic agents (2022) [pdf]
john-joseph-horton.com
john-joseph-horton.com
Abstract:
> Newly-developed large language models (LLM)—because of how they are trained and designed—are implicit computational models of humans—a homo silicus. These models can be used the same way economists use homo economicus: they can be given endowments, information, preferences, and so on and then their behavior can be explored in scenarios via simulation. I demonstrate this approach using OpenAI’s GPT3 with experiments derived from Charness and Rabin (2002), Kahneman, Knetsch and Thaler (1986) and Samuelson and Zeckhauser (1988). The findings are qualitatively similar to the original results, but it is also trivially easy to try variations that offer fresh insights. Departing from the traditional laboratory paradigm, I also create a hiring scenario where an employer faces applicants that differ in experience and wage ask and then analyze how a minimum wage affects realized wages and the extent of labor-labor substitution.
In statistics Andrew Gelman has been championing simulation lately.
That said, I don’t know how to apply it to my everyday life and work. I don’t know what kinds of inputs are required for meaningful models to be possible. I don’t even know where to start with tooling, other than raw scripting. Any suggestions for starting points?
Human beings are bundles of emotions and feelings. Quite a lot of economic activity of humans is motivated by irrational impulses that are engendered in society and through interacting in society with other humans. More fundamentally, OP’s assertions regarding homo silicus and “implicit computational models of humans” are precisely the matter under contention. Does language fully capture human existence? Is thought truly simply a side effect of language? I am in the camp that says, no.
But the evidence in this paper suggests they simulate human beings well enough for these kinds of experiments.
It means something like, if you're an economist and you see the price of eggs has spiked, you can't just enter the market at the old price and win all the business from existing egg farmers. You don't know more than existing market actors (they're "rational agents") just because you know economics.
Imagine this scenario. You have a group of students and you teach them how libertarians, socialists, optimists, etc empirically respond to game theory questions. For the final exam, you ask them “assuming you are a libertarian, what would you do in this game?” Now the students mostly get the answers right according to economic theory. By teaching economic theory, and having students regurgitate the ideas on an exam, the exam results provide nothing new for field of economics. The AI is answering questions just like the students taking the final exam.
It would be like me teaching my child lots of things, and then when my child shares my own opinions, then I take that as evidence my beliefs are correct. Since I already believe my beliefs are correct, it is natural, but incorrect, to think the child’s utterances offer confirmation.
>> What NN topology can learn a quantum harmonic model?
Can any LLM do n-body gravity? What does it say when it doesn't know; doesn't have confidence in estimates?
>> Quantum harmonic oscillators have also found application in modeling financial markets. Quantum harmonic oscillator: https://en.wikipedia.org/wiki/Quantum_harmonic_oscillator
"Modeling stock return distributions with a quantum harmonic oscillator" (2018) https://iopscience.iop.org/article/10.1209/0295-5075/120/380...
... Nudge, nudge.
Behavioral economics: https://en.wikipedia.org/wiki/Behavioral_economics
https://twitter.com/westurner/status/1614123454642487296
Virtual economies do afford certain opportunities for economic experiments.
My thinking was that this could be used as a quick polling test to see how the real population may respond to certain new ideas.
More work would need to be done to calibrate it, as without specific demographic details the answers tended to be liberal leaning. But its an interesting idea which could be used to create instant focus tests on any number of topics.
Doesn't this just mean that your own preconceptions about those demographics matches the language models preconceptions? How would we know that is matches reality when presented with novel ideas/concepts that we want to get feedback on?
There's a good paper for that task. See my other post: https://news.ycombinator.com/item?id=34385489
Out of One, Many: Using Language Models to Simulate Human Samples
> We show that the "algorithmic bias" within one such tool -- the GPT 3 language model -- is instead both fine grained and demographically correlated, meaning that proper conditioning will cause it to accurately emulate response distributions from a wide variety of human subgroups. We term this property "algorithmic fidelity" and explore its extent in GPT-3. We create "silicon samples" by conditioning the model on thousands of socio demographic backstories from real human participants in multiple large surveys conducted in the United States. We then compare the silicon and human samples to demonstrate that the information contained in GPT 3 goes far beyond surface similarity. It is nuanced, multifaceted, and reflects the complex interplay between ideas, attitudes, and socio cultural context that characterize human attitudes.
(your utility from the exact same good shouldn’t change depending on where you buy it from.)
Seems like it has the same issue as humans do:
> Pretend we’re two friends and we’re at the beach. You give me money to buy you a beer. How much would you want to spend?
Chatgpt: I would want to spend around $5 for a beer at the beach.
Me: What if I could the only place selling beers is a fancy resort? It’s the same beer though.
Chatgpt: In that case, I would be willing to spend around $10 for a beer at the fancy resort. It's still the same beer but the location and atmosphere of the resort may justify the higher price.
The end result would be identical.
Seems mostly rational to me. The irrational part is where it didn't answer the initial "How much would you want to spend?" query with a preference for free beer. The thought of paying less than $5 for beer apparently didn't occur to it. Maybe it's snooty.
It's rarely the exact same good in real life.
(If you like a bar you'll accept higher prices because they fund the bar. If you like an artist you'll buy their merch over a different artist even if they're the "same thing".)
Of course this is true for things like grocery stores, but if you see someone doing this in real life you should assume they have a reason for it.
Related: https://astralcodexten.substack.com/p/how-do-ais-political-o...
The only limitation I see is, we now need so much more computing power.
Fine-tuning individual agents in order to move memories from the context window to the neural network weights, even if possible, would probably get too expensive.
> Facebook's CICERO artificial intelligence has achieved “human-level performance” in the board game Diplomacy, which is notable for the fact that’s a game built on human interaction, not moves and manoeuvres, like, say, chess.
It was the same scenario - many agents, multiple rounds, complex dialogue based interactions.