Your LLM Is a Capable Regressor When Given In-Context Examples
arxiv.org
arxiv.org
Agree it both kind of makes sense (regression is the best way to predict the next token in this context) and kind of ironic (LLMs can do high school regression but can’t do elementary school long digit arithmetic).
I'd like to try it on a dataset like Titanic. That might be a more interesting experiment.
Or to avoid training data bias, maybe a completely random dataset generated by some more-sophisticated means, like the scikit-learn data generator which can introduce clusters, irrelevant features, etc.
I tried with linear regression with irrelevant features (e.g., NI 1/2 -> 1 informative variable, 2 total variables). The models still perform reasonably well.
All results can be seen in this heatmap: https://github.com/robertvacareanu/llm4regression/blob/main/...
I wonder if this paper will make it through peer review? Now that could be an interesting result.
The traditional OLS estimator in Excel has all sorts of classical optimality properties when its assumptions (normality, linearity) are true, so no fancy neural net can outperform it even in principle (the only way to outperform it would be to have an informative prior for how you generated the data set's parameters). So if the LLM's beat, or even matched, Excel in that case, they would be thinking too narrowly.
Same with any other models where we know the form of the answer up to some unknown parameters.
If we want parameter estimates, that means we already have a functional form in mind. In that case, we get to use established statistical theory to design optimal estimators, whether by Bayesian or other methods. Black box neural magic wouldn't help (or it might help indirectly, in computationally intractable cases).
What we would want the LLM's to do, ideally, is explore the space of known/possible 'patterns' and perform well in situations where the underlying relationship exists, is not known in advance, and is known not to have a simple form we can describe. Much like they (and we!) produce text without being able to describe why they are producing that particular text, we would expect them to make those predictions without being able to explain them in terms of parameters and functions - not without a whole other layer of explainability machinery.
Using 'established statistical theory to design optimal estimators' isn't trivial for most people. Black box magic might still be useful for them.
I've played with math examples (a while back, not with recent models) where they make errors, but seem to get the magnitude right, so perhaps easy to find the closest points (or roughly closest) to interpolate between.
Literal interpolation would be use a lot more often, except you can't practically do it in higher dimensions (even 2 isn't trivial).
What we are currently seeing in the comments are people trying random things and then saying “it doesn’t work”.
For example, for Friedman #1, GPT-4 predicts 12.89 while the true value is 11.69 (https://chat.openai.com/share/177571ad-3845-46a1-952f-963647...)
For Original #1, GPT-4 predicts 83.63 while the true value is 80.39 (https://chat.openai.com/share/808da995-99e6-444a-94da-fc7cd5...)
https://chat.openai.com/share/6217cd86-2b0f-41b2-a36a-2558dd...
`The task is to provide your best estimate for "Output". Please provide that and only that, without any additional text.`
Examples of this are available in Appendix J.
Sidenote, but all the experiments were ran with the API. There are some differences between the Chat and the API, for example the Chat can generate and execute code. I shared Chats since they are easy to look at and to try.
If you have access to an API key, I made some google colabs:
Colab links:
- GPT-4 Example: https://colab.research.google.com/drive/1Bk9uBCBvzuX00Rex-t1...
- GPT-4 Small Eval: https://colab.research.google.com/drive/1_-uHvW2oLtcCXz0c-G_...
- Claude 3 Opus Example: https://colab.research.google.com/drive/105jUAGanp7ZLG-Q9Hei...
- Claude 3 Opus Small Eval: https://colab.research.google.com/drive/1-IH68TUuqf_CZyptSSr...
Colab links:
- GPT-4 Example: https://colab.research.google.com/drive/1Bk9uBCBvzuX00Rex-t1...
- GPT-4 Small Eval: https://colab.research.google.com/drive/1_-uHvW2oLtcCXz0c-G_...
- Claude 3 Opus Example: https://colab.research.google.com/drive/105jUAGanp7ZLG-Q9Hei...
- Claude 3 Opus Small Eval: https://colab.research.google.com/drive/1-IH68TUuqf_CZyptSSr...
I wonder if testing a single point is masking a larger llm error.
I also added google colab examples:
A small example with GPT-4: https://colab.research.google.com/drive/1Bk9uBCBvzuX00Rex-t1...
A small scale eval with GPT-4: https://colab.research.google.com/drive/1_-uHvW2oLtcCXz0c-G_...
Also, feel free to DM or email me if you’d like to chat more about it!
https://arxiv.org/abs/1808.00508
They can interpolate but extrapolation is hard because of the constraints from nonlinearities
For example, slightly adapting the code from the `README.md` for re-creating a barplot similar to Figure 1:
``` print(cdf.groupby(by=['dataset', 'model']).apply(lambda x: np.sqrt(((x['pred'] - x['gold']) ** 2).mean()))) ```
The code above prints (the dataset is Original #3):
```
dataset model
original3 Claude 3 Opus 11.862656
DBRX 32.186706
GPT-4 31.672733
Gradient Boosting 21.985396
KNN 55.043003
Linear Regression 29.987077
Random Forest 31.688074
```Colab links:
- GPT-4 Example: https://colab.research.google.com/drive/1Bk9uBCBvzuX00Rex-t1...
- GPT-4 Small Eval: https://colab.research.google.com/drive/1_-uHvW2oLtcCXz0c-G_...
- Claude 3 Opus Example: https://colab.research.google.com/drive/105jUAGanp7ZLG-Q9Hei...
- Claude 3 Opus Small Eval: https://colab.research.google.com/drive/1-IH68TUuqf_CZyptSSr...