> I will address your central point: "LLMs are not capable of original work because they are essentially averaging functions." Mathematically, this is false: LLMs are not computing averages.
For someone who thinks "Gotcha" debating is pointless and tiring, you sure read where I said: "[J]ust because you use a program to take the average of billions of integers which you can't possibly understand all of, doesn't mean you don't understand what the "average()" function does" ...and thought "Gotcha!" and responded to that, without reading the very next sentence where I said: "Obviously LLMs are a lot more complex than an average, but they aren't beyond human understanding."
If I had to succinctly describe my core point it would be:
We understand all the inputs (training data + prompts) and we understand all the code that transforms those inputs into the outputs (responses), therefore we understand how this works.
> We have no idea how GPT is able to translate
It is able to translate because there are massive amounts of translations in its training data.
> or how its able to follow prompts given in another language
Because there are massive amounts of text in that language in its training data.
> why GPT-like models start learning how to do arithmetic at certain sizes
I'm pretty sure that isn't actually proven. I doubt that it's not primarily a function of size, but rather a function of what's in the training data. If you train the model on a dataset which doesn't contain enough arithmetic for the GPT model to learn arithmetic, it won't learn arithmetic. More data generally means more arithmetic data (an absolute value, not a percentage), so a larger dataset gives it enough arithmetic data to establish a matchable pattern in the model. But it's likely that if you, for example, filtered your training data to get only the arithmetic data and then used that to train the model, you could get a GPT-like model to do arithmetic with a much smaller dataset.
I say "primarily" because the definition of "arithmetic data" is pretty difficult to pin down. Textual data which doesn't contain literal numerical digits, for example, will likely contain some poorly-represented arithmetic i.e. "one and one is two" sort of stuff that has all the potential meanings of "and" and "is" mucking up the data. A dataset might have to be orders of magnitude larger if this is the sort of arithmetic data it contains.
In each of these cases, there are certainly some answers we (you and I) don't have because we don't have the training data or the computing power to ask. For example, if we wanted to know what kind of data teaches the LLM arithmetic most effectively, we'd have to acquire a bunch of data and train a bunch of models and then test their performance. But that's a far cry from "We have no idea". Given what we know about how the programs work and what the input data is, we can reason very effectively about how the program will behave even without access to the data. And given that some people do have access to the training data, the idea that we (all humans) can't understand this, is very much not in evidence.
> I suggest you look at the paper I linked about emergent capabilities without trying to nitpick aspects that can be used to argue against my point.
I had read the paper before you linked it, and did not think it was a particularly well-written paper, because of the criticism I posted earlier.
I think using the phrase "emergent capabilities" to describe "An ability is emergent if it is not present in smaller models but is present in larger models." is a poor way to communicate that idea, which has been seized upon by media and misunderstood to mean that something unexpected has occurred. If you understand how LLMs work, then you know that larger datasets produce more capabilities. That's not unexpected at all: it is blindingly obvious that more training data results in a better-trained model. They spent a lot of time justifying the phrase "emergent capabilities" in the article, likely because they knew that the public would seize upon the phrase and misinterpret it.
If you don't believe I read the paper, you'll note that my doubt that GPT's ability to do arithmetic actually is a function of size originally actually came from the paper, which notes "We made the point [...] that scale is not the only factor in emergence[.]"
There's a separate issue with what you've said in this conversation. It seems that you believe we don't understand the models, because we didn't produce the weights. Is that an accurate representation of your belief? Note that I'm asking, not telling you what you believe: please don't tell me what my central point is again.