Did the training data incorporate how popular (e.g. likes or upvotes) each sample was as a proxy for quality? Or can you achieve this performance just by looking at averages on a large enough data set?
Did the training data incorporate how popular (e.g. likes or upvotes) each sample was as a proxy for quality? Or can you achieve this performance just by looking at averages on a large enough data set?
See also https://en.wikipedia.org/wiki/Attention_schema_theory (unrelated to the notion of "attention" used in ML)
Where is this massive repository of well structured code with good clear variable names that it’s tapping into?
That's one of the issues with these models, we say they produce "good" output but really they're producing output that is "good" from one specific point of view that happens to be expressed in code and introduces a large bias into their outputs.
If it was all a computer program it’d be acting like ELIZA.
Just because generative text models in the past(like ELIZA) were bad doesn't mean that the algorithms we have now are much more than better versions of the same.
The results were beyond impressive. I don't know how it does it, but it does it really well.