It's interesting how they say that their model has high "zero-shot" performance, even though it was trained on millions of examples.
To me, zero-shot would mean a model whose parameters was never tuned through training examples at all...
To me, zero-shot would mean a model whose parameters was never tuned through training examples at all...
I can see how the machine learning definition is similar, but the whole idea of retraining and fine-tuning is foreign to a lot of subfields in GOFAI.
I guess attentional context is almost this, but LLMs don’t update their base model after a one shot inference.