> I feel like extensive pretraining goes against the spirit of generality.
What do you mean by generality?
Pretraining is fine. It is even fine for pursuit to AGI. Humans and every animal has "baked in" memory. You're born knowing how to breath and have latent fears (chickens and hawks).
Generalization is the ability to learn on a subset of something and then adapt to the entire (or a much larger portion) of the superset. It's always been that way. Humans do this, right? You learn some addition, subtraction, multiplication, and division and then you can do novel problems you've never seen before that are extremely different. We are extremely general here because we've learned the causal rule set. It isn't just memorization. This is also true for things like physics, and is literally the point of science. Causality is baked into scientific learning. Of course, it is problematic when someone learns a little bit of something and thinks they know way more about it, but unfortunately ego is quite common.
But also, I'm a bit with you. At least with what I think you're getting at. These LLMs are difficult to evaluate because we have no idea what they're trained on and you can't really know what is new, what is a slight variation from the training, and this is even more difficult considering the number of dimensions involved (meaning things may be nearly identical in latent space though they don't appear so to us humans).
I think there's still a lot of ML/AI research that can and SHOULD be done at smaller scales. We should be studying more about this adaptive learning and not just in the RL setting. One major gripe I have with the current research environment is that we are not very scientific when designing experiments. They are highly benchmark/data-set-evaluation focused. Evaluation needs to go far beyond test cases. I'll keep posting this video of Dyson recounting his work being rejected by Fermi[0][1]. You have to have a good "model". What I do not see happening in ML papers is proper variable isolation and evaluation based on this: i.e. hypothesis testing. Most papers I see are not providing substantial evidence for their claims. It may look like this, but the devil is always in the details. When doing extensive hyper-parameter tuning it becomes very difficult to determine if the effect is something you've done via architectural changes, change in the data, change in training techniques, or change in hyperparameters. To do a proper evaluation would require huge ablations with hold-one-out style scores reported. This is obviously too expensive, but the reason it gets messy is because there's a concentration on getting good scores on whatever evaluation dataset is popular. But you can show a method's utility without beating others! This is a huge thing many don't understand. Worse, by changing hyper-parameters to optimize for the test-set result, you are doing information leakage. Anything you change based on the result of the evaluation set, is, by definition, information leakage. We can get into the nitty gritty to prove why this is, but this is just common practice these days. It is the de facto method and yes, I'm bitter about this. (a former physicist who came over to ML because I loved the math and Asimov books)
[0] https://www.youtube.com/watch?v=hV41QEKiMlM
[1] I'd also like to point out that Dyson notes that the work was still __published__. Why? Because it still provides insights and the results are useful to people even if the conclusions are bad. Modern publishing seems to be more focused on novelty and is highly misaligned from scientific progress. Even repetition helps. It is information gain. But repetition results in lower information gain with each iteration. You can't determine correctness by reading a paper, you can only determine correctness by repetition. That's the point of science, right? As I said above? Even negative results are information gain! (sorry, I can rant about this a lot)