Human training begins at birth.
Evolution might result in better architecture and init(inductive biases), but that's a separate thing than training.
Human training begins at birth.
Evolution might result in better architecture and init(inductive biases), but that's a separate thing than training.
No reason beyond compute we couldn't do something similar. Ie find good architectures by evaluating them using multiple random weights, and evolve those archigectures that on average gives the best results.
Then over time add a short training step before evaluating.
Is this true? My understanding is that people are born with many pre trained weights. Was the evolutionary convergence of those weights not itself training?
No, inductive biases are not training.
I'm saying that better models(ie: better inductive biases) or non-language data is needed to advance LLMs and somehow we've arrived at "evolution is training." I'm not sure how that's relevant to the point.
And the more inductive bias we've shoved into models, the worse they've performed. Transformers have a lot less bias than either RNNs or CNNs and are better for it. Same story with what preceded both.