Don’t many/most state of the art models take many months to train on far more data than humans need for similar tasks?
Also, while e.g. GPT4 is quite capable across many tasks - humans seem to average towards learning robust _learning techniques_ themselves. Learning a new subject becomes easier thanks to somehow tracking and encoding learning strategies that are robust to learning other unrelated topics.
Humans generally need 18 years of pre training followed by 4-6 years of fine tuning before they can “one-shot” many difficult tasks. That’s way more training than any machine learning model I’m aware of.
Even for tasks like reading the newspaper and summarizing what you read, you probably had to train for 10-12 years.
A 3 year old has 3 years of multimodal training data and RLHF + a few billions of years of evolution that have primed and biased our visual and cognitive systems.
That requires a lot more data than machine models that literally zero inherent bias. Assuming you want a true apples to apples comparison.
If you want to be pedantic then only 6% of the human brain is the visual cortex. But AlexNet is also an inefficient model so something like an optimized ResNet is 100x as efficient to train. So now you're at 10.5kwh and 1.5kwh for the baby and model respectively.
You can argue details further but I'd say the energy cost of both is fairly close.
The networks we have are trained once and work efficiently for their training dataset. They are even robust to outliers in the distribution of that dataset. But they aren’t robust to surprises, unmentioned changes in assumptions/rules/patterns about the data.
Even reinforcement learning is still struggling with this as self-play effectively requires being able to run your dataset/simulation quickly enough to find new policies that work. Humans don’t even have time for that, much less access to a full running simulation of their environments. Although we do generate world models and there’s some work towards that I believe.
Again happy to be corrected.
If you want to be pedantic then 6% of the human brain is the visual cortex but then you also have to argue that AlexNet is horribly inefficient to train. So you cut the brain cost to 6% and the model cost to 1%. They're still within an order of magnitude (favoring the model) which I'd say is pretty close in terms of energy usage.
Humans are very bad at this task; it takes a massive effort to learn this many birds. In fact it's a great counterexample to human few shot learning ability...