E.g., https://openai.com/blog/evolution-strategies/
One of the remarkable results is that convergence rate for each parameter is not strongly dependent on the number of parameters.
With the neural networks we train, there is actually a highly engineered process required to create the state in which a network can be trained and inferences be driven through it. The necessary combination of hardware that is able to perform and persist operations on information and the algorithms required to do so in a way to yield this outcome is an extremely complex set of pre-conditions that we wouldn't expect to find in the computing equivalent of a primordial soup.
With natural evolution, there is no obvious agency or intent behind it. Who is there to care whether or not life started on Earth, and/or who is driving the laws of nature such that the constructive, generative process of genetic evolution actually 'works' as well as it does? Seemingly nobody. Yet this process is able to create systems that operate on scales that we can only dream of. Look up YouTube videos on ATP Synthase for example. It's a nanomachine in every sense of the word. It uses the proton equivalent of a water wheel to spin a little machine that grabs a molecule of ADP, a molecule of inorganic phosphate, then literally snaps them together with mechanical leverage to make ATP. This little miracle machine that powers most of life on earth was built in literal and figurative darkness...it's so damn small light can't see it, and there was nobody there to appreciate its beauty until we came along billions of years later.
Ultimately I'm not surprised natural evolution works, I'm surprised at its speed and efficacy.
As reproduction produces different variants of the same organism, some variations help the organism while others do not. Organisms with the helpful mutations will be more likely to pass those onto their offspring. Organisms with detrimental variations will be less likely to pass those variations to their kids.
It’s not a precise process like gradient descent but when there are billions (trillions?) of organisms evoking simultaneously and independently, it makes more sense how the complexity of biology has come about.
Similarly the details of the evolutionary 'training' also tend not to matter much, just about any algorithm that prefers better instances (by whatever metric) with a very slightly higher probability will converge after a reasonable number of generations.
Exponential processes are always surprising. If you have a trait or parameter than confers only a 1% chance of helping survival, it will have an effect of (1.01)^100 =2.7x after only 100 generation. After 1000 generations the effect is 21,000x.
Branches are retried, so if some branch is an improvement and is cut by bad luck, it may be luckier a few million years later.
This does not guarantee that the "best" solution is found, but also avoid looking in the 4^300 combinations. Also, there are shorter genes, some proteins have ~50 amino acids (~150 bases), and some useful short amino acids chains have a length of 20 or even less (~60 bases or less). It's possible to start with a short versions that does something slightly useful, and slowly increase the length an efficiency.
This book is even better because it actually talks a lot of about horizontal gene transfer s role in prokaryotic evolution, and borgs might actually be something involved in a similar process. Kinda prescient if you ask me!
1. Parallelism. The number of all sorts of organisms going through mutations is large. Like 10^40 kind of large.
2. Time. This has been going on at a very rapid tempo for quite a while.
3. Evolutionary pressure. In every generation, harmful mutations are radically weeded out so your search space is dramatically reduced at each generation.
While the first two could be roughly estimated, the third one involves non-linearity very sensitive to estimation errors. So I don’t think anybody can _prove_ this is how we ended up with Angela Merkel but it’s not implausible either and nobody has a better idea
1. Genes don't have to be optimal, or even close to optimal to work. They just have to be good enough. 2. Nature loves to copy. Large segments of DNA can be copied by a number of mechanisms and randomly placed elsewhere in a genome. So once nature "discovers" (for example) a DNA-binding motif, that motif can be added to other genes, and now you have a diverse set of DNA-binding proteins, which will continue evolving on their own.
The book's about how nature manages to actually explore the vast, vast genetic space and harvest its bounties cumulatively, while under the constrain that every "step" must be a viable organism with offspring.