Evolution Is the New Deep Learning
sentient.ai
sentient.ai
1) One of the biggest reasons they fell out of favor for more "mathematical" approaches was that no one could really explain why exactly they worked. It makes sense on the surface that "survival of the fittest" and doing something akin to multiple stochastic gradient descents would work, but no one has really been able to produce a mathematical proof as to why.
Since other folks are producing good examples of "explainable AI", I don't know how Genetic Algorithms/programming could be made 'explainable' as to why they achieved an optimal solution other than hand-waving to how evolution works in nature.
2) The most important thing to define is the fitness function, this defines what the search space looks like and how easily a globally optimal solution can be derived. For a good example of an interesting search space that a genetic program would have a difficult time with, see Schwefel functions [0]. Back when I researched these things closely, my intuition was that reality rarely fits neatly into good fitness functions and I felt that at the point you are understanding the problem, you may just be better off with a direct approach, which leads to
3) Genetic programming should only really be considered when there are no known alternatives or they are way too computationally expensive.
In either case, I would welcome a resurgence in a topic I once knew quite well, though I haven't been in that field for a few years now.
[0] https://jamesmccaffrey.files.wordpress.com/2011/12/schwefels...
My impression is that they fell out of favor precisely because don't actually use any gradients, and end up converging on good maxima slower than you could if you used the gradients from the net. Am I off base here?
If your mutations aren't small, or your parameters are not continuously valued, or your fitness function is hard to differentiate analytically, genetic algorithms might still come out ahead.
For "find" you could discuss convergence rates vs. the points that are being converged to. If you add randomness at every iteration you're not even converging, at which point annealing rate becomes a related issue.
For "global optima" are we talking about training error, or test error, or some other kind of function value?
I'm not up to date on theoretical research about this topic, but as far as I recall there are some interesting demonstrations on realistic problems showing that all the different "local" optima resulting from different random initializations are actually all "connected", i.e. there exists a nondecreasing route how you can get from a worse "local optimum" to the better one, so in reality it's not a local optimum, it's just that it's nontrivial to find the path to a better optimum if the space is very highdimentional.
1. It's hard to get stuck in multidimensional space
2. There are more saddles than convex local optima
3. There are many local optima, but they are all useful
4. Something related to spin glass theory (which I don't understand)
5. There is no theory, or we haven't found it yet; all we
know is that it works in practice and the theory will have to catch up laterFurther, the mindboggling size of the high-dimensional spaces make me all but guaranteed that not a single non-trivial neural network made by homo sapiens has ever been in global maxima. But I have nothing but my hunch on this.
Consider a neural net that produces a single output given `N` inputs– it's basically a function `f(x) = f(x_1, x_2, …, x_N)` A local minimum x* has the gradient `∇f(x) = (0, 0, …, 0)` and an NxN Hessian of the form `[H•f(x)]_{ij} = (∂^2 f)/(∂x_i ∂x_j)`.
The critical point is a local maximum if the Hessian is positive definite at x and a local minimum if it's negative definite; it's a saddle point otherwise. This corresponds to the Hessian having a mixture of positive and negative eigenvalues.
Heuristically, we have N eigenvalues, and absent prior information we can expect that the probability of each eigenvalue being positive is 1/2, and 1/2 for the eigenvalue being negative instead. So the probability of an arbitrary critical point being a local minimum is `(1/2)^N`, which becomes extremely small as N grows large. Large values of N are kinda the bread-and-butter of deep learning, so we expect that most critical points we'll encounter are saddle points. Additionally, the inherent noise in SGD means that we're unlikely to stay trapped at a saddle point, because once you're nudged away from the saddle, the rate at which you slide off it increases rapidly.
So if your net appears to have converged, it's probably at a local optimum with a reasonably deep basin of attraction, assuming you're using the bag of tricks we've accumulated in the last decade (stuff like adding a bit of noise and randomizing the order in which the training data is presented).
As for 3 & 5, we kinda cheat because if your model appears to have converged but is not performing adequately, we step outside of the learning algorithm and modify the structure of the neural net or tune the hyperparameters. I don't know if we'll ever have a general theory that explains why neural nets seem to work so well, because you'd have to characterize both the possible tasks and the possible models, but perhaps we'll gradually chip away at the problem as we gather more empirical data and formalize heuristics based on those findings.
I'm not even sure that it's not a problem in general. I know I've watched NNs frequently get stuck in local minima even on incredibly simple datasets like xor or spirals. SGD and dropout are widely used in part because they add noise to the gradients that can help break out of local optima. But that's not a perfect method
Reinforcement learning was on my mind when I was writing about "practical supervised learning applications" because yes, RL is different in that regard. And various function calculation examples (starting with XOR) indeed do so.
However, if we're applying neural networks for the (wide and practically important!) class of "pattern recognition" tasks like processing image or language data, then it's different, and those are full fields where you can easily spend a whole career working on just one of these types of data. Perhaps there's a relation with the structure and redundancy inherent in data like this.
The reason it's not done so much is because the bandwidth of moving huge numbers of gradients or weights between computers is pretty significant. There's been all sorts of research into compressing them or reducing the precision. However this is a problem for evolutionary algorithms as well.
Still, even more complex systems still do not require particularly high bandwidth. You only need to broadcast the fitness of the individuals, and from there each node can independently recalculate and combine the best ones.
Why did they fall out of favour? Fashion, perhaps a natural break in progress - hit a wall and couldn't get further. But there is plenty of GA research going on.
Tierra obviously counts, but I was thinking of more specific examples. Say, something like this paper, which showed that adding a tiny cost-function to a network spontaneously makes it more modular:
[0] http://rspb.royalsocietypublishing.org/content/280/1755/2012...
http://users.sussex.ac.uk/~lionelb/downloads/EASy/publicatio...
1. Evolvability is Inevitable: http://journals.plos.org/plosone/article?id=10.1371/journal....
2. Extinction Events can Accelerate Evolution (2015): http://journals.plos.org/plosone/article?id=10.1371/journal....
3. Evolvability Search: Directly selecting for evolvability in order to study and produce it (2016): http://www.evolvingai.org/mengistu-lehman-clune-2016-evolvab...
You are also oversimplifying by not distinguishing local and global success. Locally, this gene will be successful. Globally, some mutations will be beneficial, and the children with those mutations will be more successful.
The gene may become "dominant" (and then gain the beneficial mutations through sexual reproduction), but it will not achieve a monopoly and not remove the existence of evolution.
It is somewhat comparable to why some people are left-handed[0].
However, what you allude to is basically the necessity of evolution of forms of "error correction" against mutations for complicated organisms. Sex plays a large part in that as well[1]. And interestingly, DNA repair, another involved mechanism, may actually help evolvability[2].
[0] https://www.youtube.com/watch?v=TGLYcYCm2FM
[1] https://www.quantamagazine.org/missing-mutations-suggest-a-r...
[2] https://www.quantamagazine.org/beating-the-odds-for-lucky-mu...
>it will not achieve a monopoly and not remove the existence of evolution.
Yes it will. It's easy to do simulations. Statistically the children of the organism with less mutations will have an advantage. Eventually the gene will reach 100% of the population. The population will stop evolving and eventually go extinct.
> It's not hypothetical.
> We haven't observed it in nature
Also, the fact that in the long term such a mutation would cause a population to go extinct eventually is not really evidence that we should not find this in the wild, because it would still be effective in the short term. Look at that recent story about the mutated lobster taking over the waters of Germany for an example.
Anyway, the crucial disagreement lies here:
> It's easy to do simulations.
Yes, if a gene would evolve that brings the mutation rate to zero, this simulation works. But achieving such a thing sounds a lot like beating the laws of thermodynamics and stopping entropy from increasing. And sure, reducing entropy locally is possible by globally increasing it - in this case that would mean increasing the energy/resource budget spent on it, but I suspect that this in itself will be at a cost too high to give an advantage.
I basically don't believe (and I fully admit that this is a subjective point of view) that the "easy simulations" are noisy, large or complex enough to reflect the messy biological reality here. Essentially, there's too much approximation going on.
And btw, evolving to extinction is not the same thing: a gene that makes the entire population male is not in itself advantageous like reducing mutations is. It is more generic than the "mutation rate zero"-gene scenario, the latter is basically a specific hypothetical example of it.
1) "explainable AI" is better in GA than in deep learning. GA gives you a structure that works and probably easier to understand than any Deep Learning model (which it is a huge math function). There are so many things that also doesn't make sense why they work in deep learning but we still use them, that is the same with GA.
2) Knowing the fitness function doesn't mean you can solve the problem. When the search space is so big you need something to search on it, and there is where GA can shine. It is also the same with Deep Learning and mostly evolutionary computation. The search space is so huge for a "brute force algorithm". You need heuristics and GA works well for some problems, the same way gradient descent works well for others too.
Genetic algorithms work on problems where some subsets of variables are approximately separable (uncorrelated) from some other subsets of variables.
F(a,b,c,d,e,f) ≈ F1(a,b,c) × F2(d,e,f)
So if you have found a good combination or a,b,c it makes sense to try how it works it any promising combinations of d,e,f. Some natural world problems really have this property.
http://citeseerx.ist.psu.edu/viewdoc/download?doi=10.1.1.394...
(There's an earlier discussion about the myth of local optima here that might or might not answer my question.)
Kind of like how nobody can really explain how the brain works, or life in general. My gut feeling is that it is hubris to think that we are going to "figure out" intelligence with increasingly sophisticated mathematical models anytime soon. We are not giving proper credit to how complex it is, and the multi-billion year developmental process that it took. We think we can just short-circuit that with some fancy math because we've had success with planetary orbits and other comparatively rudimentary phenomena.
The current industry approaches are great for extracting certain kinds of value out of large data sets, but in terms of producing a result that could even begin to be considered as interesting as life (i.e. AGI or "strong AI"), I believe we will have to rely on creating a system whose inner workings are too complex for us to understand.
In other words, going off of Arthur C Clarke's definition, life is magic. And we're trying to create something equally magical. Almost by definition, if we can analytically understand it, it's not going to be interesting enough.
Or we are simply not ready to accept that it's simply a big book of heuristics fine-tuned over biological eons.
It's just big. We have too many interwoven, interdependent, synergistic faculties. Input, output, and a lot of mental stuff for making the right connections between the ins and the outs. Theory of mind, basic reasoning, the whole limbic system (emotions, basic behavior, dopaminergic motivaton), the executive functions in the prefrontal cortex, all are very specialized things, and we have a laundry list of those, all fine-tuned for each other.
And there's no big magic. Nothing to "understand", no closed formula for consciousness. It's simply a faculty that makes the "all's good, you're conscious" light go green, and it's easy to do that after all the other stuff are working well that does the heavy lifting to make sense of reality.
We can't even create life from non-life. How can we begin to understand all the stuff you're talking about that's been layered on top? We don't understand this stuff well enough to just handwave it away as unimportant or trivial.
We can't really define life in the first place.
But this is not necessarily to your favour. I think it's more of an indication of how the world doesn't fit into our... anthropomorphic way of thinking. That is, everything follows the laws of physics, no magic involved. We aren't special.
Even so, we understand it just the same no matter how you define it, because what we understand is not a function of word choice or definition. It's a function of capability.
> Clearly it's not that simple or easy, or we would have done it.
We don't have the computational power yet. Not to mention the vast amount of development required. Think of the climate models, that are huge (millions of lines of code), but they're still nowhere near complete enough, and they only have to model sunlight (Earth's rotation, orbital position, albedo), clouds, flows (winds and currents), some topography (big mountains, big flats), ice (melting, freezing), some chemistry (CO2, salts). And they only have to match a simple graph, not the behavior of a human mind (eg Turing test).
So, it's not easy, even if simple.
> We can't even create life from non-life.
We understand life. Cells, RNA, DNA, proteins, mitochondria, actins, etc. It's big, it's a lot of moving parts, and we understand it, but we can't just pop a big chunk of matter into an atomic assembler and make a cell.
And I think intelligence/sentience is similar. It's big, not magic.
We understand many small and big things about life, yes.
We have little or no understanding of the complex protocols that occur within a cell. If we did, our standard manufacturing techniques would be vastly different.
We can modify DNA and RNA in interesting ways, but they are not living. It is not until we put them into an already existing living cell that we can reprogram some characteristics of that cell.
We have pretty fine understanding of cells, but our materials science and manufacturing technology is not "vastly parallel incremental molecular", but "big precise drastic pure chunk" based compared to cellular manufacturing. Not to mention protein folding and self-assembling biomachines and so on. We're getting there.
No predictive power whatsoever.
Also, I'm amazed by consciousness, by reasoning, by our cognition, by intelligence, how we apply it, day-to-day, from pure math to messy, but useful engineering, through the ugliness of realpolitik, and the beautiful and dreadful human tangle that our civilization is. The contrasts, the why-s. (Consider the so stark difference between US and Mexico especially the border towns, which is of course deceiving, as the problems don't stop at the border, the cities, states, nations are connected. The warring cartels, the corruption, the hopeless have-nots, the dealers, the addicts, the war on drugs/terror/smuggling/slavery/yaddayadda, the DEA, ATF, their foreign counterparts, the policy going against the market, the hard on crime ideology, the big data vs gerrymandering case just on the Supreme Court's plate, the pure math and reasoning behind all that again, are all connected, just harder to frame them in a "deep" picture.)
But so far, none involves any actual irreducible complexity. No magical formula, just layers upon layers of complexity and fine-tuning.
Certainly you do realise that this has been a moving goalpost for half a century? It seems that lately people started to avoid giving a concrete estimate of the power required, though. It was so easy in the 90-s! «Human visual/verbal system processes gygabytes per second/has n flops to the xth» or the like. Well, now we have that and more; how come a deeper modeling, a finer processing, a more complicated network has come to be needed?
And your examples are incorrect, for example Navier-Stokes equations plus some general physics knowledge always have allowed us to estimate how much data we need for a certain fidelity of a finite-term weather forecast. Certainly we need more for a complete climate model, but we know what we need. No such thing about the brain.
It's an easy way to score some rationality points by voicing rejection of "magic", but it's a strawman. Nobody will bother arguing for a mythical homunculus in the seat of the soul, nor even for a concise formula summing up the workings of the mind. Pick harder targets. "It's just big" or "it's just a bunch of heuristics cobbled together" is a non-explanation. The brain is not a Rube Goldberg machine that manages to produce any sort of work simply due to its excessive complexity – it is energetically economical, taking into account that neurons are living cells that need to sustain their metabolism and not merely "compute" when provided with energy. Its discrete elements aren't really small by today's standards, nor are they fast. The number of synapses is ridiculous, but since they aren't independent, at a glance it doesn't add that much complexity too (unless we abandon reason and emulate everything close to the physical level).
Yet we have failed to realistically emulate a worm. By all accounts we have enough power for 302 neurons already. There's no workload to give to overwhelm available supercomputers. It's knowledge and understanding that we lack, and it's high time to give up on the delusion that more power, naturally coming in the future, will somehow enable a creation of predictive brain model, for this would truly be magic.
We're getting pretty good at computer vision, what's lacking is the backend for reasoning, for generating the distributions for object segmentation and scene interpretation. Basically the supervisor. (As unsupervised learning is of course just means that the supervision and goal/utility functions are external/exogenous to the ML system, such as natural selection in case of evolution.)
My example illustrates that yes, we can give an upper bound on molecule by molecule climate modeling, but that's just a large exponential number, not interesting, what we're interested in is useful approximations, which are polynomial, but they being models, they need a lot of special treatment for the edge cases. (Literally the edges of homogeneous structures, like ice-water-air, water-air, water-land, air-land [mountains, big flats, etc] interfaces. And the second order induced effects, like currents, and so on.) That means precise measurements of these effects, and modelling them. (Which would be needed anyway, even if we were to do a back to the basics N-S hydrodynamics model, as there are a lot of parameters to fine-tune.)
For the brain we know the number of neurons, the firing activity, the bandwidth of signals, etc. We can estimate the upper limit in information terms, no biggie, but that doesn't get us [much] closer for the requirements of a realistic implementation.
> Yet we have failed to realistically emulate a worm.
http://openworm.org/getting_started.html#goal seems to be matter of time, not lack of understanding. ( https://github.com/openworm/OpenWorm#quickstart ) But maybe I'm not up to date on the issues.
> it's high time to give up on the delusion that more power, naturally coming in the future, will somehow enable a creation of predictive brain model, for this would truly be magic.
a) people are saying exactly this for years, that we have enough data already, we need better theories/models
b) they fail to accept that more computing power and data is the way to test and generate theories.
> The brain is not a Rube Goldberg machine that manages to produce any sort of work simply due to its excessive complexity
A Rube Goldberg machine is simple, just has a lot of simple failure modes. (A trigger fails to trigger the next part, either because the part itself fails, or the interface between parts failed.)
> Its discrete elements aren't really small by today's standards,
If you mean cells, or cortices, agreed.
If you mean functional cognitive constituents, I also agree, but a bit disagree, as they are small parts of a big mind, all interwoven, influencing, inhibiting, motivating, restricting, reinforcing, calibrating, guiding, enhancing each other to certain degrees.
So in that sense consciousness is a big matrix which gives the coefficients for the coupling "constants" between parts. A magical formula if you will. But not more magical, than the SM of physics.
Have you not been following the work of Craig Venter? Depending on your point of view, he's already done it. Even if you don't agree, you have to admit that he's probably one of the few closest to actually doing it.
If you mean bigger biology, yes, sure, we don't have a full map of functional genomics for humans, but we're getting there.
Or maybe not, maybe it's so exponentially more complex, that it'd take as much time to understand it as it took for evolution to work it out. (Especially considering that evolution played with every individual, whereas we like to constrain our data gathering to non-aggressive methods.)
The connectome is necessary, but far from sufficient, to 'understand' a brain, even one made from only 302 neurons as in C. elegans.
The rules that govern a system can create patterns, which themselves behave according to rules, but with a set of rules that was "hard to predict" from the underlying system.
I put the example of the metabolic pathways because last time checked (~2015) the most advanced things in the field were extremely simple and without any predictive power. Things like calculating the kernel of a stoichiometric matrix or the centrality of a node in the interactomic graph.
None of that means we don't understand the principles. I'd say it's pretty much like fusion. Yes, we know how the Sun works, but putting it into a bottle is a bit of a pickle, similarly with brains. (Except brains have a lot more complexity.)
It is not a question of how there is emergence, why there is magic. The answer to that is is because systematic interactions at a low level can create higher level playing fields.
So the "how" is now a technical question, what is this system, how complex is it, and at which levels can we understand it. And since this system has been learning how to avoid erasure by entropy or by competition for 3.5 billion years, it has searched quite a possibility space, namely 2^1277500000000, if we assume making a copy every day.
I don't think that's a strong claim or that it even qualifies as a claim at all. Lots of things might decompose into simple components if subjected to the right analysis, very few things definitely won't - for example many clever people have spent a great deal of time attempting to reduce quantum and cosmic scale physics to simple intuitively founded laws... If Dennett's claim is that the human brain is the same order of object as the universe I can accept it only if we agree that all objects share the same order. Where does that get us?
[1] Apologies, I don't know what reductible means, but guessed typo - I'm open to education though and unworried by typos!
Now we have data and people for some reason want to claim that a theory with magical super complex and not-even-yet-describable and very-very-irreducible element(s) is a better fit than a good old box full of tiny yet specialized parts fine-tuned to work together over millions of years.
So in that sense the experiment is to enumerate the basic (built-in) functional components of the mind and corresponding implementational level machinery, and the of course the reverse (try to enumerate the implementation components and match them with functions) can generate important data (is there a function that has no implementation?).
That said, since the claim is that there's no magical component in the mind, and that's kind of hard to prove, but easily falsifiable, just find a/the magical component.
The problem is the same as with the soul, and the self, and so on.
I would expect this to yield new insights. So at a minimum, I’d think we can learn more now than we have been able to before. Maybe that will lead to a great increase in understanding of intelligence or a slight one, but it will lead to an increase in our understanding of the phenomena.
Computer simulations can give insight into simple systems like water flows, etc. Simulating more complex systems like a single living cells or any system built on livings cells would require systems that would not really be worth building. It would be simpler to use cells directly.
We did it already. Compter understand language, translate it, react to it. They can recognize items on a picture. Is there a task left which can't be done by computers better and faster than by humans?
>Almost by definition, if we can analytically understand it, it's not going to be interesting enough.
I think current ML is magic. I understand the math behind it. But still, I'm amazed every time when the training is over and it actually works like intended. Everything which is big enough is more than the sum of it's parts.
Are you serious? You think we're done?
What about when it doesn't work as intended and fails ridiculously, even though it usually works perfectly well?
http://www.labsix.org/physical-objects-that-fool-neural-nets...
Imagine some sort of Robocop deciding to neutralize someone for holding a turtle toy.
Your comment sounds like Lord Kelvin proclaiming that physics is "over" a couple years before people figured out there were huge holes in the theory which eventually led to quantum mechanics. Our understanding of intelligence is probably less complete than our understanding of physics _back then_.
2. ML is really underwhelming if you measure it against actually intelligent behavior. Figuring out cool regression mechanisms is neat and all, but that's what it is, and it has nothing to do with intelligence in the general sense, much like expert systems had nothing to do with actual domain knowledge, they were just one of the most primitive models, low-hanging fruit that we could exploit.
All the tasks humans still earn money doing. And given that we're nowhere near full automation, I'd say it's quite a few tasks.
Industry/Economy lacks behind state-of-the-art technology by decades.
A better example is plumbing. How would you go about automating a human plumber who handles all sorts of piping and crawl spaces in a large variety of settings?
Sure, plumbing is a much more challenging example. But it is an economic problem. The cost to automate plumbing is much higher than the utility of it.
I assume you're referring to his "any sufficiently advanced technology is indistinguishable from magic"? If so, you're misrepresenting it, because he's clearly saying it's not magic, it just appears that way to the unadvanced. And there's a big difference between "appears to be" and "is".
Also 'appears to be' != 'indistinguishable', the latter is far closer to 'is' imo.
Speak for yourself.
I agree that when there's a fast, perfect solution, it doesn't make sense to use genetic algorithms. But when finding solutions to a non-general problem (optimize CNC tooling to produce a list of orders, each of which has a series of operations that require a certain amount of time, on certain machines, and require being moved from machine to machine, such that you produce the most on-time orders for high-priority clients), genetic algorithms can work very well.
That is the first time I have heard that claim, and since we have a large body of knowledge describing how evolution works (that sampo description is one the clearest I've seen) and how it can be optimized, I imagine you are talking about some other problem.
Is it about predicting the causes of some learned trait? Is there some interesting research on that?
Even more because it's clearly not a property of genetic algorithms in general, but a very powerful effect that one aims into achieving with genetic algorithms and a good domain modeling. I don't really understand what is the meaning of something like that being impossible to prove.
So I was surprised to see them make a return about a decade later. Hopefully there is a little more rigor this time around.
NFL theorems are, should I say, purely theoretical and provide no insight on real-world problems. Say we try to find a function that is an optimal solution to something. NFL theorems consider the space of all possible functions, the overwhelming majority of which are discontinuous. Whereas real life problems tend to have functions that are at least more or less continuous.
Honestly, there's no justification to be using NFL theorems to explain why we can't optimize well on real world tasks.
Edit: And such high Kolmogorov complexity function constitute most possible objective functions -- i.e. exponentially more than the number of objective functions that don't have high Kolmogorov complexity. And all real world objective functions have comparatively low Kolmogorov complexity.
Yeah. You have a function, so basically a long array of numbers, and you want to find the maximum. If the data in the array has some structure, like it's sampled from a sine wave or something, you can use some strategies to find the maximum. Like gradient descent, or binary search. Something.
But if the array is filled with random numbers, looking at other arrays elements give absolutely no hint on what might be in an array element you haven't yet looked at. So there doesn't exist any more efficient strategies to find the maximum number, than linear or random search.
And the space of all possible functions mostly consists of discontinuous functions that are, for all purposes, just samples of random noise.
This is all NFL theorems say. I really don't understand how they got be such a big deal.
For classic old school computers yes. I'm not so sure about quantum computers. Consider: https://en.wikipedia.org/wiki/Grover%27s_algorithm
I mean, yeah, sure, not knowing if we've hit the global minimum on some optimization problem may not matter as humans, because no one gives a damn, but that is a completely arbitrary line, not a mathematical one. Too often I hear people say "well, NFL is irrelevant" as an argument for why they are mathematically correct, or why their results are mathematically significant, and that's simply a load of crap. Maybe you're right, maybe you're wrong, but you're basically just throwing darts at a dartboard.
Maybe that is how it looked from the outside. But the field continued as normal with no major problems, solving lots of industry problems with little hype.
Say you want to try many various changes like the title of your page, the color of the background, the position of your buy button etc. We solve this problem by trying out random variations of these websites - like A/B testing with more candidates - and by crossing the best performing ones to create a new generation of websites. This helps us find good performing variations in a very big search space.
This would be hard to do with deep learning as we start with no data at all, measuring the performance is quite noisy and you can't compute a gradient to know how to evolve your website. There is no smoothness between a title and another.
Also you could try to make a linear model for to see what effect each change has, but that doesn't take into account all the dependencies that can be complex, for example what title goes well with what background color. Evolutionary computation helps implicitly optimize without having to formulate a model.
With your example, let's say two parents with top/blue and bottom/red are chosen, then their offspring will either be top/red or bottom/blue because we make sure the same website isn't tested twice.
More generally each feature of the offspring will be randomly picked from the features of the two parents.
So two parents A/A/B and B/C/B can give the offspring A/C/B or B/A/B. There is also some random mutations that are possible with a low likelihood.
link to the product: https://www.ascend.ai/
Thompson Sampling / Multi-armed bandit algorithms are traffic allocation strategies to use the least amount of traffic to find the optimal variant.
So say you have version A/B/../N variants of a website, you can split the traffic equally or use a a bandit algorithm and find the best performing variant of it.
But that only helps you test N variants. If you want to test for 4 titles and 3 color buttons, 3 layouts and 5 images you already have 4x3x3x5=180 different variations and you can't test them all. Evolutionary algorithms can help you search through a much bigger search space:
First you try 10 variations, and allocate traffic equally or through a bandit algorithm. Then you combine the top performing candidates multiple times to create a new generation and start over again.
Evolution helps you find which candidates to try in a large search space and then a bandits algorithm can help you allocate traffic optimally.
You have to be careful with A/B testing with Bandit algorithms though. If the conversion rate changes over time or if you have visitors that don't always convert instantly you have to take that into account: https://www.chrisstucchio.com/blog/2015/dont_use_bandits.htm...
I don't think they account for potentially changing conversion rates over time or delayed conversions.
Aside from that, I'd be curious to see how these two approaches compare in a real-life situation.
The allocation of traffic based on the evolving optimal search space (blue button for visitors from Facebook) can then be driven through an MAB or something similar.
My domain is basically a combinatorical problem on large, sparse graphs. I am currently focused on DL with lots of custom feature engineering. I am making progress but its slow.
In fact we have many such problems in our own universe. I can give many examples of things NNs and GAs don't work well on right now. Those are just ignored by the research.
Solomonoff induction is the best attempt to try to formalize machine learning. And in theory any machine learning algorithm that works, works because it approximates Solomonoff induction somehow. But proving anything approximates Solomonoff induction is absurdly difficult or impossible. Because it's incomputable and involves the search space of all possible computer programs.
i would be surprised that "this stuff" would be exception to the unreasonable effectiveness of mathematics. mathematics underpins virtually every observed phenomenon, including theoretical physics, computer science, economics. in fact, the mathematical structure of any physical theory often points the way to further advances in that theory and even to empirical predictions.
to not expect that "this stuff" should not have any mathematical foundation is a fantastically naive view.
I formed a similar impression in my PhD research.
The thing to do is change the problem, transform the space/dimensions, so better solutions were spatially proximate. But then, I'd be solving the problem.
Another approach is to have the computer do this, seach the space of search spaces. But this higher-level space is even less likely to have informative gradients.
OTOH, the data was in terms of a language, which would have introduced its own artefacts. A better search space would compensate for those, and might have been easy to find.
> I don't know how Genetic Algorithms/programming could be made 'explainable' as to why they achieved an optimal solution other than hand-waving to how evolution works in nature.
This is quite confused. You're comparing different levels of the systems. In neural networks, we would like to know why a numerical model (which has been optimised by gradient descent) gives the outputs it gives. In GAs, (1) the objects being created usually aren't numerical models -- think instead of solutions to TSPs; (2) the reason the object is good is rather easy to see -- one just has to look at the objective function and verify that that the object has the desired properties; (3) we don't really care what other objects were considered during the search process.
Minor point: whether the result of evolution is easy to understand or not depends on the representation (encoding). Even GA results can be difficult to understand if they describe complex objects. In the case of GP, bloat (rapid increase in average program size in your population) can make results very difficult to interpret.
With either GA or ANN, you don't know that the algorithm has achieved the optimal solution and I would be surprised if you could ever prove that in the general case.
What you can know is that your algorithm performs well, e.g. by testing how well it plays go or drives a car or recognises faces, or whatever.
Let me first say that I'm not certain this is the case for the purposes lined out in the article, because I'm not entirely sure what they're trying to do.
Unless you have a crossover operator that really makes sense for your fitness function and problem, a GA is basically nothing but a bunch of SA processes running in parallel. In that case you would prefer SA, because it has less hyperparameters to tune, convergence is better defined, and there are proven methods to tune the hyperparameters.
It's not that hard if you have a working GA optimizer, to test how well it does with SA as a baseline, because the algorithms are fairly similar. Unfortunately not many demos do this baseline comparison and I'm pretty sure most of them won't do significantly better with GAs.
I enjoyed it much more than - what feels like - a quickly thrown together marketing piece with no real value for the reader.
Ken Stanley and Risto Miikkulainen original NEAT (NeuroEvolution of Augmenting Topologies) paper: http://nn.cs.utexas.edu/downloads/papers/stanley.ec02.pdf
Ken Stanley's novelty search page, and a link to his book, "Why Greatness Cannot Be Planned: The Myth of the Objective": http://eplex.cs.ucf.edu/noveltysearch/userspage/
Risto Miikkulainen's Evolving Deep Neural Networks paper: https://arxiv.org/abs/1703.00548
Ken Stanley & team's work at Uber, with links to some recent papers: https://eng.uber.com/deep-neuroevolution/
I couldn't bother reading the original article, but designing neural networks via evolutionary algorithms is a very interesting concept I wasn't aware of.
The OP is by Risto Miikkulainen who was Stanley's collaborator on NEAT and should not be summarily dismissed.
However, these overview articles do not include the newest research on evolving deep learning networks. Three such papers are introduced at https://www.sentient.ai/sentient-labs/ea-1/; there are other recent ones at https://research.googleblog.com/2018/03/using-evolutionary-a... and https://eng.uber.com/deep-neuroevolution/. It is a rapidly developing area.
What utter nonsense. Genetic Algorithms do exactly the same thing that Deep Learning methods do: optimize a function for a particular criterion. Genetic Algorithms are useful when taking gradients is not viable, as with RL methods - and RL methods can use Deep Learning! Seriously misleading.
Also, this: "Remarkably, although several human-designed LSTM variations have been proposed, they have not improved performance much—LSTM structure was essentially unchanged for 25 years. Our neuroevolution experiments showed that it can, as a matter of fact, be improved significantly by adding more complexity, i.e. memory cells and more nonlinear, parallel pathways." - this is not at all a new idea? How can you even pretend what you're doing is novel.
The entire article repeatedly relates everything in the context of what DL accomplishes and only introduces EC as an better optimized version.
Which doesn't lead one to think they are significantly different approaches but variations or tweaked versions of the same general thing...
This is just wrong.
We've had "learning-to-learn" algorithms for a few years now. Including LSTMs that can be learn the gradients for other LSTMs, and deep-RL algorithms that can optimize neural networks.
Its hard to say its inventing new things... there is still a clear goal - a loss - and we are optimizing it; poorly in the case of genetic algorithms.
I would say genetic methods are simply poorly described reinforcement learning problems, which means that there is a 1-to-1 mapping between Deep Learning and Genetic Algorithms.
PS. Making a distinction between "Genetic Algorithms" and "Genetic Programming" is like calling Deep Learning "Differential Programming" -- changing the name of a thing does not change the thing itself.
Also note that in the 90s things produced by genetic programming went patented, because they were novel algorithms.
An episode in this case is an iterative call to the Deep RL system until it outputs <STOP> at which step you give it a reward (negative number of collisions say), and before that you can send the actions to your stack machine.
Once you're satisfied with the final number, just concatenate all the produced actions to get your final program.
I just don't see a fundamental difference, especially once you start doing things like Asynchronous Actor-Critic et al. And then you can start doing Monte Carlo Tree Search on top of that since you have a simulator available...
But that's just my guess. And even so, I am not terribly convinced that would lead to astonishingly new activation functions that are better than the known ones.
Genetic algorithm: Shift parameters in a batch of random ways, evaluate using fitness function, keep the best.
There is nothing more fundamentally 'creative' about one than the other; these are both just attacks on the problem "here is this function, find parameter settings that get higher scores in it".
As you say, GAs evolve parameters to the fitness function. But GP does, in fact, evolve programs.
They are very distinct schools of work that evolved independently. Some other major schools are Learning Classifier Systems, Evolution Strategies and Evolutionary Programming.
The high-level name these days is Evolutionary Computation. In turn typically lumped under Nature-Inspired Computing, or sometimes Metaheuristics.
http://www.genetic-programming.org/combined.php
Here's a fun article from Popular Science on Koza and his "invention machine" that readers might enjoy:
https://www.popsci.com/scitech/article/2006-04/john-koza-has...
The difference is that for a GA, that can be any function, whereas for NNs it must be decomposable into a sum of partial objectives, and differentiable. Moreover, the result of training an NN is a function mapping a vector to a vector, whereas the result of training a GA could be an object drawn from any search space.
Most DL work parameter tunes a structure that is hacked by human guesswork.
The two approaches have strengths and weaknesses, and in no way does this make EC in general superior to DL, but perhaps different.
E.g. if you try the LSTM Music Maker they link to (https://www.sentient.ai/sentient-labs/ea/lstm-music/) and enter a melody, the resulting 'improvisation' will neither have any of the hallmarks of your initial input nor obey any of the conventions of any genre of music I recognise. It'll just spew out a random-seeming spray of notes. Compared to, say, JukeDeck (https://www.jukedeck.com/make/track-generator/essential), which uses bog-standard AI, it's got a long way to go.
The challenge with all GA work is finding the right fitness function. If you don't understand your domain well enough to define a good fitness function you're going to be wasting your time, and no amount of algorithmic and/or AI magic is going to help you.
Evolutionary approaches have always had one big feature in their favour: they are far more fun to work with.
They produce all these fascinating oddities, like the one that learnt to outsmart it’s opponents at infinite tic-tac-toe by playing coordinates 102312 and 47875 and watching them run out of memory trying to build a data structure for the board.
They are also far easier to combine with human expertise, i. e. “I can do this! Let’s throw some of my ideas into the gene pool”.
But empirically it’s hard to deny that neural nets have been used for some incredible things over the last years. I think it’s the rare hype that is deserved.
I guess the idea of them being evolved is that you don't _need_ to understand them, but evaluation performance is definitely a concern and doesn't seem to be addressed much. Interesting early work though!
I think that evolutionary approaches are a lot better at the exploration phase, but weak at exploitation; gradient methods tend to be the opposite.
I think some sort of combination of the two, like Max Jaderberg's "Population based training of neural networks" [1], is the way to go.
[1]: https://deepmind.com/blog/population-based-training-neural-n...
I'm not professionally versed in NNs myself, but what would be the closest equivalent to recreating this sort of training with tools available today? (in Keras or R for instance)
Was this ONLY trained by Bach Chorales? Everything sounds so fugue.
Schmitt, Lothar M (2004), Theory of Genetic Algorithms II: models for genetic operators over the string-tensor representation of populations and convergence to global optima for arbitrary fitness function under scaling, Theoretical Computer Science 310: 181–231
It took a while though, and it's possible alternative approaches would have been produced a similar result more efficiently.
[1] https://en.wikipedia.org/wiki/Recurrent_laryngeal_nerve#Evid...
We can't easily do that with real evolution but it's configurable with evolutionary algorithms, as with most machine learning algorithms. The hard part is finding those parameters.
'Evolutionary' techniques, are, in some sense, 'trial and error'. For many tasks where we already have efficient algorithms, this isn't particularly useful. However, where we DONT have efficient algorithms, or know what the concept of what an 'efficient' algorithm even is, trial and error techniques, like evolutionary algorithms, seem like great choices.
Death to CAPTCHA
In contrast, evolution strategies have a search distribution, and compute a gradient update to that distribution from an entire generation. So it's really more 'search gradients' than evolution. The only thing 'evolutionary' about it is that the minibatches are called generations.
"The AI is processing your melody...", sure, but it can be processed without AI and you could get a much better output. At least please keep the output notes on the same input scale...
evolutionary algorithms are used in practice when there is no real alternative. eg. bin packing for CNC/laser machines (self-plug svgnest.com)
http://www.springer.com/gp/book/9780387332543
He liked to have his students compare random search to evolutionary algorithms. There tended to not be a huge difference. I think that's why they are not so widely used when there is any kind of better method around. Probably in most cases you'd want to just understand the domain better.
Hence, I don't believe EAs are the next big thing.
Essentially, variation comes from many other sources than random mutation on DNA. Random mutation itself is seen as mostly a bad source of variation, leading to destruction of the genome.
http://extendedevolutionarysynthesis.com/about-the-ees/why-i... explicitly rejects the idea of a revolution regarding mutation and other genetic sources of variation:
> How can the EES seek 'profound change' and yet disavow 'revolution'?
> All recognized causes of both evolution (e.g. natural selection, genetic drift, mutation, etc) and inheritance (e.g. genes), as well as the vast body of empirical and theoretical findings generated by the field of evolutionary biology, are accepted by the EES.
> Hence the EES does not entail a rejection of current understanding within the field, and does not require revolution. The EES seeks only to supplement the existing causal framework through recognition of additional causes of evolution (e.g. developmental bias) and inheritance (e.g. epigenetic inheritance), entirely complementary to those long-established within the field. Nevertheless, should these additional processes prove to be important, the conceptual change to evolutionary biology could very well be fundamental.
Although not in your source, I did find the following. Known as Drake's rule: as a genome's size increases, the mutation rate also tends to decrease. https://en.wikipedia.org/wiki/Genome_size#Drake's_rule. A recent update to Drake's rule suggests that mutation rate is also negatively correlated with population size too, not just genome size (https://www.ncbi.nlm.nih.gov/pmc/articles/PMC3494944/). And another interesting article about mutation in reproduction: https://www.nature.com/articles/ncomms15183 (somatic mutation is two orders of magnitude higher than germline mutation). This is intriguing, though I have to wonder how the genome ever grows if mutation supposedly becomes so strict over time as to become a non-factor (transposons, gene duplication?). It's worth noting that mutations do occur more frequently in non-coding sections of the genome, which may not be accounted for in the studies I've linked to.
Still, regardless of my sources, I would like to see you provide support from specifically EES proponents for the devaluation of mutation as a source of genetic variation, as you implied.
... is the ability to optimise a black-box objective.