Learning a hierarchy
blog.openai.com
blog.openai.com
The logical extreme of this thinking would be agents that actually maximize entropy of future actions as the only objective function, like in [1]
[1] http://paulispace.com/intelligence/2017/07/06/maxent.html
At the time, I remember being excited to hear of a physics paper[1] concerning an inverted pendulum, where they solved the system for some dynamic forces which would keep the system at the position of maximum instability, and claimed that it was, in some sense, a description of dynamic intelligence. The analogy there is that this is the unique position from which the pendulum can be efficiently made to move quickly in any 'required' direction (the 'goal'.)
I still think that idea has some merit, but putting together a coherent formalization of it seems really tricky and requiring some genius far beyond my own meager pondering.
[1] I found the article: https://physics.aps.org/articles/v6/46
It's possible to create rules that operate on an already-trained network and push it in this direction without totally destroying what it's learned, by "fuzzing" the original network to generate a bunch of input/output pairs, and then using that dataset to retrain smaller sub-networks. For instance, if you have a 5 layer network that you've trained on a classification task, you can often use that network as a teacher to train a smaller network to do pretty damn well on the same classification task, even in some cases where training the smaller network directly would have been very difficult. There are several reasons that this trick can work, not the least of which is that in a sense it is a way to expand the training set dramatically.
NB: the above approach is probably not how you'd implement this, there are less crude methods to incentivize shallower levels to have more activation than deeper ones that would probably work better
I can easily imagine a phased training strategy that oscillates between a) learning new things by making the deeper layers more malleable and the shallower ones fairly rigid, and b) compressing all the data by opening up the shallow layers to change and replaying input/output into itself. I have no idea if there are any benchmarks around this sort of thing, though, typically benchmarks have fixed goals so ability to retrain for additional tasks is not really measured.
[1] http://www.actinginbalance.com/intelligence-is-adaptability/
It's that the spoils of the new economy are accumulating in a way that completely forgets the middle 90% of the country. Kevin is obviously really smart, but has access to things I don't even have in a state school by virtue of being a sharp high schooler in Palo Alto, much less when I was in high school.
> Life is long. If his location gives him access to opportunities that you don't have, figure out a way to get access to those opportunities and execute on it once you graduate from college. Many prominent
> Silicon Valley people came from small towns in the mid-west (Marc Andreessen, Evan Williams) or immigrated from poor political situations abroad (Sergey Brin, Jan Koum, Elon Musk).
Being in the Bay Area already gives you huge advantage over most of the population, especially when you compare to less developed countries.
that's what I do the rest of the day because it's part of my job
I more mean the hardware access part - at 15 my parents would have never given me their debit card to spend hundreds of dollars on GCP GPUs - good luck training GANs on a laptop CPU!
> but has access to things I don't even have in a state school by virtue of being a sharp high schooler in Palo Alto, much less when I was in high school.
and then you go on to say
> I more mean the hardware access part - at 15 my parents would have never given me their debit card to spend hundreds of dollars on GCP GPUs - good luck training GANs on a laptop CPU!
That has literally nothing to do with location as you seem to allude to in the earlier post. It has nothing to do with the spoils of an economy being distributed unequally. Maybe if the hardware was only accessible in certain parts of the country, sure your point makes sense. But anybody with money could've bought it.
So your post now reads as "I'm going to blame me not achieving as much as Kevin on my parents for not spending money on me when I was young."
That article was encouraging, if anything. It shows exactly how available educational resources to the field of AI have become that a 15-year old can have access to them and make significant progress. it shows if you take initiative, you can actually go ahead and get things done.
Palo Alto is one of the wealthiest cities in the country.
Furthermore, outside of AI there is so much fun to be had in the world that it's probably not worth being discouraged by discoveries made by some preternatural high schooler. https://xkcd.com/1024/
https://www.wired.com/story/meet-the-high-schooler-shaking-u...
Kind of reminds me of the Soar system except using Deep learning instead.
It will be interesting to see how it performs with more tiers in the hierarchy, and with more structured tasks.
Controlling a virtual arm to play a board game for example.
https://s3-us-west-2.amazonaws.com/openai-assets/MLSH/mlsh_p...
Also, for small numbers of sub-policies, would Monte Carlo playouts be faster. Where we are searching over the next step the Any may encounter. Which presumably is a finite set of possible "wall-floor" configurations ;)
In any case, great work! Always love watching OpenAI vids...
My intuition from working on various computer vision tasks, is that animal brains would do this more by rotating the perspective at the post-optic synapses, rather than having a generalized plan. We still only know how to "move up"; we just change the angle we're understanding the scene from, and so change what "up" means.
It's (somewhat)obvious that this is an idea worth trying. But that doesn't mean actually getting it to work is easy.
Good work nonetheless but for god's sake give it six legs and make it black