When A.I. Matures, It May Call Jürgen Schmidhuber ‘Dad’
nytimes.com
nytimes.com
Here are some notes in case anyone's interested -- [0] and [1]. I also recommend his TEDx talk.
[0] https://www.dropbox.com/s/v6bbktuywoqv2w3/jurgen-talk-notes....
[1] https://www.dropbox.com/s/ux53ism7fsxgo8x/jurgen-csail-talk2...
I don't know enough about AI research to judge the value of Dr. Schmidhhuber's contributions, but I've seen his name multiple times in past HN discussions.
He's been an AI researcher for a while. I think his biggest contribution is understanding the vanishing gradient problem and creating the LSTM architecture. LSTMs are widely used in both industry and academia, and many neural net architectures that aren't LSTMs are heavily inspired by the LSTM idea.
One of his students, Alex Graves, is a researcher at DeepMind who is seen as one of the top people in RNNs.
That's an excellent criterion.
I often wonder if many famous past intellectuals were mere celebrities where I can't recall a single achievement. And if one can't name a famous true idea in an current academic field, perhaps the field itself is worthless.
The OP just has not heard of any accomplishments, but anyone with a little expertise in deep (reinforcement) learning knows about the major contributions to the field by Schmidhuber.
Using this criterion you are using popular media and fields you know not much about, to brush away the accomplishments of respectable scientists. Don't base your skepticism on your own lack of knowledge: that makes it selective -- You can not cut through the bullshit, if you don't know how to wield a sword.
I can't help but wonder if the sole reason AGI doesn't exist is because it hasn't been figured out yet.
While that statement sounds obvious on the face of it, the implication is that we may already possess both the sufficient computational resources and human intelligence to realize its creation.
What would you propose as an alternative? If nothing, fine, but how can we relate, when we only have a single best (or if you prefer: flawed) thing?
> It can't handle modelling itself
It can add its internal states to the environment, and hence model these internal states.
> the large dimensionality of its hypothesis space
Its bounded by how many compressors / programs are available to it. Calculating the length of these programs that are consistent with the environment is feasible.
> handle imprecise stimulus information
I'm not sure if you mean imprecise stimuli here or imprecise sensors.
> can be made arbitrarily "stupid"
This can be seen as a flaw, or as a simple property (or even a feature). Something that can be optimal, and "stupid", does not detract much from its ability to be optimal.
> before we get to mere computational limitations or the uncomputability issue
Just because it is uncomputable, does not mean the theory is flawed. Sure, it is not practical, and we like practical things, but it is still valuable to have such a theory. Especially when approximations do yield practical applications.
Marcus Hutter on these issues: http://hunch.net/?cat=14
I tend to favor Karl Friston's "free-energy minimization" theory of the brain. For specifying tasks in engineering situations, I like the KL-control paradigm: the agent's task is to minimize (via the normal mechanisms of active inference) the Kullback-Liebler divergence between its induced distribution over latent variables (ie: causal models of the world) and some "target" distribution.
>It can add its internal states to the environment, and hence model these internal states.
No, it can't. I'd have to do a bunch of work to sketch some proofs, but without the ability to consider multiple observable random variables and use hierarchical Bayes on them, AIXI will not be able to detect that certain environmental states are actually equivalent to its own internal states. Hell, since Solomonoff Induction is incomputable, it's not even in its own hypothesis space, so it can never locate a program that generates itself.
>Its bounded by how many compressors / programs are available to it. Calculating the length of these programs that are consistent with the environment is feasible.
See below, please.
>I'm not sure if you mean imprecise stimuli here or imprecise sensors.
Both. From the perspective of any possible reasoning device, the problem is simply low likelihood precision (that is, high variance/entropy of the likelihood function). If I have so much noise in my sensory stimulus that all I see is 50% "heads" and 50% "tails" in my input bits, without any ordering to those bits, then I do not have the information to infer any complex causal structure (ie: the real world) behind those bits.
The point being: when the environment is noisy, it forces the mind to favor simpler explanations, even when those explanations are not the Truth, due to the parameter space over possible Truths (for example, random seeds added to some causal structure) being too large and spreading out the probability mass too thinly. Since Solomonoff Induction deals with algorithmically random programs as its hypothesis space, this means that any noise in the environment sufficient to add one bit of random seed to the shortest generating program cuts the probability mass allocated to the correct hypothesis by half.
>This can be seen as a flaw, or as a simple property (or even a feature). Something that can be optimal, and "stupid", does not detract much from its ability to be optimal.
AIXI is only optimal for a given Turing machine (ie: state-machine with alphabet). The "arbitrarily stupid prior" thing is just to say that for any program, we can create a Turing machine for which that program takes up arbitrarily much tape-space, thus making that program arbitrarily improbable in the Solomonoff Measure over that Turing machine's programs.
>Just because it is uncomputable, does not mean the theory is flawed. Sure, it is not practical, and we like practical things, but it is still valuable to have such a theory. Especially when approximations do yield practical applications.
Hence why I called them "mere" computational issues.
(edit: apparently people are cool with this. That's weird. I'm not claiming the parent is untrue, it's just that you shouldn't talk smack about people in public without presenting evidence if you don't use your real name.)
From what does the supposed requirement to use one's real name derive?
The internet is protected from the forces that threaten physical forums where each voice is a real person. You cannot pay off the internet to talk nice about you, and you can't follow anon home and beat him up for speaking negatively. It sounds like Jurgen has accumulated some bad karma on the internet; this is not baseless accusation as much as it is keeping ourselves at a healthy level of skepticism regarding someone who has already lost our trust. The cowards with loud, strawman accusations are usually appropriately downvoted when their criticisms are out of place.
And btw, I think his sense of humour - and brilliance - comes across best in person. I recommend attending one of his talks or watch a video. But just do it once, otherwise you'll hear the same jokes again.
I think it would be interesting for him to talk about the relationship between his belief that a deterministic universe theory is possible, with his practice of using statistical learning algorithms. Some people might view ANNs (including RNNs) as good learning algorithms for sorting out statistical patterns of probabilistic systems, but not as helpful for analyzing deterministic systems. But I think there is some good insight to be had from exploring the value of statistical learning algorithms on deterministic systems.
Seriously though, what title will it bestow on Wolfram?
This is not a good article please don't upvote it.
If you think LTSM is interesting, I am 100% with you... but there must be for sure much better articles about LTSM or the author himself that you can pick from to share here.
This article doesn't add a lot of value, takes a lot of effort to make it's point and feels like a huge waste of time when you finish reading it.
Those are not mutually exclusive. If you want to actually claim that Chris McKinstry was only insane, and his project without any scientific merit, you should make the case for it. Brushing projects away, only because the conceiver of them was mentally ill, is not nice at all.
My point is about the how the article is titled and written not the researcher.