So how do you optimize a layer? Do you still use gradient descent? So you are have a per layer loss with a positive and negative component and then do gradient descent?
So then what is the label for each layer? Do you use the same label for each layer?
And what does he mean by the forward pass not being fully known? I don't get this application of the blackbox between layers. Why would you want to do that?