What is the optimization algorithm for these if not backpropagation? This page does mention the use of hyperparameters for optimization and compares itself to an adaptive stochastic gradient descent method, which makes me think that it is using backpropagation.
Ber ber ber, I think you're right. I was assuming a Cascade-Correlation DNN.
Definitely doing the standard back-prop algorithm. Check out the code, it's pretty straightforward and the whole lib is about 2,500 lines