(1) The NN uses two convolutional units and a fully connected softmax layer. Relative to Inception V3 or highway networks, this is _not_ a very deep architecture. I was looking for a balance between accuracy and training time (trained on my MBP).
(2) I looked into other neural architectures (LSTM and several fully connected) without much difference in performance. If you take a look at the physionet google group, there are a number of other methods evaluated (logistic regression, SVMs, etc.)
(3) I did vary hyperparameters and saw a dropout of ~0.45 and frequency cut-off of 4Hz performed the best for this specific architecture. That said, I imagine the best performing features would be a concatenation of the output from several filters across a range of thresholds. Then the burden of deciding feature importance falls onto the learner.
http://arxiv.org/pdf/1507.06228v2.pdf http://physionet.org/challenge/2016/#forum