Can anyone explain how the "floats" that are the output of the last layer correspond to the individual digits?
So I guess during training you're telling it that correct answers should be 1 and the incorrect answers should be 0.
Do the encoding choices that you make regarding the input / output of a neural network influence its performance at all? Maybe for MNIST the way you have it is the most common approach?