And what were the results of these experiments?
What error rate can you reach with the smallest network architecture you tried for example?
For example: did you notice than increasing or decreasing network size required significant changes in other hyperparameters? Are small networks learning faster at the beginning of training before they start to plateau?