Unfortunately I don't remember the exact numbers, but I think it was a couple percentage points worse than we were able to get with the large models.
For example: did you notice than increasing or decreasing network size required significant changes in other hyperparameters? Are small networks learning faster at the beginning of training before they start to plateau?