I think you misunderstand. Especially
> this model contains a smaller model, that has similar performance to the trained large model, without training
The point is the opposite. There is a small net X within big net Y, such that training only X gives the same performance as training all of Y.