There is lots of work being done in model compression (quantization, simple factorization tricks, better conv kernels like depthwise separable convs, etc). We won’t let that happen!
Don't forget that folks were saying this about the web when images / rich media were becoming prevalent!
Details please. What techniques are used to reduce the model size?