I'm fairly certain the point of the gp was that the number of parameters are also not additive.
just because the first selection layer is very thin doesn't mean that the network cannot be considered composable
which happens to be consistent with what openai is telling us