It’s probably more memetics (constrained by cognitive load) than genetics that select the features of languages, and there are likely many other factors that determine the complexity of languages and grammatical structures: For example, there is a sweet spot between minimal symbol count and minimal lengths of the words. In the extremes you have either short words but many symbols, or few symbols but long words. At the same time the number of distinguishable phonemes and therefore symbols are, of course, restricted by the sounds that are producible by the average vocal tract.
Secondly, the communication channels thought→vocalization→hearing→thought or even thought→typing→reading→thought are inherently very noisy, so you end up with a lot of redundancy, like particles, introduction and transition phrases.
And lastly, I think that there are always some words and phases that are not shaped by efficiency/cognitive load, but rather by whether it is fun or fashionable to talk in a certain way. There is certainly some cultural variance that can be orthogonal to efficiency.