In practice, these requirements are obviously at odds with each other. You have to balance the freedom available to a speaker with the mental capacity of a listener. If you want to accomodate speakers of languages that allow free word order due to case marking, as well as languages that have a fixed word order, you might end up with case marking and marks to indicate a specific fixed word order. But then all of these markings would have to be understood by the listener, who'd have trouble keeping the various forms straight.
However, you might be able to quantify the learning effort required and optimize for that. This is obviously hard for grammar and semantics, but could be done for phonetics.
Instead of trying to intersect the phoneme sets in use, or giving up and choosing arbitrarily, you could collect statistics on the ability to produce and distinguish various phones, and then optimize for communicability. A speaker can pronounce a phoneme as well as the allophone that is easiest to pronounce for them, a listener can distinguish a pair of phonemes as well as the most similar combination of allophones from them. Then calculate the expectation over all pairings of speaker and listener, maximize the bandwidth accounting for the required error correction, and you get your perfect phoneme set. I'm actually interested what the result might look like, whether there will be a few, clearly separate phonemes, or a huge number with potential overlap, that needs to be compensated by unambiguous vocabulary.
It also occurred to me that this might somehow apply to constructing programming languages, but I'm hazy about the details. Maybe something like Python's "one way to do it" vs other languages supporting a variety of styles?