I also have a hard time believing Netflix would be ok putting up money to capture a picture/series all in 16mm film (or that anyone would suggest the commercial success or artistic vision demanded it).
So it doesn't stifle creativity, it future proofs Netfix for the lucrative high end market, AND Netflix is paying (at least indirectly) for that level of quality. I don't see the problem here.
The problem is that the camera requirements are defined, presumably to try to have some lower bound on the quality of the visuals, there is no requirement for audio such that it sounds good on stereo speakers, which is probably what the majority of consumers are using.
The idea that Netflix have such a strict definition on camera quality is also a bit of a farce, given how woeful the image looks after it goes through their incredibly overenthusiastic level of compression, but that's neither here nor there.
Compare to Disney+. As much as I dislike Disney, they have 4K Dolby vision for no extra cost, and quality is excellent. Prime and Netflix could both take examples out of their book. (Some of the compression in recent Amazon shows has also been very bad. For example, the trees in the background of the wheel of time were often a mess of compression artifacts even at 4K.)
That last point means that dialogue is likely to be quieter than sound effects and music, since dialogue is usually on the centre channel.
Who in the world only has 2 speakers? /s
The technology is old enough to be long out of patent either way.
Metadata only isn't a panacea
This is great if you have one.
However if its downmixed on the fly (ie done on the client side) the center channel is often just played out of both speakers. This means that sound can be muffled, because its drowned out by the incidental noise that comes from the left/right channel.
Well what other option do you have when there are only two speakers? Only play it out of one?
The defaults for automatic sound mixing will almost always be wrong. And they will differ in how they are wrong from consumer box to consumer box.
OTOH, a lot of the models end up trained on features that are very different from what humans hear.
When you master for 5.1, because you have the centre speaker, you can have dialogue at 100% volume, then out of left and right, you can have incidental noise, be that music or "atmosphere" also at 100%
(its been a while since I've mastered in 5.1) However, two channels of 100% volume (well 0dbu) is louder than just one channel. Which means that if you have lots of music, wind or other foley it'll drown out the dialogue.
It requires artistic choices from the sound team to make work properly.
> This means that sound can be muffled, because its drowned out by the incidental noise that comes from the left/right channel.
As always, it depends on the device. Dynamic range compression seems to be a relatively common feature, usually as an option described (inaccurately) with something like "Reduce Loud Sounds" like it is on the Apple TV.
Consumer hardware can only guess. A sound engineer can know.
Thing is most studios now do only one mix for theaters and also most streaming providers don't give many fucks about audio. It's not an innovation problem it's a product problem.