I think there's a lot of subtlety being missed between dynamic range and resolution. For example, I think your JND (just noticeable difference) assumption is incorrect; my understanding of the scientific consensus is that a 1db difference at any volume is consciously perceivable to "normal" human hearing (and 3db to virtually all humans), and a 0.2db difference has been shown to be subconsciously perceivable by most. This places the required bit depth (not counting dithering) of an audio recording spanning all human perception WELL beyond 96db, and probably somewhere in excess of 120-140db! This is consensus, BTW; hence most CD-apologists people defer to the dithering argument :)
Also, I did finish reading that article up to the point about bit depth, and unfortunately there's quite a bit of either bad and misleading data in there unfortunately. While the article does contain some very good true info, it's sadly riddled with enough bad/outdated claims that I can't really recommend it to anyone as a reputable/trustworthy source of truth. It strikes me very much as if the author is unaware of how imprecise our approximate understanding of psycho-accoustics actually is, when it makes extremely overconfident claims like "CD quality will be enough FOREVER".
For example, the point made about near infrared being invisible to all humans is obviously false: most humans can't, but a certain percentage of blue-eyed humans can see near infra-red (including remote control IR). I've known one such person personally, and have tested and confirmed this thoroughly. I agree that it's silly to try to make a TV that emits these frequencies (and cameras that capture them), but my point here is simply: there's a lot of bad info in that article, perhaps shamefully so for an article making such bold and confident assertions under the name of "science".
Regarding bit depth, the article doesn't really even try to dispute the fact that 96db is insufficient; it just says with dithering, we shouldn't have to worry about it. I'd love to agree, but the combination of the meta-analysis that shows it's not sufficient is all I need to prove that your linked article is simply wrong.
There many be a wide range of reasons why the meta-analysis found high res audio to sound slightly better, but none of that invalidates the validity of the results themselves.
For example, maybe most DACs aren't very good at replicating dithered subtleties (I'm just speculating examples here) -- supposing such a common problem exists, then we could debate whether it makes more sense to improve DACs with fancy improvements to dither reconstruction filtering -- OR -- we could just encode at least 140db of dynamic range in the sample bit depth in the first place! The latter seems a far simpler and less-overengineered solution to me.
Dithering is fine when it works, but let's be honest: fundamentally, dithering is a kind of compression algorithm that encodes a greater dynamic range within a signal with an artificially constrained dynamic range (but with the side-effect that not all DACs will handle it equally well). In the modern digital world, there's no need to rely on dithering any more in audio signals, in the same way that we don't see dithered 256-color images much any more. Why don't we leave compression up to the compression algorithms, rather than promoting ACDs/DACs that are glorified compression/decompression hardware? Modern lossless compression algorithms are vastly better anyway, if bitrate savings is the goal.