I think educators are too afraid in general of getting into some real accurate phonetics IMO. Lowest hanging fruit seems to be phonetic diagrams of where the tongue/lips should be. Perhaps normal in other peoples experience?
Again, I'm curious if someone solved it. Theoretically this is within modern video/audio AI, but that doesn't mean we have the models or even organized data.