>Visual sharpness at the expense of wider-scale coherence (see: sliding/floating walking woman in Tokyo demo or tiny people next to giant people in Lagos demo)
Wider-Scale coherence is still much better than previous models and has consistently been improving. It's not "visual sharpness at the expense of coherence". At worst, the models are learning wider-scale coherence slower.
Not everything is equally difficult to learn so it follows that some aspects will lag behind others. If coherence weren't improving you might have a point but it is so...