So I guess while "realism" (or believability) is really good now, prompt adherence has much room for improvement.
(though put it another way, realism has always been "solved" if the model gets to output whatever it wants as long as it looks realistic, though now it looks less like a malfunction and more like an inattentive human mistake or oversight, so even when it gets it wrong it's hard to tell it's wrong without knowing what the prompt was)