Sure, here are a couple of examples of ECP violations removing ambiguities.
1a. How often did you tell John that he should take out the trash?
b. How often did you tell John why he should take out the trash?
(1a) can either be a question about frequency of telling or frequency of trash
disposal, whereas (1b) can only be a question about frequency of telling. I asked
GPT-4 to explain how each sentence was ambiguous and it seemed to entirely miss the
embedded readings (the ones about frequency of trash disposal) for
both sentences, while finding some other ambiguities that were spurious
(such as suggesting erroneously that (1b) could be a question about how many
different reasons you gave John in a single instance).
Similarly, (2a) has both a de re and a de dicto reading, whereas (2b) has
only a de re reading:
2a. How many books did Bill say that Mary should read?
b. How many books did Bill explain why Mary should read?
That is, (2a) can be asked either in a scenario where Bill has said "read 10 books!" or
in a scenario where Bill has said "read Book A, Book B and Book C!" without
necessarily counting the books himself. (2b), on the other hand, only
has the second kind of interpretation. I've had mixed results with GPT-4 in this case (depending on exact choices of vocabulary, etc.), but it certainly makes some mistakes. For example, it says that (2b) can mean "John explained the reason for a certain number of books that Mary bought".
As the sibling comment points out, it would not show very much if
GTP-4 did correctly determine these ambiguities as it has had access to much more data
than a child. You would also need to show that the same statistical techniques
would work when applied to a realistic dataset.