You are however welcome to do some statistics with actual samples of talking to people and having them show their skills at scansion, and prove my estimation wrong. Heck, I'm willing to bet less than 1 in 100 Americans is able to even define what "iambic pentameter" is. Feel free to prove me wrong on that one too.
Except if your concern is the measurability of the skill: whether it's 100x and not just 10x or maybe just 8.345x better. In which case you've missed the point of using it as a stand-in for "much better".
The comment is referring to GPT's ability to scan and analyze the meter or rhythm of poetry. Scansion is the process of analyzing the poetic meter by marking the stressed and unstressed syllables in each line.
The comment implies that there were two instances where GPT failed to accurately analyze the meter of two different poems in a single day. This may suggest that GPT is not proficient in scanning poetry or that it may not be the best tool for analyzing the meter of poetic works.