Hence it should be clear to anyone that without engine assistance it is totally impossible to determine for 10^6 games if a certain material difference was "clean" or not. Most likely he just checked for a stability over n consecutive moves, at least this is the usual way. Others have noted more problems, that might help put his highly deviating, washed out results in order.
I really doubt if there was any attempt to check if there's some "strategic compensation" - how would you do that at a large scale? I doubt that even running a solid chess engine evaluation on all these positions is feasible, you need something where you can simply/cheaply filter positions from the database and then just count the winrate.
In simple mathematical terms a clean material advantage is:
Overall eval >= material eval
If overall eval < material eval, then compensation(other player) > 0 = "the opponent has some compensation for the pawn"
Knowing how the word is used in chess is not enough to know where the line was drawn in the data here, which is the basis for the entire analysis shown.
You can of course give up material for positional advantages as well, but from the article I don't think the author analyzed that. It would be difficult to accurately measure that anyway, some positions can be up +5 points according to the computer but only if you find 20 perfect moves in a row. Needless to say, most humans would not find those moves especially in the lower ELO brackets that the article analyzed.