Seems they need to compare against "dumb" code completion. It seems that even when they are error-free, "large" AI-code-completions are just boilerplate that should be abstracted away in some functions rather than inserted into your code base.
On a related note, maybe they should measure number of code characters that can be REMOVED by AI rather than inserted!