First paragraph: "Examples of effect sizes include the correlation between two variables,[2] the regression coefficient in a regression, the mean difference, or the ..."
> Who would consider the same expected difference to mean the same effect size when the population distributions are narrow and when they're wide?
For the case where the unit of measurement has a sensible, relevant interpretation (e.g. if a study measured dollar value of two interventions), and attempts to capture uncertainty via CI or another means, I would consider it one meausure of effect size.
The key to understanding effect size is that its most common focus is around making measures over arbitrary scales become scale invariant. But scales are not always arbitrary.
Daniel Lakens has a great article on ES and puts the motivation for calculating them well..
> First, they allow researchers to present the magnitude of the reported effects in a standardized metric which can be understood regardless of the scale that was used to measure the dependent variable. Such standardized effect sizes allow researchers to communicate the practical significance of their results (what are the practical consequences of the findings for daily life), instead of only reporting the statistical significance (how likely is the pattern of results observed in an experiment, given the assumption that there is no effect in the population).
https://www.frontiersin.org/articles/10.3389/fpsyg.2013.0086...
To the degree that the thing measured has an inherently meaningful scale to the researchers (eg sometimes dollars, time), then it is already an effect size measure.
There is some nuance here, since the level on which you might want to interpret something (eg the spread could be part of the value, especially to an individual who will have 1 and not many codebases).
You might also want to render it comparable to other studies that measure something else, and it's unclear how to convert it to your scale, so you use standardized ES measures (eg many meta analyses).
But an important point is that whether something qualifies is really a question of the scale. (What the best ES is for your specific question is another important issue!)