This places exactly zero constraints on the utility of IQ tests. Instead of asking why group A scores higher than group B, I can simply ask why group B is impaired relative to group A. Apparently that should be enough to convince you of the validity of the comparison.
> The variance of measured IQ for the same individual over time is very high.
The test-retest correlations of IQ batteries are typically in the 0.8-0.9 range. For example, the Wechsler test [1]. This makes IQ probably the single most reliable measure in the social sciences. Perhaps you're referring to the Wilson effect, but that's not at issue here.
[1] https://images.pearsonclinical.com/images/pdf/wisciv/WISCIVT...