With IQ tests, the only thing you know that if 2 people fill it out and one of them scores higher, they have a higher IQ. Based on this you can sort people into a list by increasing IQ, but the standard distribution is implied not discovered.
This is like having a double-sided scale and a bunch of weights - you can similarly compare the weights to each other, and sometimes the arm of the scale will lean to the left, sometimes to the right, by a little or lot - you can postulate that the weights are normally distributed are normally distributed, but they are absolutely not required to be (I can choose them any way I like), so your assumption would be wrong. We know this because we have a direct, not just a comparative measure for weight.
I could make up an imaginary 'weight point' scale based on these comparisons, and say weights A is 5 WP heavier than B and C is also 5 WP heavier than A.
But A might be 100g, B might be 1g, and C might be 1kg.
This is what I think of when I see studies clamining the difference between 2 groups was 5 IQ points.
It's not nothing, but IQ is already a little squishy. No one's IQ is a single number. But the article also goes into problems with the study and other potential issues.
Basically, they're saying there is this pattern in the data as recorded, but there are multiple confounding factors and issues with collecting the data in the first place.
(It's also just fine if we disagree about this --- researchers do too!)
Even height and weight change throughout the day. People are typically taller and lighter in the morning than in the evening. Weight especially is variable, it can fluctuate up to 5 to 6 pounds.
If someone is 15 points above average, they are in 84th percentile, or in top 16%.