http://www.pnas.org/cgi/content/abstract/97/9/4932
You can't discount that IOI statistic becuase it's an outlier. Every single participant at the IOI is an outlier in cognitive ability. Do you know about the theory of outliers?
There's this thing called the central limit theorem. It says that if you have a lot of small independent variables, randomly assigned some value, then the mean of all these variables (or, by the same token, the sum of the variables) is distributed approximately normally. But suppose the variables are not small, or they're not independent. Then the central limit theorem doesn't hold, and what you have, almost all of the time, is an outlier -- that's why there are often many more outliers than you'd predict in a given population, using a small sample.
Now, I'm not saying that g is zero. I said that psychometrics is a non-science, in the same sense that a lot of the social sciences are non-sciences (you can find papers which try to show a causal effect of insurance regulation on premium prices, ignoring profits entirely, for example).
The fact that g is non-zero can be readily explained by the following simple observation -- most academic subtests, including IQ's, rely on skills that are either practiced as a group, or on skills that are shared between subtests. One example is focus, in general. Another is visualisation. Another is working memory. And so on and so on.
Many of these skills are also practiced in situations, like school, where if one does well in one area, they do well in another. If you're the teacher's pet, you get more attention. If you're known as the bad kid, you're immediately discounted (and I've been on both sides). If you're poisoned against a learning environment, you just won't put any effort in.
So it's no mystery to me that g is non-zero. The point is that the field of psychometrics is totally absent of content. There's no objective test for the validity of a test, for example -- the best they have is g-loading. Over the years, this means that tests have become higher and higher g-loaded. Now this could mean that the tests are getting better, or it could be that the subtests only look different, they are becoming more similar in content.
I've been studying these tests, the actual tests, since I was twelve. It didn't take long before I figured out how poor they were at answering research questions, or questions of individual ability. If you get the chance, try to look up the history of the Stanford-Binet, or Terman's kids, or actually take a look at the scoring method behind most of these things. They're totally full of crap...