According to Figure 14 it seems to be the case, with the difference between the two group means being large enough to be significant at p < 0.002 (according to their numbers). A good threshold is usually considered to be < 0.05, so this is a pretty strong result. (The p-value is the probability this difference is simply a statistical fluke and not actually meaningful.)
Here [1] is a discussion on how to pick a good sample size for an experiment. As you can see in the table to the right, if the effect is strong a sample size of just a few dozens can be perfectly acceptable!
Here the biggest threat to the validity of these results is not the sample size, but whether one can generalize the case of the pilots in-flight (a very specific task performed by a very specific type of individuals). As trentmb says it's not clear whether pilots are very representative of the general population in this situation, although it also looks like there have been multiple studies done on the subject [2] which seem to confirm the benefits of napping as well. For what it's worth it definitely corroborates by own experience as well so I'm actually quite curious to see where Napwell will go.
[1] http://en.wikipedia.org/wiki/Sample_size_determination#Requi...