If half of all wallets get returned, then the variance of the number returned is on the order of 40 * 1/2 * (1-1/2) = 10, which means a standard deviation on the order of 3 wallets returned, or ~ 8% return rate. So the "puppy" and "family" figures might be out by about that much.
The cute-baby category had a measured return rate of 88%, which means a variance of something like 40 * 0.88 * 0.12~=4.2, for a standard deviation of ~ 2 wallets or about 5% in the return rate.
So if these results are unlucky to the tune of two 2-sigma errors pointing in the "right" direction, the puppy category might really be as good as 69%, and the baby category might really be as bad as 78%.
So, at least as far as simple sampling error goes, the "baby beats puppies" result seems pretty robust.
(No need to tell me about all the oversimplifications in the above. I know.)