Your results might seem spurious because of the small sample size, but when aggregating the results of all the participants they will have enough data to be able to conclude how many people did act like you did with apparent preferences due to chance, and how many actually where "biased" in some way.