For example, something that should be 15-2 votes, ends up being 95-82. You'll be applying a large number of votes to both sides, and pushing everything towards a 50/50 rating. This doesn't help anyone, the goal is A/B testing, and you're making it more difficult to get accurate data. 15-2 shows a lot of promise for A, but then you add 80 random votes to both sides, and 95-82 seems like a tie.
It's all in the presentation. If you just highlight one image and stamp "WINNER" next to it, most people won't even look at the numbers. Crowning a winner is more important than being scientifically accurate.
I suppose the question then becomes: will the end user notice or care if the votes are random? Do the votes need to come from humans at all?
Instead of showing the user 95-82, or 25-12 you tell the user "others prefered shirt A 2-to-1 or shirt b 66% to 33% or whatever might be appropriate.
Anyway, again, I'n not suggesting any particular techniques are the right ones, just that there (almost certainly) viable techniques available to mitigate the problem.