The problem is that with max 1 token it is not 100% guaranteed that it will be a yes/no or a score. This model (and engine) solves that by taking the probability of the yes/no token. For 1 question it is faster but not a whole lot, the power comes from being able to ask more then 1 questions for merely a few ms extra per question.
E.g. 1 question would be 300ms but 30 questions would be 360ms total. Most of the time comes from the image decoding.