One solution is to come up with a new benchmark yourself.
Manually benchmarking it by coming up with 20 questions and feeding it to a pair of models and blindly choosing the best result can give you a pretty good figure.
And that can probably be done in under 20 mins of human time.