do you somehow control how non-trivial the queries are? The LLM generates them, right?
what if every engine returns garbage, or on the other hand, handles them too well?
building a benchmark like this in a genuinely fair way seems extremely hard to me. I’m very curious about the details, of course within what you can share.