Why do you care about the public benchmarks at all?
The way companies and effective individual developers use OpenRouter, is that you first create your evaluation framework/benchmark, for your specific tasks and use cases, and make that real easy to run and use various models and providers with it.
Then you run this to gather data. Then you use said data to figure out what works and what the quality/cost tradeoff you want to make is. Then you lock that down in production while you keep iterating on your benchmark to make it match with real-world use cases and keep adding the new models that pop up.
I don't think anyone serious is just willy-nilly making individual requests against OpenRouter and similar platforms, get a "feel for a provider" then use only that provider. Not only would it be wildly inefficient, but also you need hard numbers to compare so you can make informed choices.
For this process and workflow, OpenRouter is great, because adding/changing providers and models is essentially changing two strings, rather than having a adapter for each platform you want to try out.
If you just want best accuracy requests from SOTA models for your agent you run locally or whatever, then don't use OpenRouter, it doesn't make much sense, but use the provider the model maker has available, as almost all of them run their own endpoints.