Happy to — please open an issue at
https://github.com/fstandhartinger/jevbench/issues with an endpoint or runnable code, the model and licence, and whether it was trained on the public items. Every entrant runs through the same harness, including the sealed set.
I definitely prefer open models I can run locally, because sending our private held-out set of tasks to an external API produces some headaches on ur side and basically means we have to rotate the test set a lot, to avoid contaimination.
We can call API hosted models though, we'll just flag the leaderboard entries appropriately.