Promptbench: A Unified Library for Evaluating and Understanding LLMs
github.com
github.com
Also, technical report on arxiv https://arxiv.org/abs/2312.07910
I think that something that would make the results better would be highlighting which models perform best for specific tests (ie color coding) and explaining the tests via some hover info.
Also with some fine tuned models training to get higher scores on specific tests, I don't know how valuable these tests are in comparison to chatbot arena's elo ranking https://huggingface.co/spaces/lmsys/chatbot-arena-leaderboar...