* run LLM evaluations systematically and at scale
* share the data with the public in a rigorous and transparent way
We use the UK government's Inspect [1] library to run the evaluations.
As soon as I saw this news on HN, I evaluated Mistral Small 3 on MATH [2] level 5 (hardest subset, 1,324 questions). I get an accuracy of 0.45 (± 0.011). We sample the LLM 8 times for each question, which lets us obtain less noisy estimates of mean accuracy, and measure the consistency of the LLM's answers. The 1,324*8=10,584 samples represent 8.5M tokens (2M in, 6.5M out).
You can see the full transcripts here in Inspect’s interactive interface: https://epoch.ai/inspect-viewer/484131e0/viewer?log_file=htt...
Note that MATH is a different benchmark from the MathInstruct [3] mentioned in the OP.
It's still early days for Epoch AI's benchmarking work. I'm developing a systematic database of evaluations run directly by us (so we can share the full details transparently), which we hope to release very soon.
[0]: https://epoch.ai/
[1]: https://github.com/UKGovernmentBEIS/inspect_ai