GPT-4 Turbo: 31.0
Claude 3 Opus: 27.3
Mistral Large: 17.7
Mistral Medium: 15.3
Gemini Pro: 14.2
Qwen 1.5 72B Chat: 10.7
Claude 3 Sonnet: 7.6
GPT-3.5 Turbo: 4.2
Mixtral 8x7B Instruct: 4.2
Llama 2 70B Chat: 3.5
Qwen 1.5 14B: 3.1
Nous Hermes 2 Yi 34B: 1.5
Notes: 0-shot. Maximum possible is 100. Partial credit is given if the puzzle is not fully solved. There is only one attempt allowed per puzzle. In contrast, humans players get 4 attempts and a hint when they are one step away from solving a group. Gemini Advanced is not yet available through the API.
What I found interesting is how this benchmark reveals a large capabilities gap between the top, large models and the rest, in contrast to existing over-optimized benchmarks.